Google File System (GFS)GFSMasterMasterClientGFSMasterClientC3COC1C1COChunkseversC3C4C3C4C5Master manages metadataData transfers happen directly between clients/chunkserversFiles broken into chunks (typically 64 MB)Chunks triplicated across three machines for safety
Google File System (GFS) • Master manages metadata • Data transfers happen directly between clients/chunkservers • Files broken into chunks (typically 64 MB) • Chunks triplicated across three machines for safety Replicas Master GFS Master GFS Master Client Client C0 C1 C0 C3 C4 C3 C1 C5 C3 C4 Chunksevers
atypicalMapReduceMapReduce: Easy-to-usecomputationprocessesmanyterabytesofdataonMany Google problems: “Process lots of data thousands of machinesMany kinds of inputs:Want to use easily hundreds or thousandsof CPUsMapReduce:framework that provides (for certain classes of problems)Automatic&efficientparallelization/distribution-Fault-tolerance,I/Oscheduling,status/monitoring-UserwritesMapandReducefunctionsHeavily used:~3000 jobs, 1000sof machine days each dayBigTablecanbeinputand/oroutputfor MapReducecomputations
MapReduce: Easy-to-use Cycles Many Google problems: “Process lots of data to produce other data” • Many kinds of inputs: • Want to use easily hundreds or thousands of CPUs • MapReduce: framework that provides (for certain classes of problems): – Automatic & efficient parallelization/distribution – Fault-tolerance, I/O scheduling, status/monitoring – User writes Map and Reduce functions • Heavily used: ~3000 jobs, 1000s of machine days each day BigTable can be input and/or output for MapReduce computations a typical MapReduce computation processes many terabytes of data on thousands of machines