Essential Googleby Zhongyuan Wang
Essential Google by Zhongyuan Wang
OutlineMotivation &GoalsProblems Solution: BigTableFile Systemvs Database Google's Database: Google Base· Conclusion & Outlook·References
Outline • Motivation & Goals • Problems • Solution:BigTable • File System vs Database • Google’s Database:Google Base • Conclusion & Outlook • References
OutlineMotivation&GoalsProblems Solution: BigTableFile System vs Database Google's Database: Google BaseConclusion & OutlookReferences
Outline • Motivation & Goals • Problems • Solution:BigTable • File System vs Database • Google’s Database:Google Base • Conclusion & Outlook • References
MotivationLots of (semi-)structured data at Google- URLs:: Contents, crawl metadata, links, anchors, pagerank,.- Per-user data:. User preference settings, recent queries/searchresults, .- Geographic locations: Physical entities (shops, restaurants, etc.), roads, satellite imagedata, user annotations,Scale is large- Billions of URLs, many versions/page (~20K/version)Hundreds of millions of users, thousands of q/sec100TB+ of satellite image data
Motivation • Lots of (semi-)structured data at Google – URLs: • Contents, crawl metadata, links, anchors, pagerank,. – Per-user data: • User preference settings, recent queries/search results, . – Geographic locations: • Physical entities (shops, restaurants, etc.), roads, satellite image data, user annotations, . • Scale is large – Billions of URLs, many versions/page (~20K/version) – Hundreds of millions of users, thousands of q/sec – 100TB+ of satellite image data
GoalsWant asynchronous processes to be continuously updatingdifferent pieces of data- Want access to most current data at any timeNeed to support:- Very high read/write rates (millions of ops per second)- Efficient scans over all or interesting subsets of data- Efficient joins of large one-to-one and one-to-many datasets: Often want to examine data changes over time- E.g. Contents of a web page over multiple crawls
Goals • Want asynchronous processes to be continuously updating different pieces of data – Want access to most current data at any time • Need to support: – Very high read/write rates (millions of ops per second) – Efficient scans over all or interesting subsets of data – Efficient joins of large one-to-one and one-to-many datasets • Often want to examine data changes over time – E.g. Contents of a web page over multiple crawls