Design Google's Web Crawler
Problem
Design a distributed web crawler that can crawl billions of web pages and keep the index fresh.
Requirements
Functional:
- Crawl the entire web starting from a seed set of URLs
- Avoid re-crawling recently crawled pages
- Respect robots.txt
- Handle duplicate content (same page, different URL)
Non-functional:
- Crawl 1 billion pages/day
- Politeness: max 1 request/sec per domain
- Freshness: important pages re-crawled every 24h
Discussion points
- URL frontier design (priority queue + politeness buckets)
- DNS resolution caching
- Distributed deduplication (Bloom filters, simhash)
- HTML parsing and link extraction pipeline
- How to prioritise re-crawl frequency (PageRank signal)
added …