Design Google's Web Crawler

Problem

Design a distributed web crawler that can crawl billions of web pages and keep the index fresh.

Requirements

Functional:

  • Crawl the entire web starting from a seed set of URLs
  • Avoid re-crawling recently crawled pages
  • Respect robots.txt
  • Handle duplicate content (same page, different URL)

Non-functional:

  • Crawl 1 billion pages/day
  • Politeness: max 1 request/sec per domain
  • Freshness: important pages re-crawled every 24h

Discussion points

  1. URL frontier design (priority queue + politeness buckets)
  2. DNS resolution caching
  3. Distributed deduplication (Bloom filters, simhash)
  4. HTML parsing and link extraction pipeline
  5. How to prioritise re-crawl frequency (PageRank signal)
added …
LeaderboardSalaryAccount