Design a Cache That Avoids Concurrent Duplicate Fetches

Problem

Design a caching layer where the critical requirement is to avoid issuing concurrent fetches for the same resource (request coalescing / thundering-herd protection).

Requirements

Functional:

  • Cache responses with TTL
  • On a miss, only ONE fetch per key; other callers wait for it
  • Serve stale-while-revalidate
  • Purge support

Non-functional:

  • Very high concurrent request rate
  • Protect the origin from stampedes
  • Low added latency

Discussion points

  1. In-flight request map (single-flight) keyed by resource
  2. Locking vs futures/promises for waiters
  3. Stale-while-revalidate and negative caching
  4. TTL and purge propagation across edge nodes
  5. Failure handling when the single fetch fails
added …
LeaderboardSalaryAccount