Design Cricbuzz / ESPNcricinfo
Problem Design a live cricket scoring system like Cricbuzz or ESPNcricinfo: model live matches, ingest ball-by-ball scoring updates, maintain player and team statistics, and serve real-time score queries to many concurrent viewers.
Requirements
recordDelivery(matchId, delivery) -> ScoreState— append a ball, recompute derived stategetScorecard(matchId) -> Scorecard— current score, wickets, overs, batsman/bowler figuressubscribe(matchId, listener)— push updates on every ball without pollinggetPlayerStats(playerId)/getTeamStats(teamId)— aggregated across matches- Corrections must be possible: a scorer can misrecord a ball
Core design
- Entities:
Match(teams, format, current innings, status),Innings(overs, runs, wickets, batting/bowling team),Over(a group of six legal deliveries),Delivery(bowler, striker, runs, extras type, wicket info, commentary),Player,Team. - Event sourcing is the right backbone here.
Deliveryrecords form an append-only log per match — the single source of truth. Every derived value (score, run rate, required rate, batsman strike rate, bowler economy, fall of wickets) is a fold over that log. Nothing overwrites a delivery. ScoreUpdateServiceappends the delivery and updates a materializedScoreStateincrementally, rather than replaying thousands of balls per read. The log stays authoritative; the materialized view is a cache that can always be rebuilt from it.- Observer pattern for fan-out: a live-scorecard view, push-notification service, and commentary feed all subscribe to a match and are notified per ball, decoupling the scoring write path from an open-ended set of consumers.
- Single writer per match — one scorer, or an ordered event queue keyed by match ID. Cricket state is inherently sequential (a wicket changes who is on strike, which changes how the next ball is attributed), so concurrent appends to one match are not merely a race, they are meaningless. Different matches are independent and shard cleanly by match ID.
- Extras are where the domain bites: wides and no-balls don't count toward the six legal deliveries in an over, byes/leg-byes credit the team but not the batsman, and a no-ball adds a free hit. The
Deliverymodel must carry extras as structured data, not just a runs integer, or these rules can't be computed.
Discussion points
- Corrections: since the log is append-only, a misrecorded ball is fixed with a compensating/reversal event, never a mutation — which preserves the audit trail and lets the scorecard be recomputed deterministically. This is the main payoff of event sourcing here.
- Read fan-out is the actual scale problem: one write per ball versus millions of concurrent readers. Discuss pushing via WebSocket/SSE to connected clients versus caching the scorecard at the edge with a short TTL for the polling long tail, and why the read path should never touch the event log directly.
- Trade-off: event-sourced log (auditable, replayable, corrections are natural, but every read needs a projection) vs. a mutable score row (trivial reads, no history, corrections destroy data).
- Idempotency: a scorer's client retrying a delivery submission must not double-count. Deliveries need a client-supplied ID so the append is idempotent.
- Player stats span matches, so they are a separate projection updated asynchronously off the delivery stream — keeping the live path fast and accepting eventual consistency on career aggregates.
- Ordering across the fan-out: subscribers must apply balls in sequence, so events need a monotonic per-match sequence number and clients need to detect gaps.
- Edge cases: rain/DLS revised targets, super overs, innings declarations, retired hurt, and substitutions — all of which change how derived state is computed mid-match.
asked …