Design Live Order Tracking and ETA Prediction
Problem Design a real-time order lifecycle tracking and ETA prediction system for a food-delivery app: the customer watches the order move through its states and sees a live, continuously-updating delivery estimate.
Functional requirements
- Track an order through placed → confirmed → preparing → picked → delivered.
- Show the delivery agent's live location on the customer's map.
- Recompute and push a dynamic ETA as the agent moves and as conditions change.
- Estimate restaurant preparation time and fold it into the ETA.
- Notify the customer on each state transition.
Non-functional requirements
- 1.5M concurrent orders at dinner peak, each with a customer watching the map.
- Agent location refresh every 5 s → at ~300k concurrent agents that is ~60k location writes/sec, fanned out to ~1.5M WebSocket subscribers.
- ETA accuracy within ±3 minutes for 90% of orders.
- Location-to-map end-to-end latency < 5 s; state transition push < 2 s.
- Order state must never regress or be lost; a dropped location ping is acceptable.
Key components
- Ingestion pipeline: agent app → gateway → Kafka (partitioned by agent_id) → stream processor that updates the current position in an in-memory geo store and forwards to fan-out.
- WebSocket/SSE gateway tier: holds 1.5M persistent connections, sharded by order_id, with a pub/sub backplane (Redis/Kafka) so any gateway node can receive updates for the orders it hosts.
- Order state service: event-sourced state machine — transitions appended to an immutable log, current state as a projection.
- ETA service: baseline distance/speed along the road network, plus a restaurant-delay signal (live kitchen backlog), plus an ML model over historical features (time of day, traffic, restaurant, agent, weather).
- Feature store + model serving for the ML path, with a cheap deterministic fallback.
- Push/notification service driven off the state-transition event stream.
Deep dives / trade-offs
- WebSocket session management at 1.5M concurrent: connection state, heartbeats, sticky routing vs a pub/sub backplane, graceful reconnect with resume-from-sequence, and the thundering-herd problem when a gateway node dies and 50k clients reconnect at once. Contrast with long-polling for low-end devices.
- ETA modelling: naive haversine/speed is trivially wrong on real roads. Road-network routing with live traffic is better but expensive per call at this fan-out — so cache per route segment and only recompute on meaningful movement. The ML model improves accuracy but needs a fallback when features are stale, and needs monitoring for drift. Discuss why ETA should be monotonic-ish: an estimate that jumps around erodes trust more than being a minute off.
- Event sourcing the state machine: gives a replayable audit trail and lets tracking, analytics and support all derive from one log; costs you projection lag and the need for idempotent, out-of-order-tolerant consumers (restaurant and agent apps retry aggressively).
- Agent goes offline mid-delivery: distinguish GPS gap from app crash from a genuinely stalled delivery. Dead-reckon along the predicted route for a bounded window, degrade the map to "last known" rather than freezing, escalate to support past a threshold, and never let the ETA silently stop updating.
- Backpressure: at peak, prioritize state transitions over location pings — dropping a location update is invisible, dropping "delivered" is not.
asked …