Design a Food Delivery Platform

Problem Design the high-level architecture for a food-delivery platform covering restaurant listings, menus, order placement, and live delivery tracking.

Functional requirements

  • Location-based restaurant discovery and search.
  • Browse a restaurant's menu with live item availability.
  • Place an order and pay; receive an ETA.
  • Restaurant accepts and progresses the order; a delivery partner is assigned.
  • Track delivery live and receive status notifications.

Non-functional requirements

  • ~20M DAU; ~2M orders/day with ~60% landing in a 3-hour dinner window → ~1,500 orders/sec at peak.
  • Discovery/browse traffic ~50x orders → ~75k QPS reads at peak; listing p99 < 300 ms.
  • ~300k concurrent delivery partners pinging every 4-5 s → ~70k location writes/sec.
  • Order placement p99 < 500 ms; tracking freshness < 5 s.
  • Orders and payments strongly consistent and durable; listings and location may be eventually consistent.
  • 99.95% availability; a discovery outage loses browsing, an order-service outage loses revenue directly.

Key components

  • Restaurant/Catalog service: Restaurants, MenuItems; read-mostly, cached aggressively in Redis with images on CDN.
  • Search service: geo-indexed restaurant search (geohash / S2 / spatial index) plus text relevance, backed by Elasticsearch fed via CDC from the catalog.
  • Order service: owns the lifecycle state machine (placed → accepted → preparing → out for delivery → delivered), on a relational store sharded by order_id, emitting state events.
  • Delivery/Logistics service: assigns delivery partners from the geo-index, tracks live location.
  • Location store: in-memory KV/geo store for real-time partner positions — never the primary DB.
  • Payment service with an external PSP, idempotency keys, saga compensation.
  • Kafka carrying order-state events; REST/gRPC for synchronous request/response between services.
  • Core entities: Restaurant, MenuItem, User, Order, OrderItem, DeliveryPartner, Assignment.

Deep dives / trade-offs

  • Storage per workload: orders need ACID transactions and go in a relational store; live partner locations at 70k writes/sec go in an in-memory store with TTL and are never durably persisted per-ping; the catalog is read-mostly and belongs behind a cache; search needs an inverted + geo index. The lesson is one store per access pattern, and being able to justify each.
  • Service decomposition and the transaction that spans it: creating the order, authorizing payment, and reserving restaurant capacity now cross three services. 2PC is blocking and unavailable at this scale; a saga with compensating actions (release the reservation, void the auth) is the standard answer. Walk the partial-failure case where payment succeeds and the restaurant then rejects.
  • Sync vs event-driven: order placement must be synchronous (the user is waiting on a real answer), but status fan-out, notifications, and analytics must be event-driven or they add latency to and shared fate with the critical path. State which calls are which and why.
  • Geo-indexing: geohash is simple and prefix-searchable, but cell boundaries split nearby points, so you query neighbours too. Quadtrees adapt to density (dense metros vs sparse suburbs) at higher update cost. The read/write ratio decides it — restaurant locations are near-static, partner locations are not, so the two indexes have opposite requirements and should not be the same index.
  • Read/write asymmetry: at 50:1 the browse path dominates capacity planning and is nearly all cacheable; the order path is small, uncacheable and the only part that must not lose data. Size and isolate them separately so a browse-traffic spike cannot exhaust the connection pool the order path needs.
asked …
LeaderboardSalaryAccount
Design a Food Delivery Platform · 2dbi