Design a Dish Recommender System for Zomato
Problem Design a recommender system that suggests dishes — not just restaurants — to a user on a food-delivery app's home feed and search surface.
Functional requirements
- Return a personalized ranked list of dishes for a given user, location, and time of day.
- Respect hard constraints: only dishes from restaurants currently open and deliverable to the user's address.
- Handle cold-start for brand-new users and newly listed dishes.
- Keep the feed diverse — not ten variants of the same biryani.
Non-functional requirements
- ~50-100M users; tens of millions of dishes across hundreds of thousands of restaurants.
- ~10-20k feed requests/sec at peak meal hours; p99 end-to-end under ~200ms.
- Candidate generation must cut millions of dishes to a few hundred in tens of milliseconds.
- Embeddings and ranking models retrained daily; availability and popularity features refreshed in near real time (seconds to minutes).
Key components
- Objective definition first: engagement (clicks) vs. order conversion vs. long-term retention/GMV — the choice determines the labels you train on.
- Candidate generation: collaborative filtering / two-tower embeddings over order history, content-based signals from dish and cuisine attributes, plus popularity and trending recall sources, unioned together.
- Ranking model blending relevance, freshness, price sensitivity, rating, restaurant availability, and predicted delivery time.
- Feature pipeline and store: user taste profile from order history, dish/restaurant aggregates, session context — with train/serve consistency.
- Filtering and business layer: serviceability, stock-outs, dietary preferences, and diversity/de-duplication rules applied post-ranking.
- Evaluation: offline recall@k and NDCG, then A/B testing against a popularity-based baseline on conversion and retention.
Deep dives / trade-offs
- The candidate-generation/ranking split: why you cannot score millions of dishes per request, and how multi-source recall trades coverage against latency.
- Cold-start: popularity and location priors for new users, content/attribute embeddings for new dishes, and exploration to avoid a rich-get-richer loop.
- Objective mismatch: optimizing clicks yields clickbait dishes, conversion favours cheap familiar orders — how to weigh a multi-objective ranking.
- Position and popularity bias in logged training data, since the model learns from what it previously chose to show.
asked …