Design an A/B Testing Framework
Problem Design a framework that consistently buckets users into experiment variants (A/B/n) for feature rollout, so the same user reliably lands in the same variant across requests and services.
Functional requirements
- Deterministically assign a user to a variant for a given experiment.
- Support n-way splits and arbitrary traffic allocation (e.g. 5% → 50% ramp).
- Run many experiments concurrently without them correlating with each other.
- Support overrides (force a user into a variant for QA) and sticky assignments.
- Expose the assignment to every service handling the request.
- Log exposures for downstream analysis.
Non-functional requirements
- ~100M users, ~500 concurrent experiments; ~200k assignment evaluations/sec.
- Assignment must be O(1), local, and add < 1 ms — it runs on every request in every service, so a network call is not an option.
- Zero cross-service disagreement: two services computing the variant for the same user and experiment must always agree, without sharing state.
- Ramping an experiment from 5% to 10% must not move any of the original 5% into a different variant.
- Config propagation < 30 s; exposure logging is fire-and-forget and must never block the request.
Key components
- Assignment function: variant = bucket(hash(experiment_id + ":" + user_id)) — a stable identifier (user ID or device ID) salted with the experiment key, hashed (murmur/xxhash) into a large space, then mapped to buckets.
- Bucket space: hash into a large fixed number of buckets (e.g. 10,000) and assign contiguous bucket ranges to variants, so allocation changes move only the boundary ranges.
- Config service holding experiment definitions (variants, allocations, targeting, status), pushed to a local in-process cache in every service.
- Override store: an explicit user→variant map checked before the hash, for QA and sticky rollouts.
- Exposure logging: async emit to Kafka on first evaluation, feeding the analysis pipeline.
- SDK embedded in each service so assignment is a pure local computation.
Deep dives / trade-offs
- Consistent hashing / range-based buckets vs plain modulo: with variant = hash(user) % n, changing n from 2 to 3 reshuffles essentially every user — an in-flight experiment's population changes underneath it, invalidating the results and giving users a jarring UI flip. Mapping into a large fixed bucket space and assigning ranges means a 5%→10% ramp only pulls in new buckets and leaves the existing cohort untouched. This is the crux of the question: the requirement isn't balance, it's stability under configuration change.
- Salting per experiment: hashing on user_id alone means every experiment splits the population identically — the same users are always in "A" for every test, so experiments correlate and their effects become inseparable. Hashing experiment_id + user_id decorrelates them. Discuss layered/mutually-exclusive experiments when two tests genuinely conflict, and why layers need their own salt.
- Local vs remote assignment: a central assignment service gives perfect auditability and instant kill-switches, but adds a network call to every request in every service and becomes a fleet-wide SPOF. A pure function plus pushed config gives O(1) local evaluation and no shared fate — at the cost of config propagation lag, during which two services can briefly disagree. Bound the lag and make the function deterministic.
- Identity choice: device ID works pre-login but splits a user across devices and resets on reinstall; user ID is stable but unavailable to logged-out traffic. An experiment spanning the login boundary can flip a user's variant mid-session — decide the stitching rule up front.
- Overrides and stickiness: an override store is a lookup on the hot path, so keep it small and cached, and be clear that overridden users must be excluded from analysis or they bias the result.
- Exposure logging must record the assignment at decision time, not be recomputed later — config changes make retroactive recomputation lie.
asked …