Design an A/B Testing Framework

Problem Design a framework that consistently buckets users into experiment variants (A/B/n) for feature rollout, so the same user reliably lands in the same variant across requests and services.

Functional requirements

  • Deterministically assign a user to a variant for a given experiment.
  • Support n-way splits and arbitrary traffic allocation (e.g. 5% → 50% ramp).
  • Run many experiments concurrently without them correlating with each other.
  • Support overrides (force a user into a variant for QA) and sticky assignments.
  • Expose the assignment to every service handling the request.
  • Log exposures for downstream analysis.

Non-functional requirements

  • ~100M users, ~500 concurrent experiments; ~200k assignment evaluations/sec.
  • Assignment must be O(1), local, and add < 1 ms — it runs on every request in every service, so a network call is not an option.
  • Zero cross-service disagreement: two services computing the variant for the same user and experiment must always agree, without sharing state.
  • Ramping an experiment from 5% to 10% must not move any of the original 5% into a different variant.
  • Config propagation < 30 s; exposure logging is fire-and-forget and must never block the request.

Key components

  • Assignment function: variant = bucket(hash(experiment_id + ":" + user_id)) — a stable identifier (user ID or device ID) salted with the experiment key, hashed (murmur/xxhash) into a large space, then mapped to buckets.
  • Bucket space: hash into a large fixed number of buckets (e.g. 10,000) and assign contiguous bucket ranges to variants, so allocation changes move only the boundary ranges.
  • Config service holding experiment definitions (variants, allocations, targeting, status), pushed to a local in-process cache in every service.
  • Override store: an explicit user→variant map checked before the hash, for QA and sticky rollouts.
  • Exposure logging: async emit to Kafka on first evaluation, feeding the analysis pipeline.
  • SDK embedded in each service so assignment is a pure local computation.

Deep dives / trade-offs

  • Consistent hashing / range-based buckets vs plain modulo: with variant = hash(user) % n, changing n from 2 to 3 reshuffles essentially every user — an in-flight experiment's population changes underneath it, invalidating the results and giving users a jarring UI flip. Mapping into a large fixed bucket space and assigning ranges means a 5%→10% ramp only pulls in new buckets and leaves the existing cohort untouched. This is the crux of the question: the requirement isn't balance, it's stability under configuration change.
  • Salting per experiment: hashing on user_id alone means every experiment splits the population identically — the same users are always in "A" for every test, so experiments correlate and their effects become inseparable. Hashing experiment_id + user_id decorrelates them. Discuss layered/mutually-exclusive experiments when two tests genuinely conflict, and why layers need their own salt.
  • Local vs remote assignment: a central assignment service gives perfect auditability and instant kill-switches, but adds a network call to every request in every service and becomes a fleet-wide SPOF. A pure function plus pushed config gives O(1) local evaluation and no shared fate — at the cost of config propagation lag, during which two services can briefly disagree. Bound the lag and make the function deterministic.
  • Identity choice: device ID works pre-login but splits a user across devices and resets on reinstall; user ID is stable but unavailable to logged-out traffic. An experiment spanning the login boundary can flip a user's variant mid-session — decide the stitching rule up front.
  • Overrides and stickiness: an override store is a lookup on the hot path, so keep it small and cached, and be clear that overridden users must be excluded from analysis or they bias the result.
  • Exposure logging must record the assignment at decision time, not be recomputed later — config changes make retroactive recomputation lie.
asked …
LeaderboardSalaryAccount