Load Balancing Strategies
Problem Explain what a load balancer is and design the traffic-distribution layer in front of a pool of application servers: which algorithm, at which layer, and how failures are detected.
Functional requirements
- Distribute incoming requests across a pool of backends.
- Health-check backends and route around unhealthy instances.
- Support adding/removing backends during deploys and autoscaling.
- Support session affinity where the application needs it.
- Terminate TLS and route by path/host where L7 features are required.
Non-functional requirements
- ~100k requests/sec across ~200 backends; ~1M concurrent connections.
- Routing overhead < 1 ms added latency at p99.
- Unhealthy-backend detection < 10 s (e.g. probe every 2 s, 3 consecutive failures) with no requests routed to it after.
- Load skew between hottest and median backend < 15%.
- The LB tier itself must be redundant: 99.99%+ availability, no single instance in the request path.
Key components
- Algorithms: round-robin (simple, ignores backend state), weighted round-robin (accounts for heterogeneous instance sizes), least-connections (adapts to real load and long-lived requests), IP-hash / consistent hashing (affinity and cache locality), latency-based / least-response-time (routes to the fastest backend).
- Health checks: active probes (HTTP endpoint) plus passive ejection on observed error rates.
- L4 vs L7 operation: L4 forwards TCP/UDP flows; L7 parses HTTP and can route on path, host and headers.
- Service discovery feeding the backend pool.
- Redundancy: multiple LB instances behind DNS/anycast, with a floating VIP or ECMP.
- Connection draining on deregistration so in-flight requests complete.
Deep dives / trade-offs
- Round-robin vs least-connections: round-robin is stateless and perfectly fair only when every request costs the same. With mixed workloads (a 5 ms cache hit next to a 2 s report), round-robin piles slow requests onto unlucky backends while least-connections naturally adapts. Least-connections requires per-backend state, which is why it's harder across a multi-instance LB tier.
- L4 vs L7: L4 is fast, cheap, protocol-agnostic and preserves end-to-end TLS, but it cannot retry a failed request or route by URL. L7 unlocks path routing, header-based canaries, retries, and TLS termination — at the cost of CPU (it parses every request) and of terminating encryption at the edge.
- Health checks are subtle: too aggressive and a GC pause ejects a healthy node, cascading load onto the rest and triggering more ejections. Too lax and you serve errors for 30 s. Deep checks (does the DB respond?) can eject the entire fleet simultaneously when a shared dependency blips — which is why shallow liveness and deep readiness are separated.
- Session affinity vs even distribution are in direct conflict: sticky sessions give cache locality but break even balance and cause a thundering reconnect when a node dies. Stateless backends with external session state avoid the whole trade — prefer it when you can.
- The LB is the textbook SPOF: whatever you put in front of the servers now needs its own redundancy story (DNS round-robin, anycast, or a hardware pair) and its own capacity plan for 1M concurrent connections.
- Autoscaling interaction: adding a backend at peak sends it a full share of traffic instantly, against a cold cache and an empty connection pool — slow-start/warm-up weighting exists for exactly this.
asked …