A load balancer spreads requests across many identical servers, so you add capacity by adding machines and survive any one of them failing. It only works if those servers are stateless, keeping sessions and data in shared stores, which is why the two ideas arrive together in almost every design.
Building block 2 of 6 in system design building blocks
When it comes up
- Traffic beyond what one machine can serve.
- Availability targets that rule out a single point of failure.
- Deploys without downtime: take servers out of rotation, update them, put them back.
- Sudden spikes, where capacity has to grow automatically.
- Users in several regions who should reach the nearest one.
The core idea
A layer 4 balancer routes connections by address and port: fast and simple. A layer 7 balancer reads the request, so it can route by path or header, terminate TLS and retry failed requests, at more cost per request. It sends each request to a healthy server chosen by an algorithm: round robin, least connections, or a hash of a key when the same client or object should keep landing on the same server.
Health checks take failing servers out of rotation; connection draining lets a server finish its requests before it is removed. None of this works if a server holds state that others lack, so sessions move to a shared store or into signed tokens, uploads go to object storage, and any server can answer any request. Across regions, DNS or anycast sends users to a nearby region, and a balancer inside each region does the rest.
A sketch
import random
class Pool:
"""Least connections across healthy servers, ties broken at random."""
def __init__(self, servers):
self.open = {s: 0 for s in servers} # server -> open connections
self.healthy = set(servers)
def report_health(self, server, ok):
if ok:
self.healthy.add(server)
else:
self.healthy.discard(server) # out of rotation until it recovers
def pick(self):
live = [s for s in self.open if s in self.healthy]
if not live:
raise RuntimeError("no healthy servers")
fewest = min(self.open[s] for s in live)
return random.choice([s for s in live if self.open[s] == fewest])Trade-offs to name
- Layer 4 speed against layer 7 routing, TLS termination and retries.
- Sticky sessions (simple, but uneven load and lost sessions when a server dies) against stateless servers with a shared session store (an extra network hop).
- Autoscaling lag against the cost of keeping spare capacity running.
- Aggressive health checks (fast detection, but servers flapping in and out) against lenient ones (slow detection).
Common mistakes
- Drawing one load balancer and forgetting it is now the single point of failure; use a redundant pair or a managed service.
- Putting servers that keep state in memory behind round robin.
- Leaving out health checks and draining, so deploys drop requests.
- Scaling the web tier when the database is the real bottleneck.
- Retrying requests that are not safe to repeat, such as a payment without an idempotency key.
How to explain it out loud
Lead the walkthrough: "Requests arrive at a load balancer in front of a pool of stateless API servers. Sessions live in a shared cache, so any server can take any request, and I can add servers as traffic grows." Saying how state is handled before the interviewer asks is what driving the discussion means in Devana's rubric.
Then walk a failure: "If a server dies, health checks remove it within a few seconds and the others absorb its share." Reliability is a fifth of the design score, and most candidates never mention what breaks. Name the trade-off you chose, sticky or stateless, and why.
Practice questions
These come from Devana's question bank, in the order to try them. Each one starts a voice mock interview with Josh, Devana's AI interviewer, on that question, so you practice explaining the approach out loud as well as getting it right.
- Practice
Design for a demand spike during an event
mediumDoorDash · Absorbing a tenfold spike, and deciding what to shed.
- Practice
Design an upgrade with no downtime
mediumOracle · Draining, rolling and staying compatible mid-deploy.
- Practice
Design a global load balancer
hardGoogle · Routing users to regions and handling a region failing.