A cache keeps a copy of data closer to where it is used, so reads skip slower work: a database query, a call to another service, a trip across the world. It is the first answer to a read-heavy design, and the follow-up questions are always about what happens when the copy and the truth disagree.
Building block 3 of 6 in system design building blocks
When it comes up
- Reads far outnumber writes, as your estimate should show.
- The same items are requested again and again: popular posts, a celebrity's profile, a trending query.
- Latency targets the database alone cannot meet.
- Static or slowly changing content served worldwide, which belongs on a CDN.
- An expensive computation produces the same result for many requests.
The core idea
Caches sit at several layers: in the client, at the edge (a CDN), inside the application process, in a shared cache such as Redis, and in the database's own memory. The most common pattern is cache-aside: on a read, check the cache; on a miss, read the database and store the result with a time to live. On a write, update the database and delete the cached copy, so the next read fetches the new value.
The hard parts are invalidation and load. Every cached item needs a story for going stale: a TTL, an explicit delete on write, or a versioned key. Popular items create hot keys, and when one expires, many requests can miss at once and stampede the database; adding random jitter to TTLs and letting only one request refill a key both help. Eviction (usually least recently used) decides what goes when memory is full.
A sketch
import json, random
def read_profile(user_id, cache, db, ttl=300):
key = f"profile:{user_id}"
hit = cache.get(key)
if hit is not None:
return json.loads(hit) # no database work at all
row = db.load_profile(user_id) # miss: go to the source of truth
cache.set(key, json.dumps(row), ex=ttl + random.randint(0, 60)) # jitter spreads expiries
return row
def update_profile(user_id, fields, cache, db):
db.save_profile(user_id, fields) # the truth first
cache.delete(f"profile:{user_id}") # then drop the stale copyTrade-offs to name
- Freshness against hit rate: a longer TTL serves more from cache and serves staler data.
- Deleting on write against updating on write: deleting is simpler and avoids racing writers overwriting each other.
- Read-your-own-writes: after a user edits something, they should see their change even if others see a cached copy briefly.
- Memory cost against database load.
- Write-behind caching (fast writes) against the risk of losing writes that have not reached the database.
Common mistakes
- Adding a cache with no plan for invalidation.
- Letting popular keys expire all at once and stampede the database.
- Caching personalized data under a shared key, so one user sees another's data.
- Claiming a cache helps write throughput; it mainly helps reads.
- No plan for the cache being down: the database must survive the extra load, or the system must shed some.
How to explain it out loud
Describe the read path and the write path separately, in that order, and say the hit rate you expect and why: "Most reads are for a small set of popular items, so I expect a high hit rate." Tie it back to your estimate.
Then state the staleness you will accept, per use: "Up to five minutes of staleness is fine for a profile page, but not for an account balance." Choosing different answers for different data is exactly what Devana's rubric credits under trade-offs, and it is the follow-up interviewers ask most.
Practice questions
These come from Devana's question bank, in the order to try them. Each one starts a voice mock interview with Josh, Devana's AI interviewer, on that question, so you practice explaining the approach out loud as well as getting it right.
- Practice
Design media hosting for user uploads
mediumReddit · Pull-through caching with sudden popularity spikes.
- Practice
Design edge caching for dynamic sites
hardVercel · Caching pages that change, at the edge.
- Practice
Design a caching layer over object storage
hardDatabricks · A cache that has to stay consistent with its source.
- Practice
Design Google Search autocomplete
hardGoogle · Precomputed answers for very hot prefixes.