Caching

Updated October 7, 2026 · By the Devana Team

A cache keeps a copy of data closer to where it is used, so reads skip slower work: a database query, a call to another service, a trip across the world. It is the first answer to a read-heavy design, and the follow-up questions are always about what happens when the copy and the truth disagree.

Building block 3 of 6 in system design building blocks

When it comes up

  • Reads far outnumber writes, as your estimate should show.
  • The same items are requested again and again: popular posts, a celebrity's profile, a trending query.
  • Latency targets the database alone cannot meet.
  • Static or slowly changing content served worldwide, which belongs on a CDN.
  • An expensive computation produces the same result for many requests.

The core idea

Caches sit at several layers: in the client, at the edge (a CDN), inside the application process, in a shared cache such as Redis, and in the database's own memory. The most common pattern is cache-aside: on a read, check the cache; on a miss, read the database and store the result with a time to live. On a write, update the database and delete the cached copy, so the next read fetches the new value.

The hard parts are invalidation and load. Every cached item needs a story for going stale: a TTL, an explicit delete on write, or a versioned key. Popular items create hot keys, and when one expires, many requests can miss at once and stampede the database; adding random jitter to TTLs and letting only one request refill a key both help. Eviction (usually least recently used) decides what goes when memory is full.

A sketch

import json, random

def read_profile(user_id, cache, db, ttl=300):
    key = f"profile:{user_id}"
    hit = cache.get(key)
    if hit is not None:
        return json.loads(hit)                    # no database work at all
    row = db.load_profile(user_id)                # miss: go to the source of truth
    cache.set(key, json.dumps(row), ex=ttl + random.randint(0, 60))  # jitter spreads expiries
    return row

def update_profile(user_id, fields, cache, db):
    db.save_profile(user_id, fields)              # the truth first
    cache.delete(f"profile:{user_id}")            # then drop the stale copy

Trade-offs to name

  • Freshness against hit rate: a longer TTL serves more from cache and serves staler data.
  • Deleting on write against updating on write: deleting is simpler and avoids racing writers overwriting each other.
  • Read-your-own-writes: after a user edits something, they should see their change even if others see a cached copy briefly.
  • Memory cost against database load.
  • Write-behind caching (fast writes) against the risk of losing writes that have not reached the database.

Common mistakes

  • Adding a cache with no plan for invalidation.
  • Letting popular keys expire all at once and stampede the database.
  • Caching personalized data under a shared key, so one user sees another's data.
  • Claiming a cache helps write throughput; it mainly helps reads.
  • No plan for the cache being down: the database must survive the extra load, or the system must shed some.

How to explain it out loud

Describe the read path and the write path separately, in that order, and say the hit rate you expect and why: "Most reads are for a small set of popular items, so I expect a high hit rate." Tie it back to your estimate.

Then state the staleness you will accept, per use: "Up to five minutes of staleness is fine for a profile page, but not for an account balance." Choosing different answers for different data is exactly what Devana's rubric credits under trade-offs, and it is the follow-up interviewers ask most.

Practice questions

These come from Devana's question bank, in the order to try them. Each one starts a voice mock interview with Josh, Devana's AI interviewer, on that question, so you practice explaining the approach out loud as well as getting it right.

  • Design media hosting for user uploads

    mediumReddit · Pull-through caching with sudden popularity spikes.

    Practice
  • Design edge caching for dynamic sites

    hardVercel · Caching pages that change, at the edge.

    Practice
  • Design a caching layer over object storage

    hardDatabricks · A cache that has to stay consistent with its source.

    Practice
  • Design Google Search autocomplete

    hardGoogle · Precomputed answers for very hot prefixes.

    Practice

Prove it in a mock interview

A 15-minute mock interview on a question that is not on the practice list, scored out of 100. Score 70 or more and caching is marked proven on your roadmap. It counts as one of your interviews: the Free plan has 3 a month, no card needed.