Message queues and async work

Updated October 7, 2026 · By the Devana Team

A queue lets one part of a system hand work to another without waiting for it, so slow or bursty work (sending notifications, processing uploads, recording usage) stops blocking the request that caused it. The questions that follow are about guarantees: what happens when a consumer crashes, a message arrives twice, or the backlog keeps growing.

Building block 6 of 6 in system design building blocks

When it comes up

  • Work that can finish after the response: emails, thumbnails, analytics, search indexing.
  • Fan-out to many recipients, such as notifying every follower.
  • Smoothing spikes so a slow system downstream is not overwhelmed.
  • Decoupling services so one can fail or deploy without taking down another.
  • Retrying work that fails for temporary reasons.

The core idea

Producers write messages; consumers read and process them. Most queues deliver at least once: if a consumer crashes before acknowledging a message, it is delivered again. So consumers must be idempotent, meaning processing a message twice has the same effect as once, usually by recording a message id in the same transaction as the work. That combination is what people mean by effectively exactly once.

Order is only kept within a partition, so put everything that must stay in order under one key. Retries need backoff and a limit; a message that keeps failing goes to a dead-letter queue for someone to inspect, instead of blocking everything behind it. Watch the backlog and the age of the oldest message: those are the alerts that tell you consumers are falling behind.

A sketch: an idempotent consumer

def award_points(message, db):
    """Delivery is at least once: the same event can arrive twice."""
    with db.transaction() as tx:
        if tx.exists("processed_events", message["event_id"]):
            return                                   # already applied: do nothing
        tx.add_points(message["member_id"], message["points"])
        tx.insert("processed_events", message["event_id"])  # same transaction as the effect

Because the effect and the record of it commit together, a crash either keeps both or neither, and a redelivered message finds the record and stops.

Trade-offs to name

  • At-least-once delivery with idempotent consumers against the cost and limits of exactly-once features.
  • Strict ordering (one partition per key, less parallelism) against throughput.
  • A queue that deletes messages once processed against a log that keeps them for replay.
  • Synchronous simplicity (the user gets the result now) against async resilience (the user sees "processing").

Common mistakes

  • Assuming exactly-once delivery and writing consumers that double-count.
  • No dead-letter queue, so one bad message blocks everything behind it.
  • No alert on backlog or message age until users notice.
  • Processing in parallel and losing an order the business depends on.
  • Putting a queue on the path where the user needs the answer immediately.

How to explain it out loud

Split the design into the path that has to be fast and the work that can wait: "The request records the order and returns. Receipts, loyalty points and analytics go through a queue." That one sentence explains why the queue is there, which is the part interviewers listen for.

Then volunteer the failure cases: a consumer crashing mid-message, a duplicate, a poison message, a backlog that grows for an hour. Explaining how each is handled is what Devana's rubric scores under reliability and trade-offs, and it is where strong system design answers pull ahead.

Practice questions

These come from Devana's question bank, in the order to try them. Each one starts a voice mock interview with Josh, Devana's AI interviewer, on that question, so you practice explaining the approach out loud as well as getting it right.

  • Design notification fan-out

    mediumGoogle · Queue-driven fan-out to a very large audience.

    Practice
  • Design usage metering and spend limits

    mediumTwilio · Async counting that still has to enforce a limit.

    Practice
  • Design continuous data loading

    mediumSnowflake · Ingesting a steady stream without loss or duplicates.

    Practice
  • Design a managed message queue

    hardAmazon · The queue itself: durability, visibility and ordering.

    Practice

Prove it in a mock interview

A 15-minute mock interview on a question that is not on the practice list, scored out of 100. Score 70 or more and message queues and async work is marked proven on your roadmap. It counts as one of your interviews: the Free plan has 3 a month, no card needed.