# Avoiding Insurmountable Queue Backlogs

## Metadata
- Author: [[aws.com]]
- Full Title: Avoiding Insurmountable Queue Backlogs
- Category: #articles
- Summary: Queues make systems more reliable, but a buildup of messages can cause long delays. Track message age and limit or separate excess work to prevent backlogs. Use delay and dead-letter queues to manage work that must wait or cannot be processed.
- URL: https://builder.aws.com/content/3EuRcgkTP1MI0c7zM8W6HL3WIqA/avoiding-insurmountable-queue-backlogs
## Highlights
- In queueing theory, the behavior of queues when they are short is relatively uninteresting. After all, when a queue is short, everyone is happy. It’s only when the queue is backlogged, when the line to an event goes out the door and around the corner,that people start thinking about throughput and prioritization. ([View Highlight](https://read.readwise.io/read/01m3mwrxxgpe2wy1y72mb1zq2y))
- In this article, I discuss strategies we use at Amazon to deal with queue backlog scenarios – design approaches we take to drain queues quickly and to prioritize workloads. Most importantly, I describe how to prevent queue backlogs from building up in the first place. ([View Highlight](https://read.readwise.io/read/01m3mwseh7chy7sqd0q5fzy7x2))
- In the first half, I describe scenarios that lead to backlogs, and in the second half, I describe many approaches used at Amazon to avoid backlogs or deal with them gracefully. ([View Highlight](https://read.readwise.io/read/01m3mwsmfc517zky27bew87jp6))
- The duplicitous nature of queues
Queues are powerful tools for building reliable asynchronous systems. Queues allow one system to accept a message from another system, and persist the message until it is fully processed, even in the face of long outages, server failures, or problems with dependent systems. Rather than dropping messages when a failure occurs, the queue re-drives the messages until they are successfully processed. In the end, a queue increases a system’s durability and availability, at the price of occasional increased latency due to retries. ([View Highlight](https://read.readwise.io/read/01m3mwtfs80x03s43b9zdt0n4z))
- In a queue-based system, when processing stops but messages keep arriving, the message debt can accumulate into a large backlog, driving up processing time. Work can be completed too late for the results to be useful, essentially causing the availability hit that queueing was meant to guard against. ([View Highlight](https://read.readwise.io/read/01m3mwvxaq8cvn17a87a9hwf2y))
- if a failure or unexpected load pattern causes the arrival rate to exceed the processing rate, it quickly flips into a more sinister operating mode. In this mode, the end-to-end latency grows higher and higher, and it can take a great deal of time to work through the backlog to return to the fast mode. ([View Highlight](https://read.readwise.io/read/01m3mwwex9bbgxptsgkqf8ybj0))
- With a durable queue, your request can be re-driven if your function fails the first time. ([View Highlight](https://read.readwise.io/read/01m3mx6ek1k229jytczvpxdp1q))
- In queue-based systems, a component produces data by putting messages into the queue, and another component consumes that data by periodically asking for messages, processing messages, and finally deleting them once it’s done. ([View Highlight](https://read.readwise.io/read/01m3mx8519q5pbv4nkpbp7sdt2))