---
title: Capacity planning and overload recovery
url: https://doc.liz6.com/en/theory/03-queueing-theory/11-capacity-and-overload-recovery
locale: en
area: theory
tags:
- Queueing theory
- Theory
date: 2026-09-10
modified: 2026-09-10
description: Capacity planning must address normal latency, burst recovery and extra work from failures and retries. Average arrivals below average capacity are only the starting point.
---

# Capacity planning and overload recovery

Capacity planning must address normal latency, burst recovery and extra work from failures and retries. Average arrivals below average capacity are only the starting point.

## Backlogs drain using net headroom

A deterministic fluid approximation treats jobs as divisible work. With processing capacity $C$ and incoming work rate $\lambda$:

$$\dot Q=\lambda-C\text{ when }Q>0;\qquad
\dot Q=\max(0,\lambda-C)\text{ when }Q=0.$$

An initial backlog $Q_0$ drains in $Q_0/(C-\lambda)$ when subsequent traffic is constant and $\lambda<C$. The expression $Q_0/C$ requires no new arrivals. Clearance time concerns disappearance of the entire backlog; it is not the FCFS wait of one job at the current tail, because later arrivals normally join behind that job.

With normal traffic 60/s and capacity 100/s, a 120/s burst lasting 30 s adds 600 jobs. Returning to normal leaves 40/s spare capacity and takes another 15 s to clear. At capacity 70/s, the same burst accumulates 1500 jobs and only 10/s remains for recovery, taking 150 s.

## Retries consume headroom

For at most $m$ attempts, independent fixed failure probability $p$ per attempt, and a new attempt only after failure:

$$E[A]=1+p+\cdots+p^{m-1}=\frac{1-p^m}{1-p}\quad(p\ne1).$$

Only unlimited attempts give $1/(1-p)$. Real failures can be correlated and load-dependent, so constant $p$ is a work approximation. Retry delays change arrival structure; expected amplification is not a literal timeline of attempts.

The experiment starts with $Q(0)=0$ and observes through 240 s. Backlog counts attempt work; arrival rates already include expected retry amplification.

**Capacity planning and overload recovery · Experiment**

Deterministic fluid approximation: original arrivals 60/s, rising to 120/s during 20–50 s. At most three attempts with independent fixed failure probability p give instantaneous work amplification 1+p+p². No backoff or rejection; not a literal retry process or stochastic tail model.


The experiment allows at most three attempts. With no retries, its 20–50 s burst clears by 65 s. At capacity 100/s and $p=0.3$, amplification is 1.39, normal attempts 83.4/s, and burst backlog 2004. Clearing after the burst takes about 120.72 s. At $p=0.5$, normal attempts reach 105/s, leaving no positive recovery headroom; backlog keeps rising after the burst.

## Turn analysis into a capacity review

1. Define the business unit and success condition. Separate original requests, attempts, admission, success, cancellation and expiration; fast errors are not useful completed work.
2. Measure bottleneck service demand, variability and tails by class. CPU, connections, locks and remote quotas can constrain different stages.
3. Estimate stationary waiting under an appropriate model, then test with open or closed demand matching actual behavior. Include model error and measurement uncertainty in headroom.
4. Replay bursts, slow dependencies, instance failures and retries. Examine recovery time, queue age, rejection and successful latency throughout recovery.
5. Specify admission, tenant isolation, stale work, retry budgets and scaling activation delays. Backlog continues accumulating until added capacity becomes usable.

[AWS's backlog article](https://d1.awsstatic.com/builderslibrary/pdfs/avoiding-insurmountable-queue-backlogs.pdf) provides experience with isolation and recovery; [Google SRE](https://sre.google/sre-book/handling-overload/) emphasizes resource signals and variable request costs. No universal 70% or 80% utilization threshold follows.

## Connect to control and messaging

Queueing theory describes demand and waiting; control theory describes how observed state changes actions. Queue age, backlog and resource utilization can supply feedback, while sampling, smoothing, scaling delay and saturation affect recovery. Continue with [control in computing systems](../02-control-theory/09-control-in-computing-systems.md) to place this workload model inside a dynamic decision process.

Messaging also involves durability, duplicates, partitions and replay. A FIFO sketch cannot replace those semantics. Continue with [message semantics](../../distributed-systems/07-messages-and-streams/01-message-semantics.md) and [Kafka architecture](../../distributed-systems/07-messages-and-streams/02-kafka-architecture.md). Further theory includes renewal processes, priority queues, closed-network mean-value analysis, network calculus, heavy-traffic limits and state-dependent service. This series does not claim to exhaust queueing theory.

## Check your understanding

1. Backlog is 10000, capacity 1000/s and ongoing traffic 900/s. Why does clearance take 100 s instead of 10 s?

<details><summary>Reasoning</summary>

Each second removes 1000 jobs but adds 900, a net reduction of 100. Thus 10000/100=100 s. Ten seconds requires incoming work to stop. Changing traffic, capacity or job costs requires integrating the new work balance.

</details>

2. Does a larger queue or more consumers guarantee overload recovery?

<details><summary>Reasoning</summary>

Neither guarantees it. More queue space permits more waiting; more consumers help only if they increase usable bottleneck capacity. An already saturated dependency may suffer extra contention. Identify the resource and deadline objective, then validate failure and recovery behavior.

</details>
