---
title: Finite queues and admission control
url: https://doc.liz6.com/en/theory/03-queueing-theory/08-finite-queues-and-admission
locale: en
area: theory
tags:
- Queueing theory
- Theory
date: 2026-09-10
modified: 2026-09-10
description: A finite queue replaces unbounded accumulation with an admission-versus-rejection tradeoff. Bounded occupancy means storage is bounded; a system rejecting most requests can still fail its business objectives.
---

# Finite queues and admission control

A finite queue replaces unbounded accumulation with an admission-versus-rejection tradeoff. Bounded occupancy means storage is bounded; a system rejecting most requests can still fail its business objectives.

## In M/M/1/K, K includes the job in service

Keep Poisson arrivals, independent exponential service and a single FCFS server, but allow at most $K$ jobs inside. Full arrivals are rejected immediately without retries or abandonment. For offered-load ratio $r=\lambda/\mu$, the finite-state stationary probabilities are:

$$\pi_n=\frac{r^n}{\sum_{j=0}^{K}r^j},\qquad n=0,\ldots,K.$$

At $r=1$, every state has probability $1/(K+1)$; there is no singularity. With positive arrival and service rates the finite chain is positive recurrent even when offered traffic exceeds capacity. Here $r$ is not actual utilization and may exceed one.

Poisson arrivals give rejection probability $\pi_K$. Accepted throughput, busy fraction and mean residence of admitted jobs are:

$$\lambda_{\mathrm{eff}}=\lambda(1-\pi_K)=\mu(1-\pi_0),\qquad
U=1-\pi_0,\qquad E[T\mid\text{accepted}]=\frac{L}{\lambda_{\mathrm{eff}}}.$$

Here $L=\sum n\pi_n$. Using external $\lambda$ with internal occupancy incorrectly includes requests that never entered.

**Finite queues and admission control · Experiment**

M/M/1/K with μ=100/s. K includes the job in service; full arrivals are rejected without retries. The finite chain has a stationary distribution, but admission, rejection and delay jointly determine usefulness.


At the defaults $\lambda=120$/s, $\mu=100$/s and $K=5$, rejection is about 25.1%, accepted throughput 89.9/s and mean residence 33.6 ms. Larger K can reduce rejection and keep the server busier while increasing admitted-job waiting. The model has no timeout, service failure or retry: every admitted job eventually completes. Production successful throughput needs separate accounting.

## Admission, backpressure and timeouts change different things

Admission rejects or defers work before consuming protected resources. Backpressure communicates downstream limits upstream, potentially moving the wait to a client or another queue. A timeout limits how long the caller waits. None substitutes for the others.

A client timeout need not stop server work. Retrying can create concurrent copies. Propagate deadlines, check remaining budgets while queued and at dispatch, and provide cancellation or deduplication that the operation can actually honor. These mechanisms are absent from M/M/1/K, so its mean does not predict deadline success.

A 20 ms waiting budget and mean throughput 1000/s do not imply K=20 guarantees a p99 target. Work ahead, residual service, parallelism and variability matter. Reliable bounds on individual job cost can support conservative waiting bounds; otherwise use distributions or measurements.

## Effective successful work is the objective

Distinguish offered, admitted, completed and useful successful throughput, often called goodput. Fast rejections can make an aggregate latency percentile look better. Report rejection rate beside successful-request latency and state how timeouts, cancellations and unfinished requests enter the SLO.

## Check your understanding

1. With $\lambda=\mu$ and $K=1$, what are rejection and admitted mean residence?

<details><summary>Reasoning</summary>

$\pi_0=\pi_1=1/2$, so half of arrivals are rejected, accepted throughput is $\lambda/2$ and $L=1/2$. Residence is $1/\mu$. There is no waiting slot: K=1 is not one waiting position plus one service position.

</details>

2. Does changing a timeout from 1 s to 100 ms necessarily reduce server work?

<details><summary>Reasoning</summary>

No. Without canceling original work, more frequent retries can increase it. Observe attempts, completions, cancellations, duplicates and resource occupancy rather than only shorter client waits.

</details>

## Further reading

See the [MIT lecture](https://web.mit.edu/1.041/www/lectures/L8-queuing-models-2026sp.pdf) for finite-capacity accounting, [Google SRE](https://sre.google/sre-book/handling-overload/) for admission and rejection costs, and [message semantics](../../distributed-systems/07-messages-and-streams/01-message-semantics.md) for duplicate delivery and business identity.
