M/M/1: utilization and tail latency

On this page

M/M/1 is a fully solvable model of random arrivals and service. It explains nonlinear waiting near full utilization and shows why a mean alone cannot represent every user's experience.

From state balance to means

Assume exogenous Poisson arrivals at , IID exponential service at , mutual independence, one FCFS server, unlimited waiting space, and no abandonment or retries. State counts all jobs inside. Arrivals increase it at rate ; when nonempty, completions decrease it at rate .

Stationary adjacent-state balance is . With , normalizing the geometric series yields:

Consequently:

At the infinite-state chain has no normalizable stationary distribution. Continuing the formula to obtain a negative wait is invalid. Equality is not a safe equilibrium for this stochastic queue.

An exactly computable tail

An arrival sees other jobs with probability . Memoryless service makes its system time the sum of independent exponential stages. Mixing these sums with geometric weights gives:

Waiting has a different distribution: probability corresponds to immediate service.

The system-time quantile is . Thus p99 is approximately 4.605 times the mean in this model. That ratio is not a universal production rule.

Preparing the visual
M/M/1: utilization and tail latency · Experiment

Stationary FCFS M/M/1, unlimited waiting space, μ=100/s. The Wq tail equals ρ at t=0 because other arrivals wait zero time. The lower marker is current utilization. All values are exact within this model.

At /s and , mean system time is 50 ms, mean wait 40 ms and system-time p99 about 230 ms. At , they become 100, 90 and 461 ms. Mean service remains 10 ms throughout.

Derive a conditional capacity boundary

Requiring system-time p99 at most 200 ms in this model gives:

With /s, this permits /s. The result is conditional on the model. Long services, batching, several stages or a slowing dependency require another model and validation with load tests.

A utilization curve also does not capture every scaling effect. More workers may leave the bottleneck's unchanged; contention may even reduce it. Identify service occupancy and load-dependent behavior before treating service capacity as constant.

Check your understanding

  1. At 50% utilization, do half of requests have zero total latency?
Reasoning

Half can start without waiting, but they still need service. has a probability mass at zero; does not. Waiting probability is not the total-response distribution.

  1. Can an observed production mean be multiplied by 4.605 to obtain p99?
Reasoning

Only if the stationary FCFS M/M/1 assumptions apply. Real response times may be mixtures, multistage sums or dependent processes. Record the distribution and check model error instead of assuming this ratio.

Further reading

MIT Urban Operations Research, section 4.6.1 derives the state and residence-time distributions. Next, we keep mean service fixed while relaxing exponential service.