System boundaries and queueing timelines

On this page

A service that needs 10 ms on average does not necessarily return a response within 10 ms. Queueing theory separates waiting for a resource from occupying it, then studies how arrivals, service, scheduling and capacity shape the wait.

Draw the boundary first

An application entrance, connection pool, worker pool and database can each have a queue. Decide when a job enters and leaves, and whether jobs in service count. The boundary below includes the waiting and service areas; admission rejection occurs outside it.

Queueing system boundaryExternal jobs → AdmissionWaiting area Q(t)In service B(t)Exit: success / failure / cancel

For job , let be arrival, service start and departure. This chapter has uninterrupted service and completion as the only departure outcome:

is waiting time, service time and system or response time. The number inside is : counts waiting jobs and jobs in service. A single server has ; servers can have up to jobs in service.

Service time depends on the resource boundary. A connection's occupancy from checkout to return can include remote I/O waits; it is not CPU execution time. Asynchronous I/O may release a thread while retaining a connection, memory or downstream quota.

Calculate one trace

For a single, work-conserving FIFO / FCFS server:

Start empty with previous completion time zero. Arrivals s and services s produce starts and departures . Waits are , averaging 2.25 s. System times are all 4 s, even though mean service is only 1.75 s.

Write cumulative arrivals, starts and departures as . Starting empty with no removal other than completion gives , and . Vertical differences between cumulative curves are counts. Horizontal distances for the same job are times.

Preparing the visual
System boundaries and queueing timelines · Experiment

Four jobs, one FIFO server, initially empty; the first job is long and the others take 1 s. Completions precede arrivals at ties. Bars and cumulative curves share the same trace.

Predict who is affected when only the first job gets longer. Then increase the arrival gap: waiting can disappear while service bars retain their width. Completions precede arrivals at equal timestamps so a just-released server is available.

From a trace to statistics

Arrival rate and per-server service capacity have units jobs/s. and denote stationary mean counts. Actual completion throughput differs from offered traffic: an idle server does not continuously produce completions at rate .

For one stable, work-conserving server with no losses, its busy fraction is . This is not automatically the CPU utilization of a system with several resources, nor is it concurrency divided by worker count. Stability and tails require additional model assumptions introduced later.

Check your understanding

  1. A connection-pool boundary contains 100 requests, of which 20 hold connections. What are ?
Reasoning

If entry is connection request and departure is connection return, . A waiting-queue metric reports 80, not total concurrency. Holders waiting on a database still occupy the connection and belong to this model's service area.

  1. CPU utilization is 30% while response time rises. Does that rule out queueing?
Reasoning

No. Waiting can occur at connections, locks, disks or remote services. Record request, acquisition and release timestamps for the relevant resource. Adding workers can simply move more demand onto the same downstream bottleneck.

Further reading

The CMU modeling course begins with system vocabulary and performance measures. Google SRE discusses differing request costs and resource constraints. Next, the same kind of trace yields Little's law.