Control in computing systems

12 min read
On this page

A scaling decision may already be issued while new replicas are still starting. Confusing requested capacity with available capacity—or ignoring pending actions—can lead to repeated overcorrection. This chapter applies sampling and delay and constraints.

Choose states and units

Let be waiting requests, arrivals per second, per-replica capacity, available replicas and seconds. A fluid approximation is

Work is divisible, replicas homogeneous and capacity fixed. There are no individual service times, in-flight tasks, retries, failures or load-balancing details. This explains backlog, not real p99 latency.

For , backlog grows by 30 requests/s. A 20 s activation delay can accumulate 600 requests. Scaling to 7 only matches arrivals; it does not drain existing backlog. Spare capacity or falling arrivals are required.

Suggestion, command and activation differ

A teaching recommendation is

Estimated arrivals cover new work; allocates capacity to drain old work. At arrival 70, backlog 600 and clearance target 30 s, the calculation requests 9 replicas, not 7.

This is not Kubernetes HPA and has no universal delay-stability guarantee. Replica bounds, rate limits and activation delays modify its behavior. “Set desired count to 9” must not become “add another 9” at every decision.

Preparing the visual
Control in computing systems · Experiment

Decisions every 10 s; arrivals rise from 20 to 70 requests/s during 20–100 s; capacity 10 per replica, at most 12. Deterministic fluid model, no retries. Delayed commands activate in order; stabilization uses a rolling maximum. Not full HPA or a p99 model.

Arrivals rise from 20 to 70 requests/s during 20–100 s, then fall to 20. Decisions occur every 10 s; each replica serves 10 requests/s, up to 12 replicas. Commands activate in order after the selected delay. A downscale window takes the largest recent suggestion. Equal activation delays for increases and decreases are a teaching simplification.

Smoothing has a cost

Tolerance ignores small metric changes; stabilization windows resist immediate reversals; rate limits bound capacity changes. These can reduce churn but delay response or retain idle resources. A rolling maximum is not a moving average.

HPA's basic ratio calculation is

Its full algorithm also handles tolerance, missing metrics, unready Pods, multiple metrics and scaling behavior. Downscale stabilization uses the highest recommendation in its window, rather than simply requiring several consecutive low readings. Check fields and defaults against the deployed version. Official HPA behavior

Backlog, utilization and latency are not interchangeable

Low CPU can reflect I/O waits. High CPU need not imply a waiting queue. Mean and tail latency depend on service-time distributions, scheduling and load. Converting one metric into another without a model introduces measurement error into feedback.

Under appropriate stable, long-run-average conditions, Little's law relates average population, effective throughput and mean sojourn time. Instantaneous queue divided by instantaneous rate is not p99. Queueing theory will develop these conditions; our update only balances work.

Transfer the method, not the gains

TCP adjusts sending using ACKs, loss, RTT or ECN, with network delay and algorithm-specific state. Do not label every TCP algorithm a PID controller. Revisit TCP to identify observations and actions while preserving protocol details.

Power circuits use voltage/current measurements and power stages; stored energy and compensation shape their dynamics. The reusable element is the modeling method, not identical gains.

Check your understanding

At 70 arrivals/s, 10 per replica and backlog 600, how long do 7 replicas take to empty the queue?

Reasoning

They do not under constant arrivals: spare capacity is zero. Nine replicas provide 20 requests/s spare capacity, clearing it in 30 s if capacity is immediately available and nothing else changes.

Is a longer downscale window always better?

Reasoning

It may reduce repeated starts but retain idle capacity. Compare backlog, command changes and instance-seconds on the same trace, including renewed demand.

Practical validation

Record metric timestamps, sampling, decisions, commands and actual activation. Check pending-action handling. Replay steady demand, bursts, sustained overload, cold starts and missing metrics; report backlog, resource cost and action frequency. Good results support the tested conditions, not all possible workloads.