After meeting quality requirements, the service must fit cost, quota and time constraints. Account for complete tasks, find quota bottlenecks and average in-flight work, then fit retries within a deadline. Model-internal KV and computation belong to the inference article; this one covers scheduling and failures.
Cost Calculation Based on the Full Execution Chain
Tokens are a common billing metric, but not the only source of cost. Model requests may distinguish between standard input, cache writes, cache reads, and output; tools, retrieval, storage, execution containers, networking, and self-hosted compute also incur fees. Different providers and products have different accounting methods; you cannot treat "output is always five times more expensive than input" as a universal rule.
Task Cost = Sum of all model request fees + Tool and execution fees + Associated storage/retrieval fees
Assume a service's instructional pricing is 2 currency units per million input tokens and 10 units for output. For a request with 20,000 input tokens and 2,000 output tokens, the model fee is 0.04 + 0.02 = 0.06; adding a tool fee of 0.01 brings the total to 0.07. If a task requires three calls of this scale, the cost is 0.21; additional failed attempts must also be added based on the actual fees incurred. These numbers are for calculation purposes only and do not reflect current product pricing.
Optimization Method
Potential Savings
Factors to Monitor
Removing unnecessary input
Input fees and some processing time
Whether necessary evidence is lost
Prefix caching
Computation or fees for repeated prefixes
Hit conditions, write overhead, and validity period
Adjusting model and effort
Single calculation investment
Full task success rate and retries
Controlling final length
Visible output volume
Whether it covers task requirements
Non-real-time batching
Depends on service's batch pricing
Completion deadlines, per-item failures, and duplicate submissions
Cache benefits can be modeled without real pricing: let prefix length be L, standard input unit price be u, write unit price be w, and read unit price be r. If reused N times with all other requests hitting the cache, compare N×L×u with L×w+(N−1)×L×r. If calls within the TTL are too few or prefixes change frequently, the write premium may not be recovered; mechanisms and conditions are detailed in Context Engineering.
Switching to cheaper models also requires calculating the entire task. If the single-request cost is halved, but failures lead to three additional attempts, it may not be cheaper. Report the total cost amortized per successful task and record the cost of unsuccessful tasks, rather than just comparing the quote for a single request.
Deconstructing Latency and Throughput
End-to-end time includes queuing, networking, input processing, inference/generation, tool usage, and application verification. Serial dependencies add up, while parallel branches are calculated by the critical path; the statement "latency is mainly determined by output length" applies only to certain workloads.
Time Metric
Measured From
To
Purpose
Queuing Time
Entering service
Start of processing
Insufficient admission or backend capacity
Time to First Visible Text
Request start
Appearance of useful text
Waiting for input, thinking, network, and display
Generation Phase Time
Start of generation
Completion of this output
Decoding, output size, or pauses
Tool Phase Time
Actual dispatch
Result confirmation
External services and serial dependencies
Task Completion Time
User submission
Acceptance completion
Actual user experience
Streaming delivers results incrementally, typically improving perceived wait times and allowing connection keep-alive; it does not automatically increase model generation speed, remove output limits, or guarantee that all agents and networks won't timeout. You need to set connection, read idle, single-call, and total task deadlines separately; interrupting the stream after receiving partial text does not mark a half-formed structured result as complete.
Caching, shortening inputs, batching, parallelism, and reduced exploration may improve certain stages, but you must identify bottlenecks via traces. If a tool query takes 20 seconds, cutting 100 output tokens may have little impact on the overall time; if queuing is the main issue, lowering the frontend timeout alone will only create more failed retries.
Quotas, Admission, and Bounded Queuing
Quotas may simultaneously limit request count, input tokens, output tokens, and token generation speed; specific dimensions should be confirmed via service documentation and account configuration. Claude Rate limits only limits request concurrency and does not guarantee token quotas won't be exceeded; limiting only average RPM does not eliminate instantaneous bursts.
Assume allowances of 100 requests/minute, 200,000 input tokens/minute, and 50,000 output tokens/minute. If each request averages 10,000 input tokens and 1,000 output tokens, the average limits for the three dimensions are 100, 20, and 50 requests/minute respectively, with the tightest constraint being 20 requests/minute.
If the average service time is 30 seconds, under simplified stable conditions with no backlog, 20 requests/minute corresponds to an average of about 10 in-flight requests. This Little's Law example is an average capacity account; it does not mean setting production concurrency to 10 is safe: length fluctuations, tail latency, bursts, and external shared quotas require headroom and dynamic control.
fromfractionsimportFractionrpm=100input_tpm, output_tpm=200_000, 50_000input_per_request, output_per_request=10_000, 1_000rate=min(Fraction(rpm),Fraction(input_tpm,input_per_request),Fraction(output_tpm,output_per_request))mean_inflight=rate/60*30assertrate==20assertmean_inflight==10assert3*3==9# Worst-case amplification with max 3 attempts per layer
print("Average requests/minute:",rate,"Average in-flight:",mean_inflight)
Queues require length limits and wait deadlines; beyond these, defer, degrade, or reject based on business logic. Online interactions and deferrable offline tasks should be scheduled separately to prevent long batch processing from exhausting interactive resources. Continuously accepting requests during overload, making all requests slower, is often harder to recover from than early rejection; capacity management should be based on real resources rather than single request counts. Google SRE: Handling Overload
Retries Require Deadlines and Idempotency Boundaries
Error Categories Are More Useful Than Numeric Rules of Thumb
Category
Handling Principle
Invalid parameters or unsupported capabilities
Fix the request; retrying as-is is usually meaningless
Identity or permission issues
Recover via authentication flow or report; do not escalate permission bypasses
Throttling or short-term overload
Respect service hints, use limited backoff, and reduce pressure
Connection, timeout, and partial service errors
First determine if the operation can be safely replayed
Unknown write action results
Query status or use supported idempotent protocols
Quota exhaustion, long-term unavailability
Wait for conditions to change or use explicit degradation; do not retry infinitely
Do not reduce this to "only 429/5xx can be retried" or "all 5xx should be retried." Whether 408, 409, etc., are suitable for retry depends on the protocol and operation; vendor errors and SDK default strategies should also be checked separately. Claude Errors
Incorporate Backoff into the Same Task Deadline
Common backoff upper bounds are min(cap, base×2^attempt), introducing jitter within the interval to prevent multiple clients from waking up simultaneously. When the service returns Retry-After, schedule according to its semantics, constrained by the task's remaining deadline; if there is no time left to complete one attempt after waiting, do not continue starting.
For example, if a task has 8 seconds remaining, the service requires at least 6 seconds of waiting, and the next operation requires reserving 5 seconds, this retry does not fit the budget. Setting "timeout=10 seconds per request" does not mean the task will only wait 10 seconds; retries, backoff, queuing, and tool phases all consume total duration.
If the caller retries three times and the SDK internally attempts up to three times each, the worst case is nine underlying requests. Clarify which layer manages retries, count actual attempts, and set a retry budget for the entire service if necessary; retry amplification during overload makes already insufficient capacity even tighter. Google SRE: Addressing Cascading Failures
Human Confirmation Does Not Replace Idempotency
User confirmation to "create an order" authorizes one business intent, but does not guarantee that network retries won't create two orders. Idempotency keys must correspond to the same business operation and be saved and verified by the actual service; sending different parameters for the same key should be handled according to service semantics, not just by writing "do not repeat" in the model prompt.
Request timeouts do not prove the service didn't execute. If the action is complete but the response hasn't arrived, query or replay the existing result first; canceling the wait does not automatically revoke the action. This boundary is detailed in Agent Loop with complete timing.
Which quota runs out first?
Convert quotas into request rates before taking the minimum. In Little’s law, L is mean in-flight requests, not a recommended hard concurrency cap or a guarantee on queue-tail latency.
r≤min(R,IˉImax,OˉOmax),L=60rWˉ
Preparing the visual
Which quota runs out first?
Sustainable throughput cannot exceed the smallest converted quota; average in-flight work also depends on service time.
Degradation, Batching, and Operational Verification
Degradation Must Still Meet Minimum Task Requirements
Verify context length, structured output, tool formats, input modalities, and permission boundaries before switching models. Do not silently connect a text-only fallback model to an image task, nor delete key constraints to fake equivalent results just because the backup service supports a different scope.
Options include returning existing trusted parts, deferring to offline completion, using a compatible fallback, or explicitly reporting temporary unavailability. When caching old results, mark the applicable time and version, and check user access scope. When restoring services, prevent backlog tasks from replaying simultaneously to form a new burst.
Batch Submissions Manage Lifecycle by Item
Non-real-time tasks can use batch interfaces, but price discounts, completion deadlines, size limits, and cache interactions are service-specific rules. Batch end does not mean all items succeeded; distinguish success, failure, expiration, and cancellation by item ID, and only retry items that should be retried. Claude Batch processing
Batch/parallel results may be out of order; save request mappings and match outputs by ID. When resubmitting failed items, preserve the association between the original task and new attempts to avoid overwriting successful results; model batching itself does not guarantee downstream business operations execute exactly once.
Verification Targets a Controllable Complete Service
Establish operational examples including short requests, long inputs, long tool waits, throttling, stream interruption, unknown results, and fallback switching. Record success rates, p50/p95 latency, queue length, actual attempt counts, quota usage, and cost per successful item.
How much do layered retries amplify?
“Retry twice” can be confused with two total calls; this model consistently includes the initial call in attempt limits. When SDK and workflow both retry, account for their product and share deadlines and retry budgets.