More inference effort may produce useful alternatives or repeat the same error. Separate training, generation and application verification; then distinguish having a correct candidate from selecting it. Evaluate compute budgets against complete task outcomes.
Computation in Training and Inference Stages
Training adjusts model parameters; inference produces outputs given fixed parameters and inputs. Increasing inference computation can manifest as generating intermediate steps, trying multiple candidates, invoking tools to gather new information, or revising based on feedback. It does not automatically update model weights, nor does it fabricate external facts that were not originally present.
The classic approach of Chain-of-Thought (CoT) prompting involves providing demonstrations with intermediate steps in the prompt, encouraging the model to solve new problems in a similar format. The original paper observed improvements on specific models for arithmetic, commonsense, and symbolic tasks; it is not a universal law that "saying think step-by-step works for any task." Chain-of-Thought Prompting
Reasoning models often further enhance their ability to utilize inference-stage computation through specialized training. "Thinking" in APIs is a product-level feature providing reasoning control and return structures; these are not at the same level: a prompt can request the output of a solution process, but this does not imply the model internally adopted specific training methods, incurred hidden computation costs, or performed genuine external checks.
Level
Object Changed
Example
Training
Parameters and behavioral tendencies
Learning to decompose, revise, and use feedback
Single Generation
Computation process and output for this turn
Processing more intermediate steps before answering
Application Orchestration
Organization of multiple generations and tool calls
Generating patches, running tests, correcting based on failures
Applying Additional Computation to Candidates and Verification
Decomposition Makes Intermediate Results Checkable
To calculate 27×453, you can break it down into 20×453=9060 and 7×453=3171, then add them to get 12231. An alternative decomposition is 453×(30−3)=13590−1359=12231. These two calculations can cross-verify each other, but if both are executed incorrectly by the same systemic error, they may both fail. For stronger computational guarantees, deterministic arithmetic tools can be used.
This example is merely a human-checkable solution explanation, not a record of a model's internal process. It demonstrates the role of decomposition: every step has inputs, operations, and verifiable outputs; these intermediate artifacts are easier to spot errors in compared to a mere claim of "I checked carefully."
Multiple Candidates Are Only Valuable When Selection Is Effective
Self-consistency works by sampling multiple solution paths and aggregating the final answers; research has observed gains on several reasoning benchmarks. Self-Consistency However, if multiple candidates rely on the same erroneous fact, voting may confidently select the wrong answer.
Even assuming each attempt is independent with a success probability of 0.6, the probability of at least one success in three attempts is 1−0.4³=0.936. This does not mean the final success rate can be directly written as 93.6%. This number only represents the probability that the correct answer exists among the candidates; the application still needs to identify it. Furthermore, real models' repeated generations usually do not satisfy the independent error assumption.
External Feedback Turns Search into Verifiable Improvement
Taking the example of fixing a sorting function: Candidate A passes standard positive number examples but disrupts stable ordering when duplicate elements are present. Providing this failing input and the expected result to the model offers a more explicit revision signal than simply asking it to "check again." After Candidate B passes this example, it must still be checked on other cases not used for revision to avoid patching only against counterexamples.
Validators themselves have boundaries: unit tests only cover test conditions; scoring by another model introduces bias; majority voting requires comparable answer formats. How inference computation is allocated depends on task difficulty, candidate quality, and validator capability; returns cannot be linearly scaled solely by output token count. Scaling LLM Test-Time Compute Optimally
Why lose despite a correct candidate?
Treat 93.6% as correct-candidate availability, then multiply by conditional selection rate q. q is not generic judge accuracy: it is defined only when the set contains a correct solution. The shared-error extreme exposes the independence assumption.
P(final correct)=[1−(1−p)n]q
Preparing the visual
Why lose despite a correct candidate?
At least one correct candidate has probability 1−(1−p)ⁿ; final success also needs a selector that finds it.
Visible Explanations, Internal Computation, and Execution Evidence
What is Seen
What It Can Indicate
What Cannot Be Confirmed
A solution explanation
The argument and assumptions provided by the model
That it faithfully recorded all internal processes
Thinking summary
An overview of the process the service allows displaying
That summary length equals actual reasoning overhead
Text stating "tests have been run"
The model claims to have performed the operation
That the tool actually executed and succeeded
Test artifacts associated with this run
Corresponding version, command, and results
That untested scenarios are also correct
For Claude, current thinking documentation indicates that visible thinking text is a summary; structured blocks that may need to be resumed might still exist even when display is omitted. Internal reasoning billing in the service cannot be reconstructed by counting visible summary characters; these are the semantics of this interface and should not be generalized to all open models not returning process text. Claude Thinking
Applications should separately save the original response structure for protocol continuation, displayable text, and actual tool execution records. Do not reconstruct signatures or other opaque fields from summaries; feedback requirements in tool loops and history block retention policies should comply with current model documentation. A single response may contain thinking, text, and tool calls; one cannot declare the entire round of tasks complete based solely on the first paragraph of text.
If the frontend needs to display progress, it can show real request status and tool events that have already occurred; thinking summaries can serve as auxiliary information, but "planning to run tests" should not be displayed as "tests passed." The absence of text output does not independently prove a hang; request status, timeouts, and connection events must be considered together.
Three Types of Budget Control
Effort Is a Behavioral Tendency, Not a Hard Counter
Taking Claude's output_config.effort as an example, it adjusts the tendency for computation and token usage during responses, potentially affecting thinking, tool calls, and explanations. Supported tiers and their effects vary by model; one cannot interpret "high" as generating a fixed number of tokens, nor assume that lowering effort will necessarily shorten visible answers. Claude Effort
Temperature changes the candidate probability distribution, effort changes the reasoning investment, and the length requested by the user constrains the final display goal. When a "two-sentence answer" is needed, explicitly state the length requirement; when a rigorous proof is needed, provide the proposition, conditions, and verification standards, rather than using temperature as a substitute. See Tokens and Sampling for sampling mechanisms.
Single-Output Limits and Cross-Request Budgets
Control
Scope
Handling When Boundary Is Reached
Effort or reasoning tier
Behavioral tendency for this turn
Actual consumption still needs to be tracked
max_tokens and other output limits
Single response, counted per interface definition
Check for truncation and stop reasons
Model-visible budget like task budget
Allows the model to manage a task segment
Do not treat self-restraint as a mandatory termination guarantee
Application execution limit
Requests, tools, and resources for the entire task
Scheduler stops new actions and saves state
Claude task budget is a task-level budget signal for the model; its counting differs slightly from the client repeatedly sending history or billing usage; it cannot directly replace financial accounts or application hard limits. Task budgets
For example, an application allows only 12 tool calls for this task and has a total time limit of 120 seconds. This rule must be checked by the executor and cannot be assumed to be strictly followed just by writing it into the prompt. Before executing each tool, verify the remaining call count and deadline; when the budget is exhausted, preserve results, incomplete items, and recovery points. Cancelling a request does not mean already issued external operations are automatically rolled back; results must be confirmed according to the tool's own semantics.
API Configurations Must Distinguish Model Generations
Claude's manual budget_tokens and adaptive thinking have their own support scopes; current documentation lists manual mode as legacy and provides migration instructions. Do not generalize "a certain generation of models rejects manual configuration" into a universal rule for all models, nor send unsupported fields to every endpoint just to unify configurations. Extended thinking
The configuration adaptation layer should clearly define the target model, supported modes, and output limits before constructing requests; distinguish between 400 errors caused by unsupported parameters and model task failures. Theoretical articles retain control meanings; fields, default values, and compatibility tables in integration code should be based on the target model's current documentation.
Calculating Returns Using Task Success Rates
Compare Full Runs on the Same Set of Tasks
Fix the task set, input materials, tool permissions, verification rules, and timeouts, then only change the reasoning tier or candidate strategy. For classification and extraction, look at label accuracy and format; for code tasks, look at real test results; for long Agent tasks, look at whether the final goal is completed, rather than just whether each round of text looks like progress.
Below are hypothetical results for 100 tasks. Each run's consumption already includes failed attempts and verification within that strategy; for ease of calculation, all costs are unified into fictional cost units.
Configuration
Total Successes
Total Cost
Cost Per Success
A: Less computation
70
100
100/70≈1.43
B: More computation
80
200
200/80=2.50
B's success rate increased by 10 percentage points, but the cost per success also increased. The extra 100 cost brought 10 extra successes, resulting in a marginal cost of 10. Whether this is worthwhile depends on task value, quality thresholds, and latency constraints; conclusions cannot be drawn solely based on "more successes" or "more tokens spent."
fromfractionsimportFractionruns={"A":(70,100),"B":(80,200)}forname, (successes,cost) inruns.items():print(name,"Cost Per Success",round(float(Fraction(cost,successes)),2))assertFraction(200-100,80-70)==10# A Cost Per Success 1.43
# B Cost Per Success 2.5
Real comparisons should also record p50/p95 latency, timeout rates, retries, and tool costs. If the new strategy uses different inputs or a different model, changes cannot be entirely attributed to effort. When sample sizes are small, retain per-task paired results and repeat measurements for stochastic tasks to avoid determining tiers based on small fluctuations in a single run.
Adjust Computation Based on Failure Reasons, Rather Than Maximizing Uniformly
Failure Reason
What to Address First
Is Increasing Reasoning Investment Directly Effective?
Missing necessary facts
Retrieve or read credible materials
Cannot fabricate evidence out of thin air
Understood conditions but made multi-step processing errors
Complete observation should connect requests, tools, artifacts, and task-level results; see Evaluation and Observability for details. Reasoning budget is one control point in this loop; whether to turn model suggestions into controlled actions is covered in Agent Loop and Tool Use.