Inference-time compute, candidates and verification

On this page

More inference effort may produce useful alternatives or repeat the same error. Separate training, generation and application verification; then distinguish having a correct candidate from selecting it. Evaluate compute budgets against complete task outcomes.

Computation in Training and Inference Stages

Training adjusts model parameters; inference produces outputs given fixed parameters and inputs. Increasing inference computation can manifest as generating intermediate steps, trying multiple candidates, invoking tools to gather new information, or revising based on feedback. It does not automatically update model weights, nor does it fabricate external facts that were not originally present.

The classic approach of Chain-of-Thought (CoT) prompting involves providing demonstrations with intermediate steps in the prompt, encouraging the model to solve new problems in a similar format. The original paper observed improvements on specific models for arithmetic, commonsense, and symbolic tasks; it is not a universal law that "saying think step-by-step works for any task." Chain-of-Thought Prompting

Reasoning models often further enhance their ability to utilize inference-stage computation through specialized training. "Thinking" in APIs is a product-level feature providing reasoning control and return structures; these are not at the same level: a prompt can request the output of a solution process, but this does not imply the model internally adopted specific training methods, incurred hidden computation costs, or performed genuine external checks.

LevelObject ChangedExample
TrainingParameters and behavioral tendenciesLearning to decompose, revise, and use feedback
Single GenerationComputation process and output for this turnProcessing more intermediate steps before answering
Application OrchestrationOrganization of multiple generations and tool callsGenerating patches, running tests, correcting based on failures
DisplayContent seen by the userBrief conclusions, solution explanations, progress summaries

Applying Additional Computation to Candidates and Verification

Decomposition Makes Intermediate Results Checkable

To calculate 27×453, you can break it down into 20×453=9060 and 7×453=3171, then add them to get 12231. An alternative decomposition is 453×(30−3)=13590−1359=12231. These two calculations can cross-verify each other, but if both are executed incorrectly by the same systemic error, they may both fail. For stronger computational guarantees, deterministic arithmetic tools can be used.

This example is merely a human-checkable solution explanation, not a record of a model's internal process. It demonstrates the role of decomposition: every step has inputs, operations, and verifiable outputs; these intermediate artifacts are easier to spot errors in compared to a mere claim of "I checked carefully."

Multiple Candidates Are Only Valuable When Selection Is Effective

Self-consistency works by sampling multiple solution paths and aggregating the final answers; research has observed gains on several reasoning benchmarks. Self-Consistency However, if multiple candidates rely on the same erroneous fact, voting may confidently select the wrong answer.

Even assuming each attempt is independent with a success probability of 0.6, the probability of at least one success in three attempts is 1−0.4³=0.936. This does not mean the final success rate can be directly written as 93.6%. This number only represents the probability that the correct answer exists among the candidates; the application still needs to identify it. Furthermore, real models' repeated generations usually do not satisfy the independent error assumption.

External Feedback Turns Search into Verifiable Improvement

Propose CandidatesGenerate or sample candidatesExecute VerificationCompute, test, retrieve evidenceJudge Task ResultAccept or reviseOn failure, return to the proposal stage with specific counterexamples; constrained by task budgetValue of Additional Computation: Obtaining Checkable Feedback After Proposing CandidatesThe diagram shows a verification loop that applications can implement, not an observation record of the model's internal thought chain.

Taking the example of fixing a sorting function: Candidate A passes standard positive number examples but disrupts stable ordering when duplicate elements are present. Providing this failing input and the expected result to the model offers a more explicit revision signal than simply asking it to "check again." After Candidate B passes this example, it must still be checked on other cases not used for revision to avoid patching only against counterexamples.

Validators themselves have boundaries: unit tests only cover test conditions; scoring by another model introduces bias; majority voting requires comparable answer formats. How inference computation is allocated depends on task difficulty, candidate quality, and validator capability; returns cannot be linearly scaled solely by output token count. Scaling LLM Test-Time Compute Optimally

Why lose despite a correct candidate?

Treat 93.6% as correct-candidate availability, then multiply by conditional selection rate q. q is not generic judge accuracy: it is defined only when the set contains a correct solution. The shared-error extreme exposes the independence assumption.

Preparing the visual
Why lose despite a correct candidate?

At least one correct candidate has probability 1−(1−p)ⁿ; final success also needs a selector that finds it.

Visible Explanations, Internal Computation, and Execution Evidence

What is SeenWhat It Can IndicateWhat Cannot Be Confirmed
A solution explanationThe argument and assumptions provided by the modelThat it faithfully recorded all internal processes
Thinking summaryAn overview of the process the service allows displayingThat summary length equals actual reasoning overhead
Text stating "tests have been run"The model claims to have performed the operationThat the tool actually executed and succeeded
Test artifacts associated with this runCorresponding version, command, and resultsThat untested scenarios are also correct

For Claude, current thinking documentation indicates that visible thinking text is a summary; structured blocks that may need to be resumed might still exist even when display is omitted. Internal reasoning billing in the service cannot be reconstructed by counting visible summary characters; these are the semantics of this interface and should not be generalized to all open models not returning process text. Claude Thinking

Applications should separately save the original response structure for protocol continuation, displayable text, and actual tool execution records. Do not reconstruct signatures or other opaque fields from summaries; feedback requirements in tool loops and history block retention policies should comply with current model documentation. A single response may contain thinking, text, and tool calls; one cannot declare the entire round of tasks complete based solely on the first paragraph of text.

If the frontend needs to display progress, it can show real request status and tool events that have already occurred; thinking summaries can serve as auxiliary information, but "planning to run tests" should not be displayed as "tests passed." The absence of text output does not independently prove a hang; request status, timeouts, and connection events must be considered together.

Three Types of Budget Control

Effort Is a Behavioral Tendency, Not a Hard Counter

Taking Claude's output_config.effort as an example, it adjusts the tendency for computation and token usage during responses, potentially affecting thinking, tool calls, and explanations. Supported tiers and their effects vary by model; one cannot interpret "high" as generating a fixed number of tokens, nor assume that lowering effort will necessarily shorten visible answers. Claude Effort

Temperature changes the candidate probability distribution, effort changes the reasoning investment, and the length requested by the user constrains the final display goal. When a "two-sentence answer" is needed, explicitly state the length requirement; when a rigorous proof is needed, provide the proposition, conditions, and verification standards, rather than using temperature as a substitute. See Tokens and Sampling for sampling mechanisms.

Single-Output Limits and Cross-Request Budgets

ControlScopeHandling When Boundary Is Reached
Effort or reasoning tierBehavioral tendency for this turnActual consumption still needs to be tracked
max_tokens and other output limitsSingle response, counted per interface definitionCheck for truncation and stop reasons
Model-visible budget like task budgetAllows the model to manage a task segmentDo not treat self-restraint as a mandatory termination guarantee
Application execution limitRequests, tools, and resources for the entire taskScheduler stops new actions and saves state

Claude task budget is a task-level budget signal for the model; its counting differs slightly from the client repeatedly sending history or billing usage; it cannot directly replace financial accounts or application hard limits. Task budgets

For example, an application allows only 12 tool calls for this task and has a total time limit of 120 seconds. This rule must be checked by the executor and cannot be assumed to be strictly followed just by writing it into the prompt. Before executing each tool, verify the remaining call count and deadline; when the budget is exhausted, preserve results, incomplete items, and recovery points. Cancelling a request does not mean already issued external operations are automatically rolled back; results must be confirmed according to the tool's own semantics.

API Configurations Must Distinguish Model Generations

Claude's manual budget_tokens and adaptive thinking have their own support scopes; current documentation lists manual mode as legacy and provides migration instructions. Do not generalize "a certain generation of models rejects manual configuration" into a universal rule for all models, nor send unsupported fields to every endpoint just to unify configurations. Extended thinking

The configuration adaptation layer should clearly define the target model, supported modes, and output limits before constructing requests; distinguish between 400 errors caused by unsupported parameters and model task failures. Theoretical articles retain control meanings; fields, default values, and compatibility tables in integration code should be based on the target model's current documentation.

Calculating Returns Using Task Success Rates

Compare Full Runs on the Same Set of Tasks

Fix the task set, input materials, tool permissions, verification rules, and timeouts, then only change the reasoning tier or candidate strategy. For classification and extraction, look at label accuracy and format; for code tasks, look at real test results; for long Agent tasks, look at whether the final goal is completed, rather than just whether each round of text looks like progress.

Below are hypothetical results for 100 tasks. Each run's consumption already includes failed attempts and verification within that strategy; for ease of calculation, all costs are unified into fictional cost units.

ConfigurationTotal SuccessesTotal CostCost Per Success
A: Less computation70100100/70≈1.43
B: More computation80200200/80=2.50

B's success rate increased by 10 percentage points, but the cost per success also increased. The extra 100 cost brought 10 extra successes, resulting in a marginal cost of 10. Whether this is worthwhile depends on task value, quality thresholds, and latency constraints; conclusions cannot be drawn solely based on "more successes" or "more tokens spent."

from fractions import Fraction

runs = {"A": (70, 100), "B": (80, 200)}
for name, (successes, cost) in runs.items():
    print(name, "Cost Per Success", round(float(Fraction(cost, successes)), 2))
assert Fraction(200 - 100, 80 - 70) == 10
# A Cost Per Success 1.43
# B Cost Per Success 2.5

Real comparisons should also record p50/p95 latency, timeout rates, retries, and tool costs. If the new strategy uses different inputs or a different model, changes cannot be entirely attributed to effort. When sample sizes are small, retain per-task paired results and repeat measurements for stochastic tasks to avoid determining tiers based on small fluctuations in a single run.

Adjust Computation Based on Failure Reasons, Rather Than Maximizing Uniformly

Failure ReasonWhat to Address FirstIs Increasing Reasoning Investment Directly Effective?
Missing necessary factsRetrieve or read credible materialsCannot fabricate evidence out of thin air
Understood conditions but made multi-step processing errorsDecomposition, counterexamples, verification feedbackMay be effective, requires evaluation
Invalid tool parameter formatSchema, templates, and parsersNot necessarily
Looping repeated failed actionsSave state, deduplicate, add termination conditionsMay instead prolong the loop
Correct answer but slow returnReduce unproductive exploration, compare lower tiersObserve if accuracy remains stable

Complete observation should connect requests, tools, artifacts, and task-level results; see Evaluation and Observability for details. Reasoning budget is one control point in this loop; whether to turn model suggestions into controlled actions is covered in Agent Loop and Tool Use.

Continue with:Prompts and output contracts。