---
title: Inference-time compute, candidates and verification
url: https://doc.liz6.com/en/ai/01-model-foundations/04-reasoning-and-verification
locale: en
area: ai
tags:
- Models & agents
- Model Foundations
date: 2026-06-30
modified: 2026-09-10
description: More inference effort may produce useful alternatives or repeat the same error. Separate training, generation and application verification; then distinguish having a correct candidate from selecting it. Evaluate compute budgets against complete task outcomes.
---

# Inference-time compute, candidates and verification

More inference effort may produce useful alternatives or repeat the same error. Separate training, generation and application verification; then distinguish having a correct candidate from selecting it. Evaluate compute budgets against complete task outcomes.

## Computation in Training and Inference Stages

Training adjusts model parameters; inference produces outputs given fixed parameters and inputs. Increasing inference computation can manifest as generating intermediate steps, trying multiple candidates, invoking tools to gather new information, or revising based on feedback. It does not automatically update model weights, nor does it fabricate external facts that were not originally present.

The classic approach of Chain-of-Thought (CoT) prompting involves providing demonstrations with intermediate steps in the prompt, encouraging the model to solve new problems in a similar format. The original paper observed improvements on specific models for arithmetic, commonsense, and symbolic tasks; it is not a universal law that "saying think step-by-step works for any task." [Chain-of-Thought Prompting](https://arxiv.org/abs/2201.11903)

Reasoning models often further enhance their ability to utilize inference-stage computation through specialized training. "Thinking" in APIs is a product-level feature providing reasoning control and return structures; these are not at the same level: a prompt can request the output of a solution process, but this does not imply the model internally adopted specific training methods, incurred hidden computation costs, or performed genuine external checks.

| Level | Object Changed | Example |
|---|---|---|
| Training | Parameters and behavioral tendencies | Learning to decompose, revise, and use feedback |
| Single Generation | Computation process and output for this turn | Processing more intermediate steps before answering |
| Application Orchestration | Organization of multiple generations and tool calls | Generating patches, running tests, correcting based on failures |
| Display | Content seen by the user | Brief conclusions, solution explanations, progress summaries |

## Applying Additional Computation to Candidates and Verification

### Decomposition Makes Intermediate Results Checkable

To calculate 27×453, you can break it down into 20×453=9060 and 7×453=3171, then add them to get 12231. An alternative decomposition is 453×(30−3)=13590−1359=12231. These two calculations can cross-verify each other, but if both are executed incorrectly by the same systemic error, they may both fail. For stronger computational guarantees, deterministic arithmetic tools can be used.

This example is merely a human-checkable solution explanation, not a record of a model's internal process. It demonstrates the role of decomposition: every step has inputs, operations, and verifiable outputs; these intermediate artifacts are easier to spot errors in compared to a mere claim of "I checked carefully."

### Multiple Candidates Are Only Valuable When Selection Is Effective

Self-consistency works by sampling multiple solution paths and aggregating the final answers; research has observed gains on several reasoning benchmarks. [Self-Consistency](https://arxiv.org/abs/2203.11171) However, if multiple candidates rely on the same erroneous fact, voting may confidently select the wrong answer.

Even **assuming** each attempt is independent with a success probability of 0.6, the probability of at least one success in three attempts is 1−0.4³=0.936. This does not mean the final success rate can be directly written as 93.6%. This number only represents the probability that the correct answer exists among the candidates; the application still needs to identify it. Furthermore, real models' repeated generations usually do not satisfy the independent error assumption.

### External Feedback Turns Search into Verifiable Improvement

<svg viewBox="0 0 760 355" xmlns="http://www.w3.org/2000/svg" role="img" aria-label="Value of Additional Computation: Obtaining Checkable Feedback After Proposing Candidates" style="max-width:100%;height:auto" font-family="Source Han Sans CN,Microsoft YaHei,sans-serif">
<defs><marker id="reasoning-verification-arrow" markerWidth="8" markerHeight="8" refX="7" refY="4" orient="auto"><path d="M0,0 L8,4 L0,8 Z" fill="#64748b"></path></marker></defs>
<rect width="760" height="355" rx="12" fill="#f8fafc"></rect>


<g transform="translate(0 0)"><rect x="30" y="90" width="200" height="82" rx="8" fill="#e0e7ff" stroke="#c7d2fe"></rect><text x="130.0" y="128.0" font-size="15" fill="#1e293b" text-anchor="middle" font-weight="600">Propose Candidates</text><text x="130.0" y="150.0" font-size="12" fill="#475569" text-anchor="middle" font-weight="400">Generate or sample candidates</text><line x1="230" y1="131" x2="275" y2="131" stroke="#64748b" stroke-width="1.8" marker-end="url(#reasoning-verification-arrow)"></line><rect x="280" y="90" width="200" height="82" rx="8" fill="#e0e7ff" stroke="#c7d2fe"></rect><text x="380.0" y="128.0" font-size="15" fill="#1e293b" text-anchor="middle" font-weight="600">Execute Verification</text><text x="380.0" y="150.0" font-size="12" fill="#475569" text-anchor="middle" font-weight="400">Compute, test, retrieve evidence</text><line x1="480" y1="131" x2="525" y2="131" stroke="#64748b" stroke-width="1.8" marker-end="url(#reasoning-verification-arrow)"></line><rect x="530" y="90" width="200" height="82" rx="8" fill="#e0e7ff" stroke="#c7d2fe"></rect><text x="630.0" y="128.0" font-size="15" fill="#1e293b" text-anchor="middle" font-weight="600">Judge Task Result</text><text x="630.0" y="150.0" font-size="12" fill="#475569" text-anchor="middle" font-weight="400">Accept or revise</text><line x1="630" y1="177" x2="630" y2="255" stroke="#64748b" stroke-width="1.8" marker-end="url(#reasoning-verification-arrow)"></line><line x1="630" y1="255" x2="130" y2="255" stroke="#64748b" stroke-width="1.8" marker-end="url(#reasoning-verification-arrow)"></line><line x1="130" y1="255" x2="130" y2="177" stroke="#64748b" stroke-width="1.8" marker-end="url(#reasoning-verification-arrow)"></line><text x="380" y="241" font-size="14" fill="#334155" text-anchor="middle" font-weight="400">On failure, return to the proposal stage with specific counterexamples; constrained by task budget</text></g><text x="24" y="29" font-size="19" fill="#0f172a" text-anchor="start" font-weight="700"><tspan x="24" dy="0">Value of Additional Computation: Obtaining Checkable Feedback After </tspan><tspan x="24" dy="25.650000000000002">Proposing Candidates</tspan></text><text x="24" y="288" font-size="13" fill="#475569" text-anchor="start" font-weight="400"><tspan x="24" dy="0">The diagram shows a verification loop that applications can implement, not an observation record of the model's </tspan><tspan x="24" dy="17.55">internal thought chain.</tspan></text>
</svg>

Taking the example of fixing a sorting function: Candidate A passes standard positive number examples but disrupts stable ordering when duplicate elements are present. Providing this failing input and the expected result to the model offers a more explicit revision signal than simply asking it to "check again." After Candidate B passes this example, it must still be checked on other cases not used for revision to avoid patching only against counterexamples.

Validators themselves have boundaries: unit tests only cover test conditions; scoring by another model introduces bias; majority voting requires comparable answer formats. How inference computation is allocated depends on task difficulty, candidate quality, and validator capability; returns cannot be linearly scaled solely by output token count. [Scaling LLM Test-Time Compute Optimally](https://arxiv.org/abs/2408.03314)

### Why lose despite a correct candidate?

Treat 93.6% as correct-candidate availability, then multiply by conditional selection rate q. q is not generic judge accuracy: it is defined only when the set contains a correct solution. The shared-error extreme exposes the independence assumption.

$$
P(\text{final correct})=\bigl[1-(1-p)^n\bigr]q
$$

**Why lose despite a correct candidate?**

At least one correct candidate has probability 1−(1−p)ⁿ; final success also needs a selector that finds it.


## Visible Explanations, Internal Computation, and Execution Evidence

| What is Seen | What It Can Indicate | What Cannot Be Confirmed |
|---|---|---|
| A solution explanation | The argument and assumptions provided by the model | That it faithfully recorded all internal processes |
| Thinking summary | An overview of the process the service allows displaying | That summary length equals actual reasoning overhead |
| Text stating "tests have been run" | The model claims to have performed the operation | That the tool actually executed and succeeded |
| Test artifacts associated with this run | Corresponding version, command, and results | That untested scenarios are also correct |

For Claude, current thinking documentation indicates that visible thinking text is a summary; structured blocks that may need to be resumed might still exist even when display is omitted. Internal reasoning billing in the service cannot be reconstructed by counting visible summary characters; these are the semantics of this interface and should not be generalized to all open models not returning process text. [Claude Thinking](https://platform.claude.com/docs/en/build-with-claude/thinking)

Applications should separately save the original response structure for protocol continuation, displayable text, and actual tool execution records. Do not reconstruct signatures or other opaque fields from summaries; feedback requirements in tool loops and history block retention policies should comply with current model documentation. A single response may contain thinking, text, and tool calls; one cannot declare the entire round of tasks complete based solely on the first paragraph of text.

If the frontend needs to display progress, it can show real request status and tool events that have already occurred; thinking summaries can serve as auxiliary information, but "planning to run tests" should not be displayed as "tests passed." The absence of text output does not independently prove a hang; request status, timeouts, and connection events must be considered together.

## Three Types of Budget Control

### Effort Is a Behavioral Tendency, Not a Hard Counter

Taking Claude's `output_config.effort` as an example, it adjusts the tendency for computation and token usage during responses, potentially affecting thinking, tool calls, and explanations. Supported tiers and their effects vary by model; one cannot interpret "high" as generating a fixed number of tokens, nor assume that lowering effort will necessarily shorten visible answers. [Claude Effort](https://platform.claude.com/docs/en/build-with-claude/effort)

Temperature changes the candidate probability distribution, effort changes the reasoning investment, and the length requested by the user constrains the final display goal. When a "two-sentence answer" is needed, explicitly state the length requirement; when a rigorous proof is needed, provide the proposition, conditions, and verification standards, rather than using temperature as a substitute. See [Tokens and Sampling](/ai/01-model-foundations/01-tokens-and-sampling.md) for sampling mechanisms.

### Single-Output Limits and Cross-Request Budgets

| Control | Scope | Handling When Boundary Is Reached |
|---|---|---|
| Effort or reasoning tier | Behavioral tendency for this turn | Actual consumption still needs to be tracked |
| `max_tokens` and other output limits | Single response, counted per interface definition | Check for truncation and stop reasons |
| Model-visible budget like task budget | Allows the model to manage a task segment | Do not treat self-restraint as a mandatory termination guarantee |
| Application execution limit | Requests, tools, and resources for the entire task | Scheduler stops new actions and saves state |

Claude task budget is a task-level budget signal for the model; its counting differs slightly from the client repeatedly sending history or billing usage; it cannot directly replace financial accounts or application hard limits. [Task budgets](https://platform.claude.com/docs/en/build-with-claude/task-budgets)

For example, an application allows only 12 tool calls for this task and has a total time limit of 120 seconds. This rule must be checked by the executor and cannot be assumed to be strictly followed just by writing it into the prompt. Before executing each tool, verify the remaining call count and deadline; when the budget is exhausted, preserve results, incomplete items, and recovery points. Cancelling a request does not mean already issued external operations are automatically rolled back; results must be confirmed according to the tool's own semantics.

### API Configurations Must Distinguish Model Generations

Claude's manual `budget_tokens` and adaptive thinking have their own support scopes; current documentation lists manual mode as legacy and provides migration instructions. Do not generalize "a certain generation of models rejects manual configuration" into a universal rule for all models, nor send unsupported fields to every endpoint just to unify configurations. [Extended thinking](https://platform.claude.com/docs/en/build-with-claude/extended-thinking)

The configuration adaptation layer should clearly define the target model, supported modes, and output limits before constructing requests; distinguish between 400 errors caused by unsupported parameters and model task failures. Theoretical articles retain control meanings; fields, default values, and compatibility tables in integration code should be based on the target model's current documentation.

## Calculating Returns Using Task Success Rates

### Compare Full Runs on the Same Set of Tasks

Fix the task set, input materials, tool permissions, verification rules, and timeouts, then only change the reasoning tier or candidate strategy. For classification and extraction, look at label accuracy and format; for code tasks, look at real test results; for long Agent tasks, look at whether the final goal is completed, rather than just whether each round of text looks like progress.

Below are hypothetical results for 100 tasks. Each run's consumption already includes failed attempts and verification within that strategy; for ease of calculation, all costs are unified into fictional cost units.

| Configuration | Total Successes | Total Cost | Cost Per Success |
|---|---:|---:|---:|
| A: Less computation | 70 | 100 | 100/70≈1.43 |
| B: More computation | 80 | 200 | 200/80=2.50 |

B's success rate increased by 10 percentage points, but the cost per success also increased. The extra 100 cost brought 10 extra successes, resulting in a marginal cost of 10. Whether this is worthwhile depends on task value, quality thresholds, and latency constraints; conclusions cannot be drawn solely based on "more successes" or "more tokens spent."

```python
from fractions import Fraction

runs = {"A": (70, 100), "B": (80, 200)}
for name, (successes, cost) in runs.items():
    print(name, "Cost Per Success", round(float(Fraction(cost, successes)), 2))
assert Fraction(200 - 100, 80 - 70) == 10
# A Cost Per Success 1.43
# B Cost Per Success 2.5
```

Real comparisons should also record p50/p95 latency, timeout rates, retries, and tool costs. If the new strategy uses different inputs or a different model, changes cannot be entirely attributed to effort. When sample sizes are small, retain per-task paired results and repeat measurements for stochastic tasks to avoid determining tiers based on small fluctuations in a single run.

### Adjust Computation Based on Failure Reasons, Rather Than Maximizing Uniformly

| Failure Reason | What to Address First | Is Increasing Reasoning Investment Directly Effective? |
|---|---|---|
| Missing necessary facts | Retrieve or read credible materials | Cannot fabricate evidence out of thin air |
| Understood conditions but made multi-step processing errors | Decomposition, counterexamples, verification feedback | May be effective, requires evaluation |
| Invalid tool parameter format | Schema, templates, and parsers | Not necessarily |
| Looping repeated failed actions | Save state, deduplicate, add termination conditions | May instead prolong the loop |
| Correct answer but slow return | Reduce unproductive exploration, compare lower tiers | Observe if accuracy remains stable |

Complete observation should connect requests, tools, artifacts, and task-level results; see [Evaluation and Observability](/ai/04-evaluation-and-production/01-evaluation-and-observability.md) for details. Reasoning budget is one control point in this loop; whether to turn model suggestions into controlled actions is covered in [Agent Loop and Tool Use](/ai/03-agent-systems/01-agent-loop.md).

Continue with：[Prompts and output contracts](/ai/02-context-and-interfaces/01-prompt-and-output-contracts)。
