---
title: Tokens, probabilities and sampling
url: https://doc.liz6.com/en/ai/01-model-foundations/01-tokens-and-sampling
locale: en
area: ai
tags:
- Models & agents
- Model Foundations
date: 2026-06-30
modified: 2026-09-10
description: How does text become model input, and how do scores become the next token? Follow text → token IDs → logits → probabilities → selection, then compare temperature and candidate filtering. The numerical examples explain decoding; task results determine which settings are useful.
---

# Tokens, probabilities and sampling

How does text become model input, and how do scores become the next token? Follow text → token IDs → logits → probabilities → selection, then compare temperature and candidate filtering. The numerical examples explain decoding; task results determine which settings are useful.

## From text to token ID

### Tokens are vocabulary units, not equal to characters or words

Text models typically encode strings into sequences of integers, where each integer corresponds to a token in the vocabulary. A token can represent a complete word, part of a word, punctuation, a combination of spaces, or a byte fragment. A single Unicode character may span multiple tokens; conversely, multiple characters may combine into a single token. Therefore, you cannot calculate the input for all models by simply multiplying the number of Chinese characters by two.

BPE (Byte-Pair Encoding) is a common subword construction method. During training, it starts with smaller units and repeatedly merges adjacent units based on corpus statistics. Encoding uses the already trained vocabulary and merge rules, rather than retraining for each input. For example, **assuming** it has learned to merge in order `(l,o)→lo` and `(lo,w)→low`, `l o w e r` would be tokenized as `low e r`. This is merely a mechanism example and does not represent the actual tokenization of `lower` by any existing model.

Byte-level BPE uses a base representation that covers bytes, avoiding issues where a pure character vocabulary lacks certain characters, but the name "BPE" itself does not guarantee byte fallback. Tokenization may also use methods such as Unigram or WordPiece. Normalization, pre-tokenization, and special tokens in the configuration also affect the final result. [Tokenizers Model Types](https://huggingface.co/docs/tokenizers/en/api/models)

### Why counting only the body text is insufficient

Chat systems need to encode roles and message boundaries into the model input. Tool definitions, structured content, images, etc., may also generate additional counts. The string length seen by the user and the model input length belong to different representation layers: first determine how the complete request is constructed, then calculate its token count.

For example, two user messages and combining two texts into one message may have identical body character counts, but the message boundaries may differ. If the tokenizer includes normalization processing, Unicode character sequences that look identical or similar may produce different results; do not assume all tokenizers preserve or eliminate such differences. [Tokenization Pipeline](https://huggingface.co/docs/tokenizers/en/pipeline)

Server-side counts should use the interface corresponding to the target model and be calibrated with actual response usage. Claude's counting documentation also states that counting results are estimates, and counts should be recalculated after model updates; there is no cross-model conversion multiplier that applies to all Chinese, code, and tool schemas. [Claude Token counting](https://platform.claude.com/docs/en/build-with-claude/token-counting)

## From logits to probabilities

Standard autoregressive text generation predicts the next token given the current prefix. The model outputs scores (logits) for each position in the vocabulary; the decoder then selects an ID based on the configuration, appends it to the prefix, and continues. Actual inference engines may use batching or speculative decoding to optimize execution, but the semantics of conditional generation can still be understood through this step-by-step process.

<svg viewBox="0 0 760 358.1576843261719" xmlns="http://www.w3.org/2000/svg" role="img" aria-label="One autoregressive generation: scores, candidate sets, and selection are different steps" style="max-width:100%;height:auto" font-family="Source Han Sans CN,Microsoft YaHei,sans-serif">
<defs><marker id="token-generation-arrow" markerWidth="8" markerHeight="8" refX="7" refY="4" orient="auto"><path d="M0,0 L8,4 L0,8 Z" fill="#64748b"></path></marker></defs>
<rect width="760" height="358.1576843261719" rx="12" fill="#f8fafc"></rect>


<g transform="translate(0 0)"><rect x="30" y="85" width="200" height="75" rx="8" fill="#e0e7ff" stroke="#c7d2fe"></rect><text x="130.0" y="119.5" font-size="15" fill="#1e293b" text-anchor="middle" font-weight="600">Current token sequence</text><text x="130.0" y="141.5" font-size="12" fill="#475569" text-anchor="middle" font-weight="400">Prompt + Generated content</text><line x1="230" y1="122" x2="275" y2="122" stroke="#64748b" stroke-width="1.8" marker-end="url(#token-generation-arrow)"></line><rect x="280" y="85" width="200" height="75" rx="8" fill="#e0e7ff" stroke="#c7d2fe"></rect><text x="380.0" y="119.5" font-size="15" fill="#1e293b" text-anchor="middle" font-weight="600">Model forward pass</text><text x="380.0" y="141.5" font-size="12" fill="#475569" text-anchor="middle" font-weight="400">Output logits for next position</text><line x1="480" y1="122" x2="525" y2="122" stroke="#64748b" stroke-width="1.8" marker-end="url(#token-generation-arrow)"></line><rect x="530" y="85" width="200" height="75" rx="8" fill="#e0e7ff" stroke="#c7d2fe"></rect><text x="630.0" y="119.5" font-size="15" fill="#1e293b" text-anchor="middle" font-weight="600">Processing and selection</text><text x="630.0" y="141.5" font-size="12" fill="#475569" text-anchor="middle" font-weight="400">Temperature, filtering, sampling</text><rect x="280" y="220" width="200" height="62" rx="8" fill="#e0e7ff" stroke="#c7d2fe"></rect><text x="380.0" y="248.0" font-size="15" fill="#1e293b" text-anchor="middle" font-weight="600">Selected token ID</text><text x="380.0" y="270.0" font-size="12" fill="#475569" text-anchor="middle" font-weight="400">Decode into visible text</text><line x1="630" y1="160" x2="630" y2="251" stroke="#64748b" stroke-width="1.8" marker-end="url(#token-generation-arrow)"></line><line x1="630" y1="251" x2="485" y2="251" stroke="#64748b" stroke-width="1.8" marker-end="url(#token-generation-arrow)"></line><line x1="280" y1="251" x2="130" y2="251" stroke="#64748b" stroke-width="1.8" marker-end="url(#token-generation-arrow)"></line><line x1="130" y1="251" x2="130" y2="165" stroke="#64748b" stroke-width="1.8" marker-end="url(#token-generation-arrow)"></line></g><text x="24" y="29" font-size="19" fill="#0f172a" text-anchor="start" font-weight="700"><tspan x="24" dy="0">One autoregressive generation: scores, candidate sets, and selection are </tspan><tspan x="24" dy="25.650000000000002">different steps</tspan></text><text x="24" y="315" font-size="13" fill="#475569" text-anchor="start" font-weight="400"><tspan x="24" dy="0">Stopping conditions are determined by end markers, length limits, or interface rules; visible text blocks do not </tspan><tspan x="24" dy="17.55">necessarily equal one token.</tspan></text>
</svg>

For positive temperature $T$, a numerically stable way to write softmax is:

**pᵢ = exp((zᵢ − max(z))/T) / Σⱼ exp((zⱼ − max(z))/T)**

Subtracting the maximum value does not change the normalized probabilities but avoids directly calculating huge exponents. Suppose there are only three candidates A, B, C, with logits `[2, 1, 0]`:

| Temperature | A | B | C |
|---|---:|---:|---:|
| T=1 | 66.52% | 24.47% | 9.00% |
| T=0.5 | 86.68% | 11.73% | 1.59% |

<svg viewBox="0 0 760 345.36669921875" xmlns="http://www.w3.org/2000/svg" role="img" aria-label="Same logits = [2, 1, 0], temperature changes probability concentration" style="max-width:100%;height:auto" font-family="Source Han Sans CN,Microsoft YaHei,sans-serif">
<defs><marker id="temperature-prob-arrow" markerWidth="8" markerHeight="8" refX="7" refY="4" orient="auto"><path d="M0,0 L8,4 L0,8 Z" fill="#64748b"></path></marker></defs>
<rect width="760" height="345.36669921875" rx="12" fill="#f8fafc"></rect>


<g transform="translate(0 0)"><text x="28" y="109" font-size="16" fill="#334155" text-anchor="start" font-weight="600">T = 1</text><rect x="160" y="85" width="365.86" height="42" rx="8" fill="#a5b4fc" stroke="white"></rect><rect x="525.86" y="85" width="134.585" height="42" rx="8" fill="#6ee7b7" stroke="white"></rect><rect x="660.445" y="85" width="49.5" height="42" rx="8" fill="#fcd34d" stroke="white"></rect><text x="160" y="155" font-size="14" fill="#334155" text-anchor="start" font-weight="400">A 66.52%    B 24.47%    C 9.00%</text><text x="28" y="219" font-size="16" fill="#334155" text-anchor="start" font-weight="600">T = 0.5</text><rect x="160" y="195" width="476.74" height="42" rx="8" fill="#a5b4fc" stroke="white"></rect><rect x="636.74" y="195" width="64.515" height="42" rx="8" fill="#6ee7b7" stroke="white"></rect><rect x="701.255" y="195" width="8.745000000000001" height="42" rx="8" fill="#fcd34d" stroke="white"></rect><text x="160" y="265" font-size="14" fill="#334155" text-anchor="start" font-weight="400">A 86.68%    B 11.73%    C 1.59%</text></g><text x="24" y="29" font-size="19" fill="#0f172a" text-anchor="start" font-weight="700"><tspan x="24" dy="0">Same logits = [2, 1, 0], temperature changes probability concentration</tspan></text><text x="24" y="302.208984375" font-size="13" fill="#475569" text-anchor="start" font-weight="400"><tspan x="24" dy="0">Cooling down does not add new evidence, nor does it improve model knowledge; it makes existing high-scoring </tspan><tspan x="24" dy="17.55">candidates easier to select.</tspan></text>
</svg>

Cooling down amplifies relative score differences, while heating up flattens the distribution. This changes the randomness of selection, it is not a "correctness knob": if the wrong answer originally had the highest score, cooling down will select it more stably. The entire response is not written after a single sample; changes in early tokens will alter the subsequent conditional distribution.

## Candidate truncation: top-k and top-p

top-k retains the k candidates with the highest scores; top-p first sorts by probability, retains the smallest prefix where the cumulative probability first reaches the threshold p, and then renormalizes within the retained set. The size of the latter set changes with the distribution.

For the T=1 distribution above, with top-p=0.8, A alone has only 0.6652, which is not enough; adding B reaches approximately 0.9100, so A and B are retained, and C is discarded. Renormalizing gives A≈0.7311, B≈0.2689. **top-p=0.8 does not mean retaining 80% of the vocabulary, nor does it mean the highest candidate has an 80% probability.**

At T=0.5, A alone exceeds 0.8, so with the same threshold, only A is retained. Therefore, temperature and filtering order jointly affect the results. Different engines may also combine penalties, minimum retention counts, and other filters; you must check the order of specific implementations and cannot assume that parameters with the same name on different services are exactly the same. [Transformers Generation Configuration](https://huggingface.co/docs/transformers/en/main_classes/text_generation)

The independent example below reproduces the values in the table and explicitly adopts the order of "temperature first, then top-p"; no model download is required.

```python
import math

def softmax(logits, temperature=1.0):
    if not logits or temperature <= 0:
        raise ValueError("Non-empty logits and positive temperature are required")
    peak = max(logits)
    weights = [math.exp((z - peak) / temperature) for z in logits]
    total = sum(weights)
    return [w / total for w in weights]

def nucleus(probs, threshold):
    if not 0 < threshold <= 1:
        raise ValueError("Threshold should be in (0, 1]")
    kept, mass = [], 0.0
    for i in sorted(range(len(probs)), key=lambda i: (-probs[i], i)):
        kept.append(i)
        mass += probs[i]
        if mass >= threshold:
            break
    return {i: probs[i] / mass for i in kept}

p = softmax([2, 1, 0])
q = nucleus(p, 0.8)
assert list(q) == [0, 1]
assert math.isclose(q[0], 1 / (1 + math.exp(-1)))
assert nucleus(softmax([2, 1, 0], 0.5), 0.8) == {0: 1.0}
assert math.isclose(sum(p), 1.0)
assert len(nucleus(p, 1.0)) == 3
print([round(x, 4) for x in p])
print({i: round(v, 4) for i, v in q.items()})
# [0.6652, 0.2447, 0.09]
# {0: 0.7311, 1: 0.2689}
```

This experiment verifies probability transformations and does not prove that a certain threshold is better for practical tasks. The original research on nucleus sampling discussed probability tails and degeneration issues in open-ended text generation; its results cannot be directly extrapolated to conclude that all tasks such as translation, code, and classification should use the same decoding method. [The Curious Case of Neural Text Degeneration](https://arxiv.org/abs/1904.09751)

### How does probability become a selection?

Treat probabilities as intervals on [0,1): sampling locates u. Filtering must redistribute the intervals; deleting C without renormalization leaves a gap. This visual explains selection from given scores, not the probability of a correct answer.

$$
p_i=\frac{e^{(z_i-\max z)/T}}{\sum_j e^{(z_j-\max z)/T}}
$$

**How does probability become a selection?**

Top-p retains the smallest prefix reaching its mass threshold, then renormalizes; the sample position selects one interval.


## Greedy selection and determinism

Greedy takes the highest-scoring candidate at each step. The division formula above is undefined for T=0; some interfaces treat zero temperature as greedy decoding, while others require disabling sampling via a separate parameter. Mathematical limits and interface conventions should be distinguished.

Greedy only guarantees local selection. Suppose at the first step A=0.6, B=0.4, and after selecting A, the best next step probability is 0.5, while after selecting B, the best next step is 0.9. Then the probabilities of the two candidate paths are 0.30 and 0.36 respectively. Step-by-step greedy would first select A, but it did not obtain the path with the higher probability among the two. Beam search retains multiple candidates to mitigate such issues, but is still limited by search width, length bias, and scoring objectives; the highest sequence probability does not equal factual correctness.

For identical logits, argmax with fixed tie-breaking is deterministic; whether real services can produce exactly the same logits again depends on the model version, precision, kernels, batching, and determinism settings. Fixing the seed only controls the corresponding random number process and cannot lock all execution conditions. Low temperature alone cannot promise byte-by-byte reproducibility, and low effort is not a substitute for determinism.

For auditing or regression testing, save the complete input, model and generation configuration, service version information, and raw output. If you must reproduce previously displayed content, you should read the saved result from that time; this is a different guarantee from re-running the model to get the same answer.

### Does the best local choice give the best path?

Local choice compares candidates at the current prefix. Move B’s continuation past 0.75: whole-path ranking reverses while greedy still starts with A. The comparison concerns probability, not truthfulness.

**Does the best local choice give the best path?**

A starts at 0.6 versus B at 0.4; above 0.75 on B’s next step, B has the higher two-step path probability.


## Output stopping and format constraints

| Control | What it constrains | What it does not guarantee |
|---|---|---|
| EOS or end marker | Model chooses to end | The task has been completed correctly |
| Max generation length | Upper limit of output size | JSON, code blocks are definitely complete |
| stop sequence | Stop at specified boundaries | Will not mistakenly truncate the same string in the body |
| Format or schema constraints | Allowed syntax, fields, or type ranges | Field values are business-true and legal |
| Repetition penalty | Adjusts scores of already appeared tokens | All repetitions are unnecessary |

For example, when generating order JSON, the schema may require `quantity` to be an integer, but it may not know that the warehouse only has 2 items left. Outputting `{"quantity": 5}` can be syntactically valid but business-invalid. The client must still check the stop reason, parse completely, execute field validation, and verify inventory and permissions at the business layer.

Repetition penalties also have contextual costs: variable names in code, JSON fields, or fixed terms often need to appear repeatedly; overly strong penalties may destroy consistency. Before selecting sampling parameters, first confirm the output protocol and task constraints. See [Prompt Engineering and Structured Output](/ai/02-context-and-interfaces/01-prompt-and-output-contracts.md) for the complete process of structured output.

## Re-establishing baselines during migration

When switching models, separately confirm the tokenizer, chat template, supported generation parameters, and their defaults, then compare length, format pass rate, task accuracy, and result variance on the same task set. Do not infer that a parameter is definitely supported or unsupported based on "local" or "hosted"; rely on the specific model and endpoint. [Claude Migration Guide](https://platform.claude.com/docs/en/about-claude/models/migration-guide)

If the new interface no longer provides temperature, first remove unsupported fields, then use clear task constraints and evaluation to confirm output behavior. **Effort adjusts inference investment, temperature adjusts the sampling distribution; the two cannot be numerically mapped.** Asking for "four different plans" expresses the output goal but is not equivalent to increasing temperature; asking to "output only one label" is not equivalent to lowering effort. See [Reasoning and Thinking](/ai/01-model-foundations/04-reasoning-and-verification.md) for independent trade-offs in inference investment.

Continue with：[Attention, feed-forward networks and MoE](/ai/01-model-foundations/02-attention-and-moe)。
