Tokens, probabilities and sampling

On this page

How does text become model input, and how do scores become the next token? Follow text → token IDs → logits → probabilities → selection, then compare temperature and candidate filtering. The numerical examples explain decoding; task results determine which settings are useful.

From text to token ID

Tokens are vocabulary units, not equal to characters or words

Text models typically encode strings into sequences of integers, where each integer corresponds to a token in the vocabulary. A token can represent a complete word, part of a word, punctuation, a combination of spaces, or a byte fragment. A single Unicode character may span multiple tokens; conversely, multiple characters may combine into a single token. Therefore, you cannot calculate the input for all models by simply multiplying the number of Chinese characters by two.

BPE (Byte-Pair Encoding) is a common subword construction method. During training, it starts with smaller units and repeatedly merges adjacent units based on corpus statistics. Encoding uses the already trained vocabulary and merge rules, rather than retraining for each input. For example, assuming it has learned to merge in order (l,o)→lo and (lo,w)→low, l o w e r would be tokenized as low e r. This is merely a mechanism example and does not represent the actual tokenization of lower by any existing model.

Byte-level BPE uses a base representation that covers bytes, avoiding issues where a pure character vocabulary lacks certain characters, but the name "BPE" itself does not guarantee byte fallback. Tokenization may also use methods such as Unigram or WordPiece. Normalization, pre-tokenization, and special tokens in the configuration also affect the final result. Tokenizers Model Types

Why counting only the body text is insufficient

Chat systems need to encode roles and message boundaries into the model input. Tool definitions, structured content, images, etc., may also generate additional counts. The string length seen by the user and the model input length belong to different representation layers: first determine how the complete request is constructed, then calculate its token count.

For example, two user messages and combining two texts into one message may have identical body character counts, but the message boundaries may differ. If the tokenizer includes normalization processing, Unicode character sequences that look identical or similar may produce different results; do not assume all tokenizers preserve or eliminate such differences. Tokenization Pipeline

Server-side counts should use the interface corresponding to the target model and be calibrated with actual response usage. Claude's counting documentation also states that counting results are estimates, and counts should be recalculated after model updates; there is no cross-model conversion multiplier that applies to all Chinese, code, and tool schemas. Claude Token counting

From logits to probabilities

Standard autoregressive text generation predicts the next token given the current prefix. The model outputs scores (logits) for each position in the vocabulary; the decoder then selects an ID based on the configuration, appends it to the prefix, and continues. Actual inference engines may use batching or speculative decoding to optimize execution, but the semantics of conditional generation can still be understood through this step-by-step process.

Current token sequencePrompt + Generated contentModel forward passOutput logits for next positionProcessing and selectionTemperature, filtering, samplingSelected token IDDecode into visible textOne autoregressive generation: scores, candidate sets, and selection are different stepsStopping conditions are determined by end markers, length limits, or interface rules; visible text blocks do not necessarily equal one token.

For positive temperature , a numerically stable way to write softmax is:

pᵢ = exp((zᵢ − max(z))/T) / Σⱼ exp((zⱼ − max(z))/T)

Subtracting the maximum value does not change the normalized probabilities but avoids directly calculating huge exponents. Suppose there are only three candidates A, B, C, with logits [2, 1, 0]:

TemperatureABC
T=166.52%24.47%9.00%
T=0.586.68%11.73%1.59%
T = 1A 66.52% B 24.47% C 9.00%T = 0.5A 86.68% B 11.73% C 1.59%Same logits = [2, 1, 0], temperature changes probability concentrationCooling down does not add new evidence, nor does it improve model knowledge; it makes existing high-scoring candidates easier to select.

Cooling down amplifies relative score differences, while heating up flattens the distribution. This changes the randomness of selection, it is not a "correctness knob": if the wrong answer originally had the highest score, cooling down will select it more stably. The entire response is not written after a single sample; changes in early tokens will alter the subsequent conditional distribution.

Candidate truncation: top-k and top-p

top-k retains the k candidates with the highest scores; top-p first sorts by probability, retains the smallest prefix where the cumulative probability first reaches the threshold p, and then renormalizes within the retained set. The size of the latter set changes with the distribution.

For the T=1 distribution above, with top-p=0.8, A alone has only 0.6652, which is not enough; adding B reaches approximately 0.9100, so A and B are retained, and C is discarded. Renormalizing gives A≈0.7311, B≈0.2689. top-p=0.8 does not mean retaining 80% of the vocabulary, nor does it mean the highest candidate has an 80% probability.

At T=0.5, A alone exceeds 0.8, so with the same threshold, only A is retained. Therefore, temperature and filtering order jointly affect the results. Different engines may also combine penalties, minimum retention counts, and other filters; you must check the order of specific implementations and cannot assume that parameters with the same name on different services are exactly the same. Transformers Generation Configuration

The independent example below reproduces the values in the table and explicitly adopts the order of "temperature first, then top-p"; no model download is required.

import math

def softmax(logits, temperature=1.0):
    if not logits or temperature <= 0:
        raise ValueError("Non-empty logits and positive temperature are required")
    peak = max(logits)
    weights = [math.exp((z - peak) / temperature) for z in logits]
    total = sum(weights)
    return [w / total for w in weights]

def nucleus(probs, threshold):
    if not 0 < threshold <= 1:
        raise ValueError("Threshold should be in (0, 1]")
    kept, mass = [], 0.0
    for i in sorted(range(len(probs)), key=lambda i: (-probs[i], i)):
        kept.append(i)
        mass += probs[i]
        if mass >= threshold:
            break
    return {i: probs[i] / mass for i in kept}

p = softmax([2, 1, 0])
q = nucleus(p, 0.8)
assert list(q) == [0, 1]
assert math.isclose(q[0], 1 / (1 + math.exp(-1)))
assert nucleus(softmax([2, 1, 0], 0.5), 0.8) == {0: 1.0}
assert math.isclose(sum(p), 1.0)
assert len(nucleus(p, 1.0)) == 3
print([round(x, 4) for x in p])
print({i: round(v, 4) for i, v in q.items()})
# [0.6652, 0.2447, 0.09]
# {0: 0.7311, 1: 0.2689}

This experiment verifies probability transformations and does not prove that a certain threshold is better for practical tasks. The original research on nucleus sampling discussed probability tails and degeneration issues in open-ended text generation; its results cannot be directly extrapolated to conclude that all tasks such as translation, code, and classification should use the same decoding method. The Curious Case of Neural Text Degeneration

How does probability become a selection?

Treat probabilities as intervals on [0,1): sampling locates u. Filtering must redistribute the intervals; deleting C without renormalization leaves a gap. This visual explains selection from given scores, not the probability of a correct answer.

Preparing the visual
How does probability become a selection?

Top-p retains the smallest prefix reaching its mass threshold, then renormalizes; the sample position selects one interval.

Greedy selection and determinism

Greedy takes the highest-scoring candidate at each step. The division formula above is undefined for T=0; some interfaces treat zero temperature as greedy decoding, while others require disabling sampling via a separate parameter. Mathematical limits and interface conventions should be distinguished.

Greedy only guarantees local selection. Suppose at the first step A=0.6, B=0.4, and after selecting A, the best next step probability is 0.5, while after selecting B, the best next step is 0.9. Then the probabilities of the two candidate paths are 0.30 and 0.36 respectively. Step-by-step greedy would first select A, but it did not obtain the path with the higher probability among the two. Beam search retains multiple candidates to mitigate such issues, but is still limited by search width, length bias, and scoring objectives; the highest sequence probability does not equal factual correctness.

For identical logits, argmax with fixed tie-breaking is deterministic; whether real services can produce exactly the same logits again depends on the model version, precision, kernels, batching, and determinism settings. Fixing the seed only controls the corresponding random number process and cannot lock all execution conditions. Low temperature alone cannot promise byte-by-byte reproducibility, and low effort is not a substitute for determinism.

For auditing or regression testing, save the complete input, model and generation configuration, service version information, and raw output. If you must reproduce previously displayed content, you should read the saved result from that time; this is a different guarantee from re-running the model to get the same answer.

Does the best local choice give the best path?

Local choice compares candidates at the current prefix. Move B’s continuation past 0.75: whole-path ranking reverses while greedy still starts with A. The comparison concerns probability, not truthfulness.

Preparing the visual
Does the best local choice give the best path?

A starts at 0.6 versus B at 0.4; above 0.75 on B’s next step, B has the higher two-step path probability.

Output stopping and format constraints

ControlWhat it constrainsWhat it does not guarantee
EOS or end markerModel chooses to endThe task has been completed correctly
Max generation lengthUpper limit of output sizeJSON, code blocks are definitely complete
stop sequenceStop at specified boundariesWill not mistakenly truncate the same string in the body
Format or schema constraintsAllowed syntax, fields, or type rangesField values are business-true and legal
Repetition penaltyAdjusts scores of already appeared tokensAll repetitions are unnecessary

For example, when generating order JSON, the schema may require quantity to be an integer, but it may not know that the warehouse only has 2 items left. Outputting {"quantity": 5} can be syntactically valid but business-invalid. The client must still check the stop reason, parse completely, execute field validation, and verify inventory and permissions at the business layer.

Repetition penalties also have contextual costs: variable names in code, JSON fields, or fixed terms often need to appear repeatedly; overly strong penalties may destroy consistency. Before selecting sampling parameters, first confirm the output protocol and task constraints. See Prompt Engineering and Structured Output for the complete process of structured output.

Re-establishing baselines during migration

When switching models, separately confirm the tokenizer, chat template, supported generation parameters, and their defaults, then compare length, format pass rate, task accuracy, and result variance on the same task set. Do not infer that a parameter is definitely supported or unsupported based on "local" or "hosted"; rely on the specific model and endpoint. Claude Migration Guide

If the new interface no longer provides temperature, first remove unsupported fields, then use clear task constraints and evaluation to confirm output behavior. Effort adjusts inference investment, temperature adjusts the sampling distribution; the two cannot be numerically mapped. Asking for "four different plans" expresses the output goal but is not equivalent to increasing temperature; asking to "output only one label" is not equivalent to lowering effort. See Reasoning and Thinking for independent trade-offs in inference investment.

Continue with:Attention, feed-forward networks and MoE。