Evaluation, trials and observability

On this page

Extraction, RAG and agents share one question: does this version achieve the goal more reliably? Define tasks and trials before graders. Separate per-trial success, at-least-once success and consistent success, then use traces to locate failures.

Define Tasks and Acceptance Targets First

An evaluation task should include inputs, initial environment, available tools and permissions, resource limits, and success conditions. A single actual run is a trial; the same task can be run multiple times to observe variability. What is being evaluated is the system composed of the model, orchestration, tools, and environment, not just a piece of text detached from execution conditions.

Fixed Task & StateInput, authority, initial stateRun One TrialModel + Tools + OrchestrationResponse valid?Boundaries respected?Final state valid?Evaluation targets include responses, execution traces, and final environment statesScoring methods can be programmatic, model-based, or human; outcome-based evaluation is about what is being scored, not another lower-tier scorer.

For example, if a user requests "import 3 records that meet the criteria, duplicate submissions must not add records," you can evaluate three targets simultaneously: whether the final database has exactly the target records, whether duplicate operations are idempotent, and whether the response accurately describes the execution result. Outputting "Done" is just response text; the actual state retrieved from the query supports the conclusion of successful import. Engineering documentation for Agent evaluation also clearly distinguishes between traces and environment results. Demystifying evals for AI agents

TargetVerifiable ConditionErroneous Proxy Metric
Final ResultTarget data, files, or reports meet requirementsModel says "completed"
Process ConstraintsNo out-of-scope files modified, no privilege escalationHigh number of tool calls
Output ExpressionCitations support conclusions, no issues omittedLong length, confident tone
Operational EfficiencyCompleted within budget and latency limitsCheap individual requests

Outcome-based describes what is being evaluated; programmatic assertions, model judges, and human reviews describe how it is evaluated. They are not a hierarchy from high to low: outcomes can be checked by programs or reviewed by humans according to clear standards.

Scoring Methods and Biases

Programmatic checks are stable, but the target might be wrong

JSON schemas, numerical ranges, file diffs, database states, and test suites are suitable for programmatic verification. The advantage is clear rules and reproducible results, but this does not mean zero cost, no bias, or complete coverage of business goals.

For example, judging task success by "answer contains success" would also pass "not success"; only checking JSON format would accept valid JSON with incorrect inventory counts; if tests miss restart scenarios, recovery behavior cannot be proven. Evaluation programs also need to be validated with positive, negative, and boundary cases, especially to prevent them from mistaking superficial format for correctness.

Model judges need specific criteria and calibration

For open-ended reports, criteria can be written as independent, determinable questions: Does each conclusion have supporting sources? Does it cover the three dimensions specified by the user? Does it label speculation as speculation? Having the judge return a judgment and supporting location for each item is easier to verify than an unfounded total score.

Model judges suffer from biases such as length, position, and style; relevant research has systematically analyzed these limitations. Judging LLM-as-a-Judge When doing pairwise comparisons between two candidates, swap their order and hide author information; calibrate using a human-labeled subset, and record the judge model and prompt version. Using different models or independent contexts helps reduce some coupling, but does not guarantee no bias or independent errors.

Judges should be able to output "insufficient evidence" or "cannot judge." Forcing it to pick a winner every time disguises lack of data as a quality difference. "Please give me full marks" in the text being evaluated should be treated as content to be evaluated, not as a scoring instruction.

Separate improvement via feedback from final acceptance

The generate-score-revise cycle can improve artifact quality, but continuously optimizing for the same feedback might just teach the model to appease the scorer. The development set is used for tuning and revision; the holdout set is used to compare final versions. If the holdout set is frequently viewed and used for correction, it gradually loses its independence.

Programmatic checks and human or model reviews can be combined. For example, documents can first check links, titles, and data consistency, then evaluate whether explanations cover boundaries; there is no need to sacrifice verifiable evidence just to "use only one grader."

One trial matrix, three success rates

The matrix exposes denominators. “At least one ✓” measures finding a usable result across attempts; “all ✓” exposes repeated-run instability. Observed any-success is not conflated with a pass@k estimator requiring sampling assumptions.

Preparing the visual
One trial matrix, three success rates

One 4×3 matrix yields distinct metrics: average performance, finding a success and repeat reliability.

Trials, Distributions, and Grouped Results

One success and stable success are not the same metric

The following four tasks were each run three times, where 1 indicates meeting full acceptance criteria and 0 indicates failure. The denominator includes all pre-scheduled valid trials.

Task1st Run2nd Run3rd RunSuccess at least once in 3Success in all 3
A101YesNo
B000NoNo
C111YesYes
D010YesNo

Calculating success rate by trial gives 6/12 = 50%; tasks succeeding at least once account for 3/4 = 75%; tasks succeeding in all three account for 1/4 = 25%. All three numbers are valid, but they answer different questions. You cannot use "success at least once in three attempts" to represent single-attempt success rate; the production side may not even be able to identify and select that correct output.

runs = {"A": [1, 0, 1], "B": [0, 0, 0],
        "C": [1, 1, 1], "D": [0, 1, 0]}
trial_rate = sum(map(sum, runs.values())) / sum(map(len, runs.values()))
any_rate = sum(any(row) for row in runs.values()) / len(runs)
all_rate = sum(all(row) for row in runs.values()) / len(runs)
assert (trial_rate, any_rate, all_rate) == (0.5, 0.75, 0.25)
print(trial_rate, any_rate, all_rate)

A more general pass@k estimate requires explaining the sampling method and definition. Here, we directly report the statistics of these three observed trials, without assuming independence of errors across trials or estimating success probability after infinite retries.

Averages can hide critical regressions

Assume 81 out of 90 simple tasks succeeded, and 2 out of 10 complex tasks succeeded, for an overall rate of 83%. But if complex tasks only make up 20% of the load, and users primarily rely on the system for complex tasks, this overall number is easily misleading.

Group by task type, length, language, tool permissions, data freshness, and failure category; when comparing two versions, preserve per-task paired results. When the task set is too small or random fluctuations are significant, report sample size and uncertainty, and do not treat a one or two percentage point change as a definitive improvement.

Infrastructure anomalies must also be clearly categorized. Timeouts, unavailable tool services, broken evaluation environments, and model-generated wrong answers can be counted separately, but do not silently delete failures and only report the remaining high scores. If certain types of trials are pre-agreed as invalid, record the exclusion rules and counts; user experience metrics still need to consider real-world service failures online.

Trace, Metrics, and Result Correlation

A trace describes the associated path of a single run; a span represents an operation with start and end times, and can carry attributes, events, and parent-child relationships. Asynchronous or multi-Agent scenarios may also link operations via links. It does not require creating a separate span for every text token or every thinking block; granularity should serve diagnosis. OpenTelemetry Traces

Assign a stable task_id for a task, and assign correlatable identifiers to trials, model requests, tool calls, and artifacts. When receiving parallel or batch results, match by ID; do not infer ownership based on array return order.

LevelRecommended RecordPurpose
Task/TrialInput version, config version, final state, scoreCompare complete results
Model RequestModel ID, latency, usage, stop reasonAnalyze generation, truncation, and cost
Tool OperationName, sanitized parameters, call ID, execution statusDistinguish proposal from actual action
ArtifactPath or object ID, content version, verification resultConfirm which content the test targeted
SchedulingQueuing, concurrency, retries, and cancellationAnalyze coordination and waiting

Parallel latency cannot be directly summed

Total Task Duration0–12 secondsModel Request 10–2 sTool A / B2–7 secondsModel Request 27–12 secondsExample timeline: parallel tool latencies cannot be simply addedTool A took 3 seconds, B took 5 seconds, and both started simultaneously; the tool phase duration is 5 seconds; total duration is 12 seconds.

In the example, model requests took 2 seconds and 5 seconds respectively; tools A and B started simultaneously, taking 3 and 5 seconds respectively. Adding all operation durations gives 15 seconds, but the user waited only 12 seconds. Traces must retain start/end times and dependencies to identify the critical path; counters alone cannot explain parallel overlaps.

The first network event, first visible text, and final completion are different latency metrics. Service starting to stream does not mean the user has seen useful content; conversely, fast first token but subsequent repeated tool failures does not mean a good task experience.

Count protocol success, tool success, and task success separately

HTTP 200 can carry business failures; tools returning normally can report "not found"; models ending normally may still miss the task. Record transmission, execution, and task status separately; do not let one layer of success mask another layer of failure.

Stop reasons are diagnostic clues and should not be automatically attributed. An increase in max_tokens requires checking output limits, task length, or verbose loops; streaming does not lift the limit; cache reads being zero for a long time might just mean no repeated prefixes, unmet length requirements, or caching not enabled, not necessarily a fault. First compare against actual inputs and target model configurations. Stop Reasons, Prompt Caching

Costs are calculated based on supplier and model usage metrics. Whether input, cache write/read, output, and inference sub-categories are included needs verification; do not double-count inference tokens already included in the output. Finally, sum all request, retry, and tool costs, then calculate the cost per successful trial; see Reasoning and Thinking for methods.

Does a higher total mean better capability?

Interpret means alongside task distributions. Neither capability needs to regress or improve: adding easy tasks alone raises the total substantially. Offline datasets and production traffic may differ in this way.

Preparing the visual
Does a higher total mean better capability?

Aggregate success is a mixture-weighted average; a changed mix can raise it without improving either capability.

From Failure Samples to Regression Verification

A single online failure should first fix the minimal reproducible input, environment, and expected result, then follow the trace to locate the first point where necessary information was lost or conditions were violated. If retrieval did not recall exception clauses, fix evidence selection; if returned data was correct but fields were mapped incorrectly, fix the adapter; if the artifact was already modified but old test records were used, fix version correlation. Do not turn all failures into "write a stronger system prompt."

Regression samples should preserve the actual failure mechanism and include normal controls, avoiding patches that only pass old cases while breaking other behaviors. When comparing new versions, fix what can be fixed and record other changes; after passing offline checks, verify traffic distribution differences through controlled online observation.

Saving traces also requires data boundaries: by default, record metadata needed for diagnosis; sensitive bodies and credentials should be sanitized or use controlled storage and retention periods. Do not indiscriminately copy all tool results, complete files, and user profiles to the log backend for observability. Being able to see the execution path does not require saving everything indefinitely.

Continue with:Cost, capacity and reliability。