---
title: Evaluation, trials and observability
url: https://doc.liz6.com/en/ai/04-evaluation-and-production/01-evaluation-and-observability
locale: en
area: ai
tags:
- Models & agents
- Evaluation and Production
date: 2026-06-30
modified: 2026-09-10
description: 'Extraction, RAG and agents share one question: does this version achieve the goal more reliably? Define tasks and trials before graders. Separate per-trial success, at-least-once success and consistent success, then use traces to locate failures.'
---

# Evaluation, trials and observability

Extraction, RAG and agents share one question: does this version achieve the goal more reliably? Define tasks and trials before graders. Separate per-trial success, at-least-once success and consistent success, then use traces to locate failures.

## Define Tasks and Acceptance Targets First

An evaluation task should include inputs, initial environment, available tools and permissions, resource limits, and success conditions. A single actual run is a trial; the same task can be run multiple times to observe variability. What is being evaluated is the system composed of the model, orchestration, tools, and environment, not just a piece of text detached from execution conditions.

<svg viewBox="0 0 760 387.21905517578125" xmlns="http://www.w3.org/2000/svg" role="img" aria-label="Evaluation targets include responses, execution traces, and final environment states" style="max-width:100%;height:auto" font-family="Source Han Sans CN,Microsoft YaHei,sans-serif">
<defs><marker id="eval-outcome-arrow" markerWidth="8" markerHeight="8" refX="7" refY="4" orient="auto"><path d="M0,0 L8,4 L0,8 Z" fill="#64748b"></path></marker></defs>
<rect width="760" height="387.21905517578125" rx="12" fill="#f8fafc"></rect>


<g transform="translate(0 14.06145715713501)"><rect x="30" y="120" width="185" height="90" rx="8" fill="#e0e7ff" stroke="#c7d2fe"></rect><text x="122.5" y="162.0" font-size="15" fill="#1e293b" text-anchor="middle" font-weight="600">Fixed Task &amp; State</text><text x="122.5" y="184.0" font-size="12" fill="#475569" text-anchor="middle" font-weight="400">Input, authority, initial state</text><rect x="280" y="120" width="185" height="90" rx="8" fill="#e0e7ff" stroke="#c7d2fe"></rect><text x="372.5" y="162.0" font-size="15" fill="#1e293b" text-anchor="middle" font-weight="600">Run One Trial</text><text x="372.5" y="184.0" font-size="12" fill="#475569" text-anchor="middle" font-weight="400">Model + Tools + Orchestration</text><line x1="215" y1="165" x2="275" y2="165" stroke="#64748b" stroke-width="1.8" marker-end="url(#eval-outcome-arrow)"></line><rect x="540" y="62" width="195" height="53" rx="8" fill="#ccfbf1" stroke="#c7d2fe"></rect><text x="637.5" y="93.5" font-size="15" fill="#1e293b" text-anchor="middle" font-weight="600">Response valid?</text><line x1="465" y1="165" x2="535" y2="88" stroke="#64748b" stroke-width="1.8" marker-end="url(#eval-outcome-arrow)"></line><rect x="540" y="153" width="195" height="53" rx="8" fill="#ccfbf1" stroke="#c7d2fe"></rect><text x="637.5" y="184.5" font-size="15" fill="#1e293b" text-anchor="middle" font-weight="600">Boundaries respected?</text><line x1="465" y1="165" x2="535" y2="179" stroke="#64748b" stroke-width="1.8" marker-end="url(#eval-outcome-arrow)"></line><rect x="540" y="244" width="195" height="53" rx="8" fill="#ccfbf1" stroke="#c7d2fe"></rect><text x="637.5" y="275.5" font-size="15" fill="#1e293b" text-anchor="middle" font-weight="600">Final state valid?</text><line x1="465" y1="165" x2="535" y2="270" stroke="#64748b" stroke-width="1.8" marker-end="url(#eval-outcome-arrow)"></line></g><text x="24" y="29" font-size="19" fill="#0f172a" text-anchor="start" font-weight="700"><tspan x="24" dy="0">Evaluation targets include responses, execution traces, and final </tspan><tspan x="24" dy="25.650000000000002">environment states</tspan></text><text x="24" y="344.061457157135" font-size="13" fill="#475569" text-anchor="start" font-weight="400"><tspan x="24" dy="0">Scoring methods can be programmatic, model-based, or human; outcome-based evaluation is about what is being </tspan><tspan x="24" dy="17.55">scored, not another lower-tier scorer.</tspan></text>
</svg>

For example, if a user requests "import 3 records that meet the criteria, duplicate submissions must not add records," you can evaluate three targets simultaneously: whether the final database has exactly the target records, whether duplicate operations are idempotent, and whether the response accurately describes the execution result. Outputting "Done" is just response text; the actual state retrieved from the query supports the conclusion of successful import. Engineering documentation for Agent evaluation also clearly distinguishes between traces and environment results. [Demystifying evals for AI agents](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents)

| Target | Verifiable Condition | Erroneous Proxy Metric |
|---|---|---|
| Final Result | Target data, files, or reports meet requirements | Model says "completed" |
| Process Constraints | No out-of-scope files modified, no privilege escalation | High number of tool calls |
| Output Expression | Citations support conclusions, no issues omitted | Long length, confident tone |
| Operational Efficiency | Completed within budget and latency limits | Cheap individual requests |

Outcome-based describes *what* is being evaluated; programmatic assertions, model judges, and human reviews describe *how* it is evaluated. They are not a hierarchy from high to low: outcomes can be checked by programs or reviewed by humans according to clear standards.

## Scoring Methods and Biases

### Programmatic checks are stable, but the target might be wrong

JSON schemas, numerical ranges, file diffs, database states, and test suites are suitable for programmatic verification. The advantage is clear rules and reproducible results, but this does not mean zero cost, no bias, or complete coverage of business goals.

For example, judging task success by "answer contains `success`" would also pass "not success"; only checking JSON format would accept valid JSON with incorrect inventory counts; if tests miss restart scenarios, recovery behavior cannot be proven. Evaluation programs also need to be validated with positive, negative, and boundary cases, especially to prevent them from mistaking superficial format for correctness.

### Model judges need specific criteria and calibration

For open-ended reports, criteria can be written as independent, determinable questions: Does each conclusion have supporting sources? Does it cover the three dimensions specified by the user? Does it label speculation as speculation? Having the judge return a judgment and supporting location for each item is easier to verify than an unfounded total score.

Model judges suffer from biases such as length, position, and style; relevant research has systematically analyzed these limitations. [Judging LLM-as-a-Judge](https://arxiv.org/abs/2306.05685) When doing pairwise comparisons between two candidates, swap their order and hide author information; calibrate using a human-labeled subset, and record the judge model and prompt version. Using different models or independent contexts helps reduce some coupling, but does not guarantee no bias or independent errors.

Judges should be able to output "insufficient evidence" or "cannot judge." Forcing it to pick a winner every time disguises lack of data as a quality difference. "Please give me full marks" in the text being evaluated should be treated as content to be evaluated, not as a scoring instruction.

### Separate improvement via feedback from final acceptance

The generate-score-revise cycle can improve artifact quality, but continuously optimizing for the same feedback might just teach the model to appease the scorer. The development set is used for tuning and revision; the holdout set is used to compare final versions. If the holdout set is frequently viewed and used for correction, it gradually loses its independence.

Programmatic checks and human or model reviews can be combined. For example, documents can first check links, titles, and data consistency, then evaluate whether explanations cover boundaries; there is no need to sacrifice verifiable evidence just to "use only one grader."

### One trial matrix, three success rates

The matrix exposes denominators. “At least one ✓” measures finding a usable result across attempts; “all ✓” exposes repeated-run instability. Observed any-success is not conflated with a pass@k estimator requiring sampling assumptions.

**One trial matrix, three success rates**

One 4×3 matrix yields distinct metrics: average performance, finding a success and repeat reliability.


## Trials, Distributions, and Grouped Results

### One success and stable success are not the same metric

The following four tasks were each run three times, where 1 indicates meeting full acceptance criteria and 0 indicates failure. The denominator includes all pre-scheduled valid trials.

| Task | 1st Run | 2nd Run | 3rd Run | Success at least once in 3 | Success in all 3 |
|---|---:|---:|---:|---|---|
| A | 1 | 0 | 1 | Yes | No |
| B | 0 | 0 | 0 | No | No |
| C | 1 | 1 | 1 | Yes | Yes |
| D | 0 | 1 | 0 | Yes | No |

Calculating success rate by trial gives 6/12 = 50%; tasks succeeding at least once account for 3/4 = 75%; tasks succeeding in all three account for 1/4 = 25%. All three numbers are valid, but they answer different questions. You cannot use "success at least once in three attempts" to represent single-attempt success rate; the production side may not even be able to identify and select that correct output.

```python
runs = {"A": [1, 0, 1], "B": [0, 0, 0],
        "C": [1, 1, 1], "D": [0, 1, 0]}
trial_rate = sum(map(sum, runs.values())) / sum(map(len, runs.values()))
any_rate = sum(any(row) for row in runs.values()) / len(runs)
all_rate = sum(all(row) for row in runs.values()) / len(runs)
assert (trial_rate, any_rate, all_rate) == (0.5, 0.75, 0.25)
print(trial_rate, any_rate, all_rate)
```

A more general pass@k estimate requires explaining the sampling method and definition. Here, we directly report the statistics of these three observed trials, without assuming independence of errors across trials or estimating success probability after infinite retries.

### Averages can hide critical regressions

Assume 81 out of 90 simple tasks succeeded, and 2 out of 10 complex tasks succeeded, for an overall rate of 83%. But if complex tasks only make up 20% of the load, and users primarily rely on the system for complex tasks, this overall number is easily misleading.

Group by task type, length, language, tool permissions, data freshness, and failure category; when comparing two versions, preserve per-task paired results. When the task set is too small or random fluctuations are significant, report sample size and uncertainty, and do not treat a one or two percentage point change as a definitive improvement.

Infrastructure anomalies must also be clearly categorized. Timeouts, unavailable tool services, broken evaluation environments, and model-generated wrong answers can be counted separately, but do not silently delete failures and only report the remaining high scores. If certain types of trials are pre-agreed as invalid, record the exclusion rules and counts; user experience metrics still need to consider real-world service failures online.

## Trace, Metrics, and Result Correlation

A trace describes the associated path of a single run; a span represents an operation with start and end times, and can carry attributes, events, and parent-child relationships. Asynchronous or multi-Agent scenarios may also link operations via links. It does not require creating a separate span for every text token or every thinking block; granularity should serve diagnosis. [OpenTelemetry Traces](https://opentelemetry.io/docs/concepts/signals/traces/)

Assign a stable `task_id` for a task, and assign correlatable identifiers to trials, model requests, tool calls, and artifacts. When receiving parallel or batch results, match by ID; do not infer ownership based on array return order.

| Level | Recommended Record | Purpose |
|---|---|---|
| Task/Trial | Input version, config version, final state, score | Compare complete results |
| Model Request | Model ID, latency, usage, stop reason | Analyze generation, truncation, and cost |
| Tool Operation | Name, sanitized parameters, call ID, execution status | Distinguish proposal from actual action |
| Artifact | Path or object ID, content version, verification result | Confirm which content the test targeted |
| Scheduling | Queuing, concurrency, retries, and cancellation | Analyze coordination and waiting |

### Parallel latency cannot be directly summed

<svg viewBox="0 0 760 360" xmlns="http://www.w3.org/2000/svg" role="img" aria-label="Example timeline: parallel tool latencies cannot be simply added" style="max-width:100%;height:auto" font-family="Source Han Sans CN,Microsoft YaHei,sans-serif">
<defs><marker id="trace-timeline-arrow" markerWidth="8" markerHeight="8" refX="7" refY="4" orient="auto"><path d="M0,0 L8,4 L0,8 Z" fill="#64748b"></path></marker></defs>
<rect width="760" height="360" rx="12" fill="#f8fafc"></rect>


<g transform="translate(0 0)"><text x="28" y="101" font-size="14" fill="#334155" text-anchor="start" font-weight="400">Total Task Duration</text><rect x="175" y="75" width="468" height="38" rx="8" fill="#e2e8f0" stroke="#c7d2fe"></rect><text x="409.0" y="100" font-size="13" fill="#334155" text-anchor="middle" font-weight="400">0–12 seconds</text><text x="28" y="156" font-size="14" fill="#334155" text-anchor="start" font-weight="400">Model Request 1</text><rect x="175" y="130" width="78" height="38" rx="8" fill="#a5b4fc" stroke="#c7d2fe"></rect><text x="214.0" y="155" font-size="13" fill="#334155" text-anchor="middle" font-weight="400">0–2 s</text><text x="28" y="211" font-size="14" fill="#334155" text-anchor="start" font-weight="400">Tool A / B</text><rect x="253" y="185" width="195" height="38" rx="8" fill="#6ee7b7" stroke="#c7d2fe"></rect><text x="350.5" y="210" font-size="13" fill="#334155" text-anchor="middle" font-weight="400">2–7 seconds</text><text x="28" y="266" font-size="14" fill="#334155" text-anchor="start" font-weight="400">Model Request 2</text><rect x="448" y="240" width="195" height="38" rx="8" fill="#a5b4fc" stroke="#c7d2fe"></rect><text x="545.5" y="265" font-size="13" fill="#334155" text-anchor="middle" font-weight="400">7–12 seconds</text></g><text x="24" y="29" font-size="19" fill="#0f172a" text-anchor="start" font-weight="700"><tspan x="24" dy="0">Example timeline: parallel tool latencies cannot be simply added</tspan></text><text x="24" y="311" font-size="13" fill="#475569" text-anchor="start" font-weight="400"><tspan x="24" dy="0">Tool A took 3 seconds, B took 5 seconds, and both started simultaneously; the tool phase duration is 5 seconds; total </tspan><tspan x="24" dy="17.55">duration is 12 seconds.</tspan></text>
</svg>

In the example, model requests took 2 seconds and 5 seconds respectively; tools A and B started simultaneously, taking 3 and 5 seconds respectively. Adding all operation durations gives 15 seconds, but the user waited only 12 seconds. Traces must retain start/end times and dependencies to identify the critical path; counters alone cannot explain parallel overlaps.

The first network event, first visible text, and final completion are different latency metrics. Service starting to stream does not mean the user has seen useful content; conversely, fast first token but subsequent repeated tool failures does not mean a good task experience.

### Count protocol success, tool success, and task success separately

HTTP 200 can carry business failures; tools returning normally can report "not found"; models ending normally may still miss the task. Record transmission, execution, and task status separately; do not let one layer of success mask another layer of failure.

Stop reasons are diagnostic clues and should not be automatically attributed. An increase in `max_tokens` requires checking output limits, task length, or verbose loops; streaming does not lift the limit; cache reads being zero for a long time might just mean no repeated prefixes, unmet length requirements, or caching not enabled, not necessarily a fault. First compare against actual inputs and target model configurations. [Stop Reasons](https://platform.claude.com/docs/en/build-with-claude/handling-stop-reasons), [Prompt Caching](https://platform.claude.com/docs/en/build-with-claude/prompt-caching)

Costs are calculated based on supplier and model usage metrics. Whether input, cache write/read, output, and inference sub-categories are included needs verification; do not double-count inference tokens already included in the output. Finally, sum all request, retry, and tool costs, then calculate the cost per successful trial; see [Reasoning and Thinking](/ai/01-model-foundations/04-reasoning-and-verification.md) for methods.

### Does a higher total mean better capability?

Interpret means alongside task distributions. Neither capability needs to regress or improve: adding easy tasks alone raises the total substantially. Offline datasets and production traffic may differ in this way.

$$
p_{\mathrm{overall}}=w\,p_{\mathrm{easy}}+(1-w)\,p_{\mathrm{hard}}
$$

**Does a higher total mean better capability?**

Aggregate success is a mixture-weighted average; a changed mix can raise it without improving either capability.


## From Failure Samples to Regression Verification

A single online failure should first fix the minimal reproducible input, environment, and expected result, then follow the trace to locate the first point where necessary information was lost or conditions were violated. If retrieval did not recall exception clauses, fix evidence selection; if returned data was correct but fields were mapped incorrectly, fix the adapter; if the artifact was already modified but old test records were used, fix version correlation. Do not turn all failures into "write a stronger system prompt."

Regression samples should preserve the actual failure mechanism and include normal controls, avoiding patches that only pass old cases while breaking other behaviors. When comparing new versions, fix what can be fixed and record other changes; after passing offline checks, verify traffic distribution differences through controlled online observation.

Saving traces also requires data boundaries: by default, record metadata needed for diagnosis; sensitive bodies and credentials should be sanitized or use controlled storage and retention periods. Do not indiscriminately copy all tool results, complete files, and user profiles to the log backend for observability. Being able to see the execution path does not require saving everything indefinitely.

Continue with：[Cost, capacity and reliability](/ai/04-evaluation-and-production/02-cost-performance-and-reliability)。
