Extraction, RAG and agents share one question: does this version achieve the goal more reliably? Define tasks and trials before graders. Separate per-trial success, at-least-once success and consistent success, then use traces to locate failures.
Define Tasks and Acceptance Targets First
An evaluation task should include inputs, initial environment, available tools and permissions, resource limits, and success conditions. A single actual run is a trial; the same task can be run multiple times to observe variability. What is being evaluated is the system composed of the model, orchestration, tools, and environment, not just a piece of text detached from execution conditions.
For example, if a user requests "import 3 records that meet the criteria, duplicate submissions must not add records," you can evaluate three targets simultaneously: whether the final database has exactly the target records, whether duplicate operations are idempotent, and whether the response accurately describes the execution result. Outputting "Done" is just response text; the actual state retrieved from the query supports the conclusion of successful import. Engineering documentation for Agent evaluation also clearly distinguishes between traces and environment results. Demystifying evals for AI agents
Target
Verifiable Condition
Erroneous Proxy Metric
Final Result
Target data, files, or reports meet requirements
Model says "completed"
Process Constraints
No out-of-scope files modified, no privilege escalation
High number of tool calls
Output Expression
Citations support conclusions, no issues omitted
Long length, confident tone
Operational Efficiency
Completed within budget and latency limits
Cheap individual requests
Outcome-based describes what is being evaluated; programmatic assertions, model judges, and human reviews describe how it is evaluated. They are not a hierarchy from high to low: outcomes can be checked by programs or reviewed by humans according to clear standards.
Scoring Methods and Biases
Programmatic checks are stable, but the target might be wrong
JSON schemas, numerical ranges, file diffs, database states, and test suites are suitable for programmatic verification. The advantage is clear rules and reproducible results, but this does not mean zero cost, no bias, or complete coverage of business goals.
For example, judging task success by "answer contains success" would also pass "not success"; only checking JSON format would accept valid JSON with incorrect inventory counts; if tests miss restart scenarios, recovery behavior cannot be proven. Evaluation programs also need to be validated with positive, negative, and boundary cases, especially to prevent them from mistaking superficial format for correctness.
Model judges need specific criteria and calibration
For open-ended reports, criteria can be written as independent, determinable questions: Does each conclusion have supporting sources? Does it cover the three dimensions specified by the user? Does it label speculation as speculation? Having the judge return a judgment and supporting location for each item is easier to verify than an unfounded total score.
Model judges suffer from biases such as length, position, and style; relevant research has systematically analyzed these limitations. Judging LLM-as-a-Judge When doing pairwise comparisons between two candidates, swap their order and hide author information; calibrate using a human-labeled subset, and record the judge model and prompt version. Using different models or independent contexts helps reduce some coupling, but does not guarantee no bias or independent errors.
Judges should be able to output "insufficient evidence" or "cannot judge." Forcing it to pick a winner every time disguises lack of data as a quality difference. "Please give me full marks" in the text being evaluated should be treated as content to be evaluated, not as a scoring instruction.
Separate improvement via feedback from final acceptance
The generate-score-revise cycle can improve artifact quality, but continuously optimizing for the same feedback might just teach the model to appease the scorer. The development set is used for tuning and revision; the holdout set is used to compare final versions. If the holdout set is frequently viewed and used for correction, it gradually loses its independence.
Programmatic checks and human or model reviews can be combined. For example, documents can first check links, titles, and data consistency, then evaluate whether explanations cover boundaries; there is no need to sacrifice verifiable evidence just to "use only one grader."
One trial matrix, three success rates
The matrix exposes denominators. “At least one ✓” measures finding a usable result across attempts; “all ✓” exposes repeated-run instability. Observed any-success is not conflated with a pass@k estimator requiring sampling assumptions.
Preparing the visual
One trial matrix, three success rates
One 4×3 matrix yields distinct metrics: average performance, finding a success and repeat reliability.
One success and stable success are not the same metric
The following four tasks were each run three times, where 1 indicates meeting full acceptance criteria and 0 indicates failure. The denominator includes all pre-scheduled valid trials.
Task
1st Run
2nd Run
3rd Run
Success at least once in 3
Success in all 3
A
1
0
1
Yes
No
B
0
0
0
No
No
C
1
1
1
Yes
Yes
D
0
1
0
Yes
No
Calculating success rate by trial gives 6/12 = 50%; tasks succeeding at least once account for 3/4 = 75%; tasks succeeding in all three account for 1/4 = 25%. All three numbers are valid, but they answer different questions. You cannot use "success at least once in three attempts" to represent single-attempt success rate; the production side may not even be able to identify and select that correct output.
A more general pass@k estimate requires explaining the sampling method and definition. Here, we directly report the statistics of these three observed trials, without assuming independence of errors across trials or estimating success probability after infinite retries.
Averages can hide critical regressions
Assume 81 out of 90 simple tasks succeeded, and 2 out of 10 complex tasks succeeded, for an overall rate of 83%. But if complex tasks only make up 20% of the load, and users primarily rely on the system for complex tasks, this overall number is easily misleading.
Group by task type, length, language, tool permissions, data freshness, and failure category; when comparing two versions, preserve per-task paired results. When the task set is too small or random fluctuations are significant, report sample size and uncertainty, and do not treat a one or two percentage point change as a definitive improvement.
Infrastructure anomalies must also be clearly categorized. Timeouts, unavailable tool services, broken evaluation environments, and model-generated wrong answers can be counted separately, but do not silently delete failures and only report the remaining high scores. If certain types of trials are pre-agreed as invalid, record the exclusion rules and counts; user experience metrics still need to consider real-world service failures online.
Trace, Metrics, and Result Correlation
A trace describes the associated path of a single run; a span represents an operation with start and end times, and can carry attributes, events, and parent-child relationships. Asynchronous or multi-Agent scenarios may also link operations via links. It does not require creating a separate span for every text token or every thinking block; granularity should serve diagnosis. OpenTelemetry Traces
Assign a stable task_id for a task, and assign correlatable identifiers to trials, model requests, tool calls, and artifacts. When receiving parallel or batch results, match by ID; do not infer ownership based on array return order.
Level
Recommended Record
Purpose
Task/Trial
Input version, config version, final state, score
Compare complete results
Model Request
Model ID, latency, usage, stop reason
Analyze generation, truncation, and cost
Tool Operation
Name, sanitized parameters, call ID, execution status
Distinguish proposal from actual action
Artifact
Path or object ID, content version, verification result
Confirm which content the test targeted
Scheduling
Queuing, concurrency, retries, and cancellation
Analyze coordination and waiting
Parallel latency cannot be directly summed
In the example, model requests took 2 seconds and 5 seconds respectively; tools A and B started simultaneously, taking 3 and 5 seconds respectively. Adding all operation durations gives 15 seconds, but the user waited only 12 seconds. Traces must retain start/end times and dependencies to identify the critical path; counters alone cannot explain parallel overlaps.
The first network event, first visible text, and final completion are different latency metrics. Service starting to stream does not mean the user has seen useful content; conversely, fast first token but subsequent repeated tool failures does not mean a good task experience.
Count protocol success, tool success, and task success separately
HTTP 200 can carry business failures; tools returning normally can report "not found"; models ending normally may still miss the task. Record transmission, execution, and task status separately; do not let one layer of success mask another layer of failure.
Stop reasons are diagnostic clues and should not be automatically attributed. An increase in max_tokens requires checking output limits, task length, or verbose loops; streaming does not lift the limit; cache reads being zero for a long time might just mean no repeated prefixes, unmet length requirements, or caching not enabled, not necessarily a fault. First compare against actual inputs and target model configurations. Stop Reasons, Prompt Caching
Costs are calculated based on supplier and model usage metrics. Whether input, cache write/read, output, and inference sub-categories are included needs verification; do not double-count inference tokens already included in the output. Finally, sum all request, retry, and tool costs, then calculate the cost per successful trial; see Reasoning and Thinking for methods.
Does a higher total mean better capability?
Interpret means alongside task distributions. Neither capability needs to regress or improve: adding easy tasks alone raises the total substantially. Offline datasets and production traffic may differ in this way.
poverall=wpeasy+(1−w)phard
Preparing the visual
Does a higher total mean better capability?
Aggregate success is a mixture-weighted average; a changed mix can raise it without improving either capability.
{"id":"ai-task-mix","title":"Does a higher total mean better capability?","summary":"Aggregate success is a mixture-weighted average; a changed mix can raise it without improving either capability.","height":1000,"html":"<h2 data-i18n=\"heading\"></h2><p class=\"intro\" data-i18n=\"intro\"></p><div class=\"control\"><label for=\"share\" data-i18n=\"share\"></label><output for=\"share\" id=\"share-value\"></output><input id=\"share\" type=\"range\" min=\"0\" max=\"100\" value=\"90\" step=\"1\"></div><div id=\"mix\"></div><div class=\"metrics\"><div><span data-i18n=\"overall\"></span><strong id=\"overall\"></strong></div><div><span data-i18n=\"hard\"></span><strong id=\"hard\"></strong></div></div><p class=\"formula\" id=\"math\"></p><div class=\"copy\"><p id=\"explanation\" role=\"status\"></p><p class=\"reserve\" aria-hidden=\"true\" data-i18n=\"explain\"></p></div><details class=\"scope\"><summary data-i18n=\"scopeLabel\"></summary><p data-i18n=\"scope\"></p></details>","css":".intro{margin:8px 0 18px;color:var(--muted);font-size:14px}.control{display:grid;grid-template-columns:minmax(0,1fr) auto;gap:6px 12px;align-items:center;margin:14px 0;font-size:13px}.control output{color:var(--accent);text-align:right;font-variant-numeric:tabular-nums;min-width:4em}.control input{grid-column:1/-1;width:100%;margin:0}.presets{display:flex;gap:4px;align-items:stretch}.presets button{flex:1;min-width:0;overflow-wrap:anywhere;font-size:13px;min-height:44px}.actions{display:grid;grid-template-columns:repeat(2,minmax(0,1fr));gap:8px;margin:16px 0}.actions button{min-width:0;min-height:44px;font-size:13px;overflow-wrap:anywhere}.copy{display:grid;font-size:14px;margin:16px 0}.copy>*{grid-area:1/1;margin:0;overflow-wrap:anywhere}.reserve{visibility:hidden}.scope{margin-top:14px;color:var(--muted);font-size:12px}.scope summary{padding:8px 0;cursor:pointer}.scope p{margin-top:10px;overflow-wrap:anywhere}.metrics{display:grid;grid-template-columns:repeat(2,minmax(0,1fr));gap:14px;margin:18px 0}.metrics>div{border-left:2px solid var(--accent);padding:0 6px 0 10px;min-width:0}.metrics span{display:block;min-height:3.6em;color:var(--muted);font-size:12px;overflow-wrap:anywhere}.metrics strong{font-size:21px;display:block;min-height:2em;font-weight:550;overflow-wrap:anywhere;font-variant-numeric:tabular-nums}.check{display:flex;align-items:center;gap:9px;font-size:13px;margin:14px 0}.check input{width:18px;height:18px;flex-shrink:0;accent-color:var(--accent)}select,input[type=text],input[type=number]{background:var(--paper);color:var(--ink);font:inherit;font-size:14px;padding:10px;border:1px solid var(--rule);border-radius:7px;max-width:100%;min-width:0}select{width:100%;min-height:44px}.panel{padding:14px;border:1px solid var(--rule);border-radius:9px;margin:16px 0;min-width:0}.panel>header{font-size:13px;font-weight:600;margin-bottom:12px;color:var(--muted)}.panel>div{overflow-wrap:anywhere}.track-row{margin:16px 0}.track-row>span{display:block;font-size:12px;color:var(--muted);margin-bottom:8px}.track{display:flex;gap:5px;min-width:0}.cell{flex:1;min-width:0;min-height:56px;border:1px solid var(--rule);border-radius:5px;background:var(--surface);display:flex;align-items:center;justify-content:center;text-align:center;padding:6px 3px;font-size:12px;overflow-wrap:anywhere}.cell.on{background:var(--accent-soft);border-color:var(--accent)}.cell.gold{background:var(--second-soft);border-color:var(--second)}.cell.empty{border-style:dashed;color:var(--muted)}.cell.error{border-color:var(--second);text-decoration:line-through}.table{display:grid;gap:5px;font-size:12px}.tr{display:grid;gap:5px}.tr>div{background:var(--surface);border-radius:4px;padding:9px 6px;min-width:0;min-height:5em;overflow-wrap:anywhere;display:flex;align-items:center}.tr>.gold{background:var(--second-soft)}.tr>.on{background:var(--accent-soft)}.formula,.source{font:13px/1.8 ui-monospace,monospace;white-space:pre-wrap;overflow-wrap:anywhere;margin:16px 0}.formula{border-top:1px solid var(--rule);padding-top:12px;min-height:7.2em}.source{min-height:8em;background:var(--surface);padding:12px;border-radius:8px}.intervals{margin:16px 0}.interval{position:relative;height:38px;background:var(--surface);margin:6px 0;border-radius:4px;overflow:hidden}.interval i{position:absolute;height:100%;background:var(--accent-soft);border-left:2px solid var(--accent)}.interval i.gold{background:var(--second-soft);border-color:var(--second)}.interval span{position:relative;z-index:1;font:12px/38px ui-monospace,monospace;padding-left:7px}.bars{display:grid;gap:10px;margin:16px 0}.bar-row{display:grid;grid-template-columns:44px minmax(0,1fr) 62px;gap:8px;align-items:center;font:12px ui-monospace,monospace}.bar-track{height:24px;background:var(--surface);border-radius:4px;overflow:hidden}.bar-track i{display:block;height:100%;background:var(--accent);transition:width .2s}.bar-track i.gold{background:var(--second)}.bar-track i.empty{opacity:.2}.bar-row output{text-align:right}.diagram{display:block;width:100%;height:auto;margin:18px 0}.diagram text{font:13px ui-monospace,monospace;fill:var(--ink)}.node{fill:var(--surface);stroke:var(--rule)}.node.on{fill:var(--accent-soft);stroke:var(--accent)}.node.gold{fill:var(--second-soft);stroke:var(--second)}.edge{stroke:var(--rule);stroke-width:2;fill:none}.edge.on{stroke:var(--accent)}.matrix{display:grid;grid-template-columns:repeat(3,minmax(0,1fr));gap:6px;margin:16px 0}.matrix button{min-width:0;min-height:44px;font:13px ui-monospace,monospace}.matrix button[aria-pressed=true]{background:var(--accent-soft);border-color:var(--accent)}.matrix button.masked{opacity:.35}.legend{font-size:12px;color:var(--muted);min-height:3.6em;overflow-wrap:anywhere}.numeric{font-variant-numeric:tabular-nums}.scope p{line-break:strict}@media(max-width:480px){.presets button,.actions button{font-size:12px}.panel{padding:12px}.metrics{gap:10px}.metrics strong{font-size:19px}.bar-row{grid-template-columns:34px minmax(0,1fr) 58px;gap:6px}}@media(prefers-reduced-motion:reduce){*,*::before,*::after{transition:none!important;animation:none!important}}\n\n.panel>h3{font-size:13px;font-weight:600;margin:0 0 12px;color:var(--muted)}\n\n.bar-row{grid-template-columns:44px minmax(0,1fr) 82px}.bar-row output{white-space:nowrap}@media(max-width:480px){.bar-row{grid-template-columns:34px minmax(0,1fr) 76px}}\n\n","js":"const $=s=>document.querySelector(s);const set=(id,v)=>$('#'+id).textContent=v;const t=k=>viz.t(k);function cells(id,values){$('#'+id).replaceChildren(...values.map(v=>{const e=document.createElement('div');e.className='cell '+(v.cls||'');e.textContent=v.text;e.title=v.title||v.text;return e;}));}function pressed(mode){document.querySelectorAll('[data-mode]').forEach(e=>e.setAttribute('aria-pressed',String(e.dataset.mode===mode)));}const esc=s=>String(s).replace(/[&<>\"']/g,c=>({'&':'&','<':'<','>':'>','\"':'"',\"'\":'''}[c]));const fmt=s=>'{'+[...s].sort().join(', ')+'}';function graph(id,nodes,edges){const el=$('#'+id);el.setAttribute('viewBox','0 0 360 200');el.innerHTML='<defs><marker id=\"arrow-'+id+'\" viewBox=\"0 0 10 10\" refX=\"8\" refY=\"5\" markerWidth=\"5\" markerHeight=\"5\" orient=\"auto-start-reverse\"><path d=\"M 0 0 L 10 5 L 0 10 z\" fill=\"var(--muted)\"/></marker></defs>'+edges.map(e=>{const a=nodes[e[0]],b=nodes[e[1]],dx=b.x-a.x,dy=b.y-a.y,d=Math.hypot(dx,dy)||1;if(edges.some(r=>r[0]===e[1]&&r[1]===e[0])){const nx=-dy/d,ny=dx/d;return '<path class=\"edge '+(e[2]?'on':'')+'\" d=\"M '+(a.x+dx/d*22+nx*12)+' '+(a.y+dy/d*22+ny*12)+' Q '+((a.x+b.x)/2+nx*38)+' '+((a.y+b.y)/2+ny*38)+' '+(b.x-dx/d*25+nx*12)+' '+(b.y-dy/d*25+ny*12)+'\" marker-end=\"url(#arrow-'+id+')\"/>';}return '<line class=\"edge '+(e[2]?'on':'')+'\" x1=\"'+(a.x+dx/d*25)+'\" y1=\"'+(a.y+dy/d*25)+'\" x2=\"'+(b.x-dx/d*28)+'\" y2=\"'+(b.y-dy/d*28)+'\" marker-end=\"url(#arrow-'+id+')\"/>';}).join('')+nodes.map(n=>'<g><circle class=\"node '+(n.cls||'')+'\" cx=\"'+n.x+'\" cy=\"'+n.y+'\" r=\"25\"/><text x=\"'+n.x+'\" y=\"'+(n.y+4)+'\" text-anchor=\"middle\">'+esc(n.name)+'</text>'+(n.sub?'<text x=\"'+n.x+'\" y=\"'+(n.y+43)+'\" text-anchor=\"middle\">'+esc(n.sub)+'</text>':'')+'</g>').join('');}function table(id,rows){$('#'+id).classList.add('table');$('#'+id).replaceChildren(...rows.map(r=>{const row=document.createElement('div');row.className='tr';row.style.gridTemplateColumns='repeat('+r.length+',minmax(0,1fr))';r.forEach(x=>{const c=document.createElement('div');if(typeof x==='object'){c.textContent=x.text;c.className=x.cls||'';}else c.textContent=x;row.append(c)});return row;}));}function band(id,parts,total){$('#'+id).replaceChildren(...parts.map((p,i)=>{const e=document.createElement('i');e.style.width=(100*p.value/total)+'%';e.className=i%2?'gold':'on';e.title=p.name+': '+p.value;return e}));}function intervals(id,items,total){$('#'+id).replaceChildren(...items.map(p=>{const row=document.createElement('div');row.className='interval';const bar=document.createElement('i');bar.style.left=(100*p.start/total)+'%';bar.style.width=(100*(p.end-p.start)/total)+'%';bar.className=p.cls||'';const label=document.createElement('span');label.textContent=p.label;row.append(bar,label);return row;}));}function plot(id,series,xr,yr,opts={}){const e=$('#'+id),X=x=>42+(x-xr[0])/(xr[1]-xr[0])*298,Y=y=>175-(y-yr[0])/(yr[1]-yr[0])*148;let out='<defs><clipPath id=\"clip-'+id+'\"><rect x=\"42\" y=\"27\" width=\"298\" height=\"148\"/></clipPath></defs>';for(let i=0;i<3;i++){let x=xr[0]+i*(xr[1]-xr[0])/2,y=yr[0]+i*(yr[1]-yr[0])/2;out+='<path class=\"gridline\" d=\"M '+X(x)+' 27V175M42 '+Y(y)+'H340\"/><text x=\"'+X(x)+'\" y=\"194\" text-anchor=\"middle\">'+esc(opts.xfmt?opts.xfmt(x):Number(x.toFixed(2)))+'</text><text x=\"36\" y=\"'+(Y(y)+4)+'\" text-anchor=\"end\">'+esc(opts.yfmt?opts.yfmt(y):Number(y.toFixed(2)))+'</text>';}out+='<g clip-path=\"url(#clip-'+id+')\">';for(const s of series){out+='<path fill=\"none\" stroke=\"'+(s.color||'var(--accent)')+'\" stroke-width=\"2.5\" '+(s.dash?'stroke-dasharray=\"5 4\"':'')+' d=\"'+s.pts.map((p,i)=>(i?'L':'M')+X(p[0]).toFixed(2)+','+Y(p[1]).toFixed(2)).join(' ')+'\"/>';}for(const p of opts.points||[])out+='<circle cx=\"'+X(p[0])+'\" cy=\"'+Y(p[1])+'\" r=\"4\" fill=\"var(--second)\" stroke=\"var(--paper)\" stroke-width=\"1.5\"/>';if(opts.cursor!==undefined)out+='<path d=\"M'+X(opts.cursor)+' 27V175\" stroke=\"var(--muted)\" stroke-dasharray=\"3 3\"/>';e.innerHTML=out+'</g>';e.dataset.curves=JSON.stringify(series.map(s=>s.pts));}const curve=(fn,lo,hi,n=120)=>Array.from({length:n+1},(_,i)=>{const x=lo+(hi-lo)*i/n;return[x,fn(x)]});function inputs(draw){document.querySelectorAll('input,select').forEach(e=>{e.addEventListener('input',draw);e.addEventListener('change',draw)});draw()}const n=id=>+$('#'+id).value;const show=(id,v,d=3)=>set(id,Number(v.toFixed(d)));function bars(id,values,max=1){$('#'+id).className='bars';$('#'+id).innerHTML=values.map(v=>'<div class=\"bar-row\"><span>'+esc(v.label)+'</span><div class=\"bar-track\"><i class=\"'+(v.cls||'')+'\" style=\"width:'+Math.max(0,Math.min(100,v.value/max*100))+'%\"></i></div><output>'+esc(v.text??(v.value*100).toFixed(1)+'%')+'</output></div>').join('')}function softmax(z,T=1){let max=Math.max(...z),w=z.map(x=>Math.exp((x-max)/T)),sum=w.reduce((a,b)=>a+b,0);return w.map(x=>x/sum)}function draw(){let w=n('share')/100,p=w*.9+(1-w)*.2;set('share-value',n('share')+'%');bars('mix',[{label:'Easy',value:.9},{label:'Hard',value:.2},{label:'All',value:p,cls:'gold'}]);set('overall',(100*p).toFixed(1)+'%');set('hard','20%');set('math',w.toFixed(2)+' × 0.90 + '+(1-w).toFixed(2)+' × 0.20 = '+p.toFixed(3));set('explanation',t('explain'));document.body.dataset.overall=p;}inputs(draw);","audio":false,"strings":{"explain":"Raising easy-task share from 50% to 90% moves the total from 55% to 83%; hard-task success stays 20%. Compare group results and fix or explicitly reweight the evaluation mix.","hard":"Hard-group success","heading":"Does a higher total mean better capability?","intro":"Change only the easy-task share; keep both group success rates fixed.","overall":"Aggregate success","reset":"Reset","scope":"Easy success is 90%, hard success 20%; deterministic mixture means only, no sampling noise.","scopeLabel":"Model scope and assumptions","share":"Easy-task share"}}
From Failure Samples to Regression Verification
A single online failure should first fix the minimal reproducible input, environment, and expected result, then follow the trace to locate the first point where necessary information was lost or conditions were violated. If retrieval did not recall exception clauses, fix evidence selection; if returned data was correct but fields were mapped incorrectly, fix the adapter; if the artifact was already modified but old test records were used, fix version correlation. Do not turn all failures into "write a stronger system prompt."
Regression samples should preserve the actual failure mechanism and include normal controls, avoiding patches that only pass old cases while breaking other behaviors. When comparing new versions, fix what can be fixed and record other changes; after passing offline checks, verify traffic distribution differences through controlled online observation.
Saving traces also requires data boundaries: by default, record metadata needed for diagnosis; sensitive bodies and credentials should be sanitized or use controlled storage and retention periods. Do not indiscriminately copy all tool results, complete files, and user profiles to the log backend for observability. Being able to see the execution path does not require saving everything indefinitely.