A release combines code, model, prompts, tools, policy and indexes. Assemble evidence for that exact combination before choosing a canary, wider rollout or stop. This article integrates validation; detailed formulas and execution rules remain in their dedicated articles.
Fixed Candidate Versions and Task Boundaries
First, clearly define what the system accomplishes and how completion is confirmed. Query-based assistants may require answers with verifiable sources; execution-based assistants also need to check actual business states. A completed model response, HTTP 200, or tools not throwing exceptions cannot independently prove task success. Manual processing is also a distinct result branch; its proportion and cost should be tracked and cannot be silently counted as automatic success.
Candidate versions must at least allow locating application code, model identifiers and parameters, prompt templates, tool schemas, and permission policies. When using retrieval, document document snapshots, chunking and embedding configurations, and index versions; when using persistent state, document state formats and migration rules. Do not just assign a version number to the prompt while letting other dependencies change silently after verification. For services that only support floating model aliases, record the actually obtainable version information and verification time, and acknowledge that changes by the provider reduce reproducibility.
Record the responsible person, candidate version, verification time, evidence location, and conclusion for each item. Record not executed and failed items separately, but neither can masquerade as passed; for not applicable items, state the reason, and do not simply delete the checkboxes.
Quality and Model Configuration
Check Item
Evidence to Retain
Signals That Cannot Replace It
[ ] Define task success and failure
Output requirements, business post-conditions, rejection/manual transfer rules
Model self-statement "completed"
[ ] Cover representative tasks and boundaries
Fixed evaluation set, grouped results by difficulty, language, tool type, etc.
Single demonstration or overall average score
[ ] Retain independent verification set
Split between tuning set and verification set, repeated trial strategy
Repeatedly tuning to full score on the same batch of questions
[ ] Calibrate scoring methods
Program check coverage, manual review, judge misjudgment examples
"Used program scoring so it must be objective"
[ ] Verify model capabilities and parameters
Parameters actually supported by the selected endpoint, verification of stop/reject/truncation branches
Setting high effort uniformly for all models
[ ] Verify structure and business constraints
Schema validation and negative examples for quantity, status, resource ownership, etc.
JSON can be parsed
The "what to observe" and "who scores" in evaluation are two dimensions: business results can be checked by programs or evaluated by humans or model judges; do not treat outcome-based evaluation as a lower-level method placed below program scoring. One success does not prove stable repeated execution; the number of trials per task and aggregation methods should be stated. Anthropic: Agent Evaluation
Total budget for tool definitions, retrieval materials, history, and output reserves
Single document fits, but complete request exceeds limit
[ ] Verify retrieval and answer layers
Candidate recall, ranking, citation correspondence, and behavior when no evidence exists
Answer is fluent but citations do not support the conclusion
[ ] Check version and permission changes
Index and cache tests after document updates, deletions, and permission revocations
Original text revoked, but old snippets still returned
[ ] Verify session recovery
Recovery records for completed steps, pending results, and external object IDs
After recovery, unknown results are treated as unexecuted
[ ] Verify state compatibility
Old state reading, new state writing, and rollback compatibility scope
Rolled-back application cannot read new checkpoints
Taking "prohibit cross-tenant reading" as an example, do not just check if the model finally outputs someone else's content; also confirm that retrieval and tool execution phases did not bring this content into the context. Caches and long-term memory are also access paths. Cache hits do not prove the answer used the latest data; cache hit rate targets need to be judged together with freshness requirements.
Stream request establishment success counts as completion
The names and statistical methods of observability fields vary by service. Based on interface documentation, clarify which usage fields include or exclude cache, inference, and other usage; do not simply add all numbers together. Cache reads being zero for a long time may come from requests not meeting conditions, traffic characteristics, or configuration issues; do not judge the cause of failure based solely on this value.
Existing task authorizations should be passed by the system along the operation chain; actions requiring additional confirmation should be handled according to specific policies; successive manual clicks do not replace resource permissions and idempotency mechanisms. See Security and Protection, Cost, Performance, and Reliability. Boundaries for access protocols and tool discovery can be found in MCP and Skills.
Evidence Status and Version Matching
Four states can be used: pass (meeting agreed conditions), fail (not meeting), not_run (not executed), not_applicable (reasonably and confirmed by responsible process as not applicable). The existence of a report does not mean it passed, and a passed report does not mean it belongs to the current candidate. For necessary checks, missing, failed, not executed, or version mismatch should all prevent automatic release.
The local teaching example below only demonstrates this rule and will not deploy anything. Abstract combinations of multiple actual dependency versions into candidate; real systems also need to verify evidence source, content, timeliness, and scope of application, and cannot trust arbitrary callers' self-reported pass.
fromcopyimportdeepcopyREQUIRED={"quality","authorization","recovery"}defblockers(candidate,evidence):problems=[]fornameinsorted(REQUIRED):item=evidence.get(name)ifnotisinstance(item,dict):problems.append((name,"missing"))elifitem.get("candidate")!=candidate:problems.append((name,"stale"))elifitem.get("status")!="pass":problems.append((name,"not_passed"))elifnotisinstance(item.get("report"),str)ornotitem["report"].strip():problems.append((name,"no_report"))returnproblemsreports={name:{"candidate":"release-17","status":"pass","report":f"reports/{name}-17.json"}fornameinREQUIRED}assertblockers("release-17",reports)==[]assertlen(blockers("release-18",reports))==3forstatein["fail","not_run","not_applicable"]:changed=deepcopy(reports)changed["quality"]["status"]=stateassertblockers("release-17",changed)==[("quality","not_passed")]changed=deepcopy(reports)delchanged["authorization"]assertblockers("release-17",changed)==[("authorization","missing")]changed=deepcopy(reports)changed["recovery"]["report"]=""assertblockers("release-17",changed)==[("recovery","no_report")]print("Match version for release; old versions, failures, missing items, and missing reports all prevent release")
Necessary items here do not allow skipping with "not applicable". Conditional items can be judged as not applicable by release rules in clear scenarios, for example, query services with no business writes do not need refund deduplication checks, but still need their own timeout recovery checks. Candidate versions cannot reduce the list of necessary items themselves.
Which candidate do green checks cover?
Checklist status should bind to the candidate, input set and check environment, not persist as a timeless boolean. This model expands revision only, separating a past pass from currently usable evidence.
Preparing the visual
Which candidate do green checks cover?
Release gates need valid evidence for the current candidate; previous green results do not automatically carry over.
{"id":"ai-release-evidence","title":"Which candidate do green checks cover?","summary":"Release gates need valid evidence for the current candidate; previous green results do not automatically carry over.","height":1000,"html":"<h2 data-i18n=\"heading\"></h2><p class=\"intro\" data-i18n=\"intro\"></p><div class=\"actions\"><button id=\"change\" data-i18n=\"change\"></button><button id=\"quality\" data-i18n=\"quality\"></button><button id=\"authority\" data-i18n=\"authority\"></button><button id=\"recovery\" data-i18n=\"recovery\"></button><button id=\"reset\" data-i18n=\"reset\"></button></div><section class=\"panel\"><h3 data-i18n=\"state\"></h3><div id=\"state\"></div></section><div class=\"copy\"><p id=\"explanation\" role=\"status\"></p><p class=\"reserve\" aria-hidden=\"true\" data-i18n=\"explain\"></p></div><details class=\"scope\"><summary data-i18n=\"scopeLabel\"></summary><p data-i18n=\"scope\"></p></details>","css":".intro{margin:8px 0 18px;color:var(--muted);font-size:14px}.control{display:grid;grid-template-columns:minmax(0,1fr) auto;gap:6px 12px;align-items:center;margin:14px 0;font-size:13px}.control output{color:var(--accent);text-align:right;font-variant-numeric:tabular-nums;min-width:4em}.control input{grid-column:1/-1;width:100%;margin:0}.presets{display:flex;gap:4px;align-items:stretch}.presets button{flex:1;min-width:0;overflow-wrap:anywhere;font-size:13px;min-height:44px}.actions{display:grid;grid-template-columns:repeat(2,minmax(0,1fr));gap:8px;margin:16px 0}.actions button{min-width:0;min-height:44px;font-size:13px;overflow-wrap:anywhere}.copy{display:grid;font-size:14px;margin:16px 0}.copy>*{grid-area:1/1;margin:0;overflow-wrap:anywhere}.reserve{visibility:hidden}.scope{margin-top:14px;color:var(--muted);font-size:12px}.scope summary{padding:8px 0;cursor:pointer}.scope p{margin-top:10px;overflow-wrap:anywhere}.metrics{display:grid;grid-template-columns:repeat(2,minmax(0,1fr));gap:14px;margin:18px 0}.metrics>div{border-left:2px solid var(--accent);padding:0 6px 0 10px;min-width:0}.metrics span{display:block;min-height:3.6em;color:var(--muted);font-size:12px;overflow-wrap:anywhere}.metrics strong{font-size:21px;display:block;min-height:2em;font-weight:550;overflow-wrap:anywhere;font-variant-numeric:tabular-nums}.check{display:flex;align-items:center;gap:9px;font-size:13px;margin:14px 0}.check input{width:18px;height:18px;flex-shrink:0;accent-color:var(--accent)}select,input[type=text],input[type=number]{background:var(--paper);color:var(--ink);font:inherit;font-size:14px;padding:10px;border:1px solid var(--rule);border-radius:7px;max-width:100%;min-width:0}select{width:100%;min-height:44px}.panel{padding:14px;border:1px solid var(--rule);border-radius:9px;margin:16px 0;min-width:0}.panel>header{font-size:13px;font-weight:600;margin-bottom:12px;color:var(--muted)}.panel>div{overflow-wrap:anywhere}.track-row{margin:16px 0}.track-row>span{display:block;font-size:12px;color:var(--muted);margin-bottom:8px}.track{display:flex;gap:5px;min-width:0}.cell{flex:1;min-width:0;min-height:56px;border:1px solid var(--rule);border-radius:5px;background:var(--surface);display:flex;align-items:center;justify-content:center;text-align:center;padding:6px 3px;font-size:12px;overflow-wrap:anywhere}.cell.on{background:var(--accent-soft);border-color:var(--accent)}.cell.gold{background:var(--second-soft);border-color:var(--second)}.cell.empty{border-style:dashed;color:var(--muted)}.cell.error{border-color:var(--second);text-decoration:line-through}.table{display:grid;gap:5px;font-size:12px}.tr{display:grid;gap:5px}.tr>div{background:var(--surface);border-radius:4px;padding:9px 6px;min-width:0;min-height:5em;overflow-wrap:anywhere;display:flex;align-items:center}.tr>.gold{background:var(--second-soft)}.tr>.on{background:var(--accent-soft)}.formula,.source{font:13px/1.8 ui-monospace,monospace;white-space:pre-wrap;overflow-wrap:anywhere;margin:16px 0}.formula{border-top:1px solid var(--rule);padding-top:12px;min-height:7.2em}.source{min-height:8em;background:var(--surface);padding:12px;border-radius:8px}.intervals{margin:16px 0}.interval{position:relative;height:38px;background:var(--surface);margin:6px 0;border-radius:4px;overflow:hidden}.interval i{position:absolute;height:100%;background:var(--accent-soft);border-left:2px solid var(--accent)}.interval i.gold{background:var(--second-soft);border-color:var(--second)}.interval span{position:relative;z-index:1;font:12px/38px ui-monospace,monospace;padding-left:7px}.bars{display:grid;gap:10px;margin:16px 0}.bar-row{display:grid;grid-template-columns:44px minmax(0,1fr) 62px;gap:8px;align-items:center;font:12px ui-monospace,monospace}.bar-track{height:24px;background:var(--surface);border-radius:4px;overflow:hidden}.bar-track i{display:block;height:100%;background:var(--accent);transition:width .2s}.bar-track i.gold{background:var(--second)}.bar-track i.empty{opacity:.2}.bar-row output{text-align:right}.diagram{display:block;width:100%;height:auto;margin:18px 0}.diagram text{font:13px ui-monospace,monospace;fill:var(--ink)}.node{fill:var(--surface);stroke:var(--rule)}.node.on{fill:var(--accent-soft);stroke:var(--accent)}.node.gold{fill:var(--second-soft);stroke:var(--second)}.edge{stroke:var(--rule);stroke-width:2;fill:none}.edge.on{stroke:var(--accent)}.matrix{display:grid;grid-template-columns:repeat(3,minmax(0,1fr));gap:6px;margin:16px 0}.matrix button{min-width:0;min-height:44px;font:13px ui-monospace,monospace}.matrix button[aria-pressed=true]{background:var(--accent-soft);border-color:var(--accent)}.matrix button.masked{opacity:.35}.legend{font-size:12px;color:var(--muted);min-height:3.6em;overflow-wrap:anywhere}.numeric{font-variant-numeric:tabular-nums}.scope p{line-break:strict}@media(max-width:480px){.presets button,.actions button{font-size:12px}.panel{padding:12px}.metrics{gap:10px}.metrics strong{font-size:19px}.bar-row{grid-template-columns:34px minmax(0,1fr) 58px;gap:6px}}@media(prefers-reduced-motion:reduce){*,*::before,*::after{transition:none!important;animation:none!important}}\n\n.panel>h3{font-size:13px;font-weight:600;margin:0 0 12px;color:var(--muted)}\n\n.bar-row{grid-template-columns:44px minmax(0,1fr) 82px}.bar-row output{white-space:nowrap}@media(max-width:480px){.bar-row{grid-template-columns:34px minmax(0,1fr) 76px}}\n\n","js":"const $=s=>document.querySelector(s);const set=(id,v)=>$('#'+id).textContent=v;const t=k=>viz.t(k);function cells(id,values){$('#'+id).replaceChildren(...values.map(v=>{const e=document.createElement('div');e.className='cell '+(v.cls||'');e.textContent=v.text;e.title=v.title||v.text;return e;}));}function pressed(mode){document.querySelectorAll('[data-mode]').forEach(e=>e.setAttribute('aria-pressed',String(e.dataset.mode===mode)));}const esc=s=>String(s).replace(/[&<>\"']/g,c=>({'&':'&','<':'<','>':'>','\"':'"',\"'\":'''}[c]));const fmt=s=>'{'+[...s].sort().join(', ')+'}';function graph(id,nodes,edges){const el=$('#'+id);el.setAttribute('viewBox','0 0 360 200');el.innerHTML='<defs><marker id=\"arrow-'+id+'\" viewBox=\"0 0 10 10\" refX=\"8\" refY=\"5\" markerWidth=\"5\" markerHeight=\"5\" orient=\"auto-start-reverse\"><path d=\"M 0 0 L 10 5 L 0 10 z\" fill=\"var(--muted)\"/></marker></defs>'+edges.map(e=>{const a=nodes[e[0]],b=nodes[e[1]],dx=b.x-a.x,dy=b.y-a.y,d=Math.hypot(dx,dy)||1;if(edges.some(r=>r[0]===e[1]&&r[1]===e[0])){const nx=-dy/d,ny=dx/d;return '<path class=\"edge '+(e[2]?'on':'')+'\" d=\"M '+(a.x+dx/d*22+nx*12)+' '+(a.y+dy/d*22+ny*12)+' Q '+((a.x+b.x)/2+nx*38)+' '+((a.y+b.y)/2+ny*38)+' '+(b.x-dx/d*25+nx*12)+' '+(b.y-dy/d*25+ny*12)+'\" marker-end=\"url(#arrow-'+id+')\"/>';}return '<line class=\"edge '+(e[2]?'on':'')+'\" x1=\"'+(a.x+dx/d*25)+'\" y1=\"'+(a.y+dy/d*25)+'\" x2=\"'+(b.x-dx/d*28)+'\" y2=\"'+(b.y-dy/d*28)+'\" marker-end=\"url(#arrow-'+id+')\"/>';}).join('')+nodes.map(n=>'<g><circle class=\"node '+(n.cls||'')+'\" cx=\"'+n.x+'\" cy=\"'+n.y+'\" r=\"25\"/><text x=\"'+n.x+'\" y=\"'+(n.y+4)+'\" text-anchor=\"middle\">'+esc(n.name)+'</text>'+(n.sub?'<text x=\"'+n.x+'\" y=\"'+(n.y+43)+'\" text-anchor=\"middle\">'+esc(n.sub)+'</text>':'')+'</g>').join('');}function table(id,rows){$('#'+id).classList.add('table');$('#'+id).replaceChildren(...rows.map(r=>{const row=document.createElement('div');row.className='tr';row.style.gridTemplateColumns='repeat('+r.length+',minmax(0,1fr))';r.forEach(x=>{const c=document.createElement('div');if(typeof x==='object'){c.textContent=x.text;c.className=x.cls||'';}else c.textContent=x;row.append(c)});return row;}));}function band(id,parts,total){$('#'+id).replaceChildren(...parts.map((p,i)=>{const e=document.createElement('i');e.style.width=(100*p.value/total)+'%';e.className=i%2?'gold':'on';e.title=p.name+': '+p.value;return e}));}function intervals(id,items,total){$('#'+id).replaceChildren(...items.map(p=>{const row=document.createElement('div');row.className='interval';const bar=document.createElement('i');bar.style.left=(100*p.start/total)+'%';bar.style.width=(100*(p.end-p.start)/total)+'%';bar.className=p.cls||'';const label=document.createElement('span');label.textContent=p.label;row.append(bar,label);return row;}));}function plot(id,series,xr,yr,opts={}){const e=$('#'+id),X=x=>42+(x-xr[0])/(xr[1]-xr[0])*298,Y=y=>175-(y-yr[0])/(yr[1]-yr[0])*148;let out='<defs><clipPath id=\"clip-'+id+'\"><rect x=\"42\" y=\"27\" width=\"298\" height=\"148\"/></clipPath></defs>';for(let i=0;i<3;i++){let x=xr[0]+i*(xr[1]-xr[0])/2,y=yr[0]+i*(yr[1]-yr[0])/2;out+='<path class=\"gridline\" d=\"M '+X(x)+' 27V175M42 '+Y(y)+'H340\"/><text x=\"'+X(x)+'\" y=\"194\" text-anchor=\"middle\">'+esc(opts.xfmt?opts.xfmt(x):Number(x.toFixed(2)))+'</text><text x=\"36\" y=\"'+(Y(y)+4)+'\" text-anchor=\"end\">'+esc(opts.yfmt?opts.yfmt(y):Number(y.toFixed(2)))+'</text>';}out+='<g clip-path=\"url(#clip-'+id+')\">';for(const s of series){out+='<path fill=\"none\" stroke=\"'+(s.color||'var(--accent)')+'\" stroke-width=\"2.5\" '+(s.dash?'stroke-dasharray=\"5 4\"':'')+' d=\"'+s.pts.map((p,i)=>(i?'L':'M')+X(p[0]).toFixed(2)+','+Y(p[1]).toFixed(2)).join(' ')+'\"/>';}for(const p of opts.points||[])out+='<circle cx=\"'+X(p[0])+'\" cy=\"'+Y(p[1])+'\" r=\"4\" fill=\"var(--second)\" stroke=\"var(--paper)\" stroke-width=\"1.5\"/>';if(opts.cursor!==undefined)out+='<path d=\"M'+X(opts.cursor)+' 27V175\" stroke=\"var(--muted)\" stroke-dasharray=\"3 3\"/>';e.innerHTML=out+'</g>';e.dataset.curves=JSON.stringify(series.map(s=>s.pts));}const curve=(fn,lo,hi,n=120)=>Array.from({length:n+1},(_,i)=>{const x=lo+(hi-lo)*i/n;return[x,fn(x)]});function inputs(draw){document.querySelectorAll('input,select').forEach(e=>{e.addEventListener('input',draw);e.addEventListener('change',draw)});draw()}const n=id=>+$('#'+id).value;const show=(id,v,d=3)=>set(id,Number(v.toFixed(d)));function bars(id,values,max=1){$('#'+id).className='bars';$('#'+id).innerHTML=values.map(v=>'<div class=\"bar-row\"><span>'+esc(v.label)+'</span><div class=\"bar-track\"><i class=\"'+(v.cls||'')+'\" style=\"width:'+Math.max(0,Math.min(100,v.value/max*100))+'%\"></i></div><output>'+esc(v.text??(v.value*100).toFixed(1)+'%')+'</output></div>').join('')}function softmax(z,T=1){let max=Math.max(...z),w=z.map(x=>Math.exp((x-max)/T)),sum=w.reduce((a,b)=>a+b,0);return w.map(x=>x/sum)}let v=17,reports={quality:17,authority:17,recovery:17};function draw(){let ok=Object.values(reports).every(r=>r===v);table('state',[[t('candidate'),'r'+v],...Object.entries(reports).map(([k,r])=>[t(k),'r'+r+' ✓',r===v?'✓':'≠']),[t('gate'),ok?'✓':'×']]);set('explanation',t('explain'));document.body.dataset.accepted=ok;document.body.dataset.version=v;}$('#change').onclick=()=>{v=v===17?18:17;draw()};for(const key of Object.keys(reports))$('#'+key).onclick=()=>{reports[key]=v;draw()};$('#reset').onclick=()=>{v=17;reports={quality:17,authority:17,recovery:17};draw()};draw();","audio":false,"strings":{"authority":"Validate authority","candidate":"Candidate revision","change":"Switch candidate r17 / r18","explain":"Reports may still say pass, but r17 reports cannot open the r18 gate. Reusing old-candidate evidence also requires validity conditions such as expiry; this model checks revision only.","gate":"Example release gate","heading":"Which candidate do green checks cover?","intro":"Update the candidate, then produce acceptance evidence for its current version.","quality":"Validate quality","recovery":"Validate recovery","reset":"Reset","scope":"Three local example checks: quality, authority and recovery. Buttons model passing reports; real releases need full risk coverage. No deployment occurs.","scopeLabel":"Model scope and assumptions","state":"Candidate and evidence"}}
Canarying and Recovery
After passing offline acceptance, observe the new version with limited real tasks and compare with comparable old version traffic. Canarying requires pre-defined metrics, observation conditions, and stop rules; setting traffic ratio to 1% does not automatically yield sufficient evidence. Low-traffic businesses may go a long time without encountering important failure scenarios. Google SRE: Canarying Releases
Stage
Key Confirmation
Basis for Continuing or Stopping
Shadow Verification (when applicable)
Output and resource consumption under same input
Write tools should be isolated or simulated, not repeat real side effects
Meet preset observation conditions and no stop rules triggered
Gradual Expansion
Quotas, queuing, cache, and downstream capacity
Still within budget and service goals after scale growth
Rollback to Stable Version
New task routing, in-flight tasks, and state compatibility
Confirm recovery results and handle already occurred side effects separately
For example, 95 successes out of 100 tasks does not mean the real success rate is proven to be at least 95%. Under independent, identically distributed simplifying assumptions, the two-sided 95% Wilson interval lower bound for 95/100 is approximately 88.8%; real tasks also have distribution changes and correlations. Whether to expand should be combined with sample size, key groups, and business tolerance; do not use a single point estimate to replace judgment.
The rollback plan needs to answer three different questions: how new requests return to old configurations, how in-flight tasks stop or complete, and how to verify effects like refunds/sending emails that have already occurred. Rolling back the prompt does not revoke refunds; forcibly switching model or tool versions for in-progress tasks may also break their state contracts. Prioritize binding tasks to candidate versions, and clearly define upgrade, drain, and recovery strategies.
How strong is a 95% success rate?
Release evidence also needs enough sample information. More samples narrow the interval at a fixed observed proportion, but repeated similar tasks do not establish independence or replace group coverage. The interval is not a universal release threshold. NIST: Wilson interval
1+nz2p^+2nz2±znp^(1−p^)+4n2z2
Preparing the visual
How strong is a 95% success rate?
The same observed rate has a wider interval with fewer samples; zero failures do not prove zero failure probability.
{"id":"ai-wilson-interval","title":"How strong is a 95% success rate?","summary":"The same observed rate has a wider interval with fewer samples; zero failures do not prove zero failure probability.","height":1000,"html":"<h2 data-i18n=\"heading\"></h2><p class=\"intro\" data-i18n=\"intro\"></p><div class=\"control\"><label for=\"rate\" data-i18n=\"rate\"></label><output for=\"rate\" id=\"rate-value\"></output><input id=\"rate\" type=\"range\" min=\"0\" max=\"100\" value=\"95\" step=\"5\"></div><div class=\"presets\" role=\"group\" data-i18n-label=\"size\"><button data-mode=\"small\" data-i18n=\"small\"></button><button data-mode=\"medium\" data-i18n=\"medium\"></button><button data-mode=\"large\" data-i18n=\"large\"></button></div><section class=\"panel\"><h3 data-i18n=\"interval\"></h3><div id=\"band\"></div></section><div class=\"metrics\"><div><span data-i18n=\"interval\"></span><strong id=\"interval\"></strong></div><div><span data-i18n=\"count\"></span><strong id=\"count\"></strong></div></div><p class=\"formula\" id=\"math\"></p><div class=\"copy\"><p id=\"explanation\" role=\"status\"></p><p class=\"reserve\" aria-hidden=\"true\" data-i18n=\"explain\"></p></div><details class=\"scope\"><summary data-i18n=\"scopeLabel\"></summary><p data-i18n=\"scope\"></p></details>","css":".intro{margin:8px 0 18px;color:var(--muted);font-size:14px}.control{display:grid;grid-template-columns:minmax(0,1fr) auto;gap:6px 12px;align-items:center;margin:14px 0;font-size:13px}.control output{color:var(--accent);text-align:right;font-variant-numeric:tabular-nums;min-width:4em}.control input{grid-column:1/-1;width:100%;margin:0}.presets{display:flex;gap:4px;align-items:stretch}.presets button{flex:1;min-width:0;overflow-wrap:anywhere;font-size:13px;min-height:44px}.actions{display:grid;grid-template-columns:repeat(2,minmax(0,1fr));gap:8px;margin:16px 0}.actions button{min-width:0;min-height:44px;font-size:13px;overflow-wrap:anywhere}.copy{display:grid;font-size:14px;margin:16px 0}.copy>*{grid-area:1/1;margin:0;overflow-wrap:anywhere}.reserve{visibility:hidden}.scope{margin-top:14px;color:var(--muted);font-size:12px}.scope summary{padding:8px 0;cursor:pointer}.scope p{margin-top:10px;overflow-wrap:anywhere}.metrics{display:grid;grid-template-columns:repeat(2,minmax(0,1fr));gap:14px;margin:18px 0}.metrics>div{border-left:2px solid var(--accent);padding:0 6px 0 10px;min-width:0}.metrics span{display:block;min-height:3.6em;color:var(--muted);font-size:12px;overflow-wrap:anywhere}.metrics strong{font-size:21px;display:block;min-height:2em;font-weight:550;overflow-wrap:anywhere;font-variant-numeric:tabular-nums}.check{display:flex;align-items:center;gap:9px;font-size:13px;margin:14px 0}.check input{width:18px;height:18px;flex-shrink:0;accent-color:var(--accent)}select,input[type=text],input[type=number]{background:var(--paper);color:var(--ink);font:inherit;font-size:14px;padding:10px;border:1px solid var(--rule);border-radius:7px;max-width:100%;min-width:0}select{width:100%;min-height:44px}.panel{padding:14px;border:1px solid var(--rule);border-radius:9px;margin:16px 0;min-width:0}.panel>header{font-size:13px;font-weight:600;margin-bottom:12px;color:var(--muted)}.panel>div{overflow-wrap:anywhere}.track-row{margin:16px 0}.track-row>span{display:block;font-size:12px;color:var(--muted);margin-bottom:8px}.track{display:flex;gap:5px;min-width:0}.cell{flex:1;min-width:0;min-height:56px;border:1px solid var(--rule);border-radius:5px;background:var(--surface);display:flex;align-items:center;justify-content:center;text-align:center;padding:6px 3px;font-size:12px;overflow-wrap:anywhere}.cell.on{background:var(--accent-soft);border-color:var(--accent)}.cell.gold{background:var(--second-soft);border-color:var(--second)}.cell.empty{border-style:dashed;color:var(--muted)}.cell.error{border-color:var(--second);text-decoration:line-through}.table{display:grid;gap:5px;font-size:12px}.tr{display:grid;gap:5px}.tr>div{background:var(--surface);border-radius:4px;padding:9px 6px;min-width:0;min-height:5em;overflow-wrap:anywhere;display:flex;align-items:center}.tr>.gold{background:var(--second-soft)}.tr>.on{background:var(--accent-soft)}.formula,.source{font:13px/1.8 ui-monospace,monospace;white-space:pre-wrap;overflow-wrap:anywhere;margin:16px 0}.formula{border-top:1px solid var(--rule);padding-top:12px;min-height:7.2em}.source{min-height:8em;background:var(--surface);padding:12px;border-radius:8px}.intervals{margin:16px 0}.interval{position:relative;height:38px;background:var(--surface);margin:6px 0;border-radius:4px;overflow:hidden}.interval i{position:absolute;height:100%;background:var(--accent-soft);border-left:2px solid var(--accent)}.interval i.gold{background:var(--second-soft);border-color:var(--second)}.interval span{position:relative;z-index:1;font:12px/38px ui-monospace,monospace;padding-left:7px}.bars{display:grid;gap:10px;margin:16px 0}.bar-row{display:grid;grid-template-columns:44px minmax(0,1fr) 62px;gap:8px;align-items:center;font:12px ui-monospace,monospace}.bar-track{height:24px;background:var(--surface);border-radius:4px;overflow:hidden}.bar-track i{display:block;height:100%;background:var(--accent);transition:width .2s}.bar-track i.gold{background:var(--second)}.bar-track i.empty{opacity:.2}.bar-row output{text-align:right}.diagram{display:block;width:100%;height:auto;margin:18px 0}.diagram text{font:13px ui-monospace,monospace;fill:var(--ink)}.node{fill:var(--surface);stroke:var(--rule)}.node.on{fill:var(--accent-soft);stroke:var(--accent)}.node.gold{fill:var(--second-soft);stroke:var(--second)}.edge{stroke:var(--rule);stroke-width:2;fill:none}.edge.on{stroke:var(--accent)}.matrix{display:grid;grid-template-columns:repeat(3,minmax(0,1fr));gap:6px;margin:16px 0}.matrix button{min-width:0;min-height:44px;font:13px ui-monospace,monospace}.matrix button[aria-pressed=true]{background:var(--accent-soft);border-color:var(--accent)}.matrix button.masked{opacity:.35}.legend{font-size:12px;color:var(--muted);min-height:3.6em;overflow-wrap:anywhere}.numeric{font-variant-numeric:tabular-nums}.scope p{line-break:strict}@media(max-width:480px){.presets button,.actions button{font-size:12px}.panel{padding:12px}.metrics{gap:10px}.metrics strong{font-size:19px}.bar-row{grid-template-columns:34px minmax(0,1fr) 58px;gap:6px}}@media(prefers-reduced-motion:reduce){*,*::before,*::after{transition:none!important;animation:none!important}}\n\n.panel>h3{font-size:13px;font-weight:600;margin:0 0 12px;color:var(--muted)}\n\n.bar-row{grid-template-columns:44px minmax(0,1fr) 82px}.bar-row output{white-space:nowrap}@media(max-width:480px){.bar-row{grid-template-columns:34px minmax(0,1fr) 76px}}\n\n.metrics strong{font-size:16px;min-height:4em}","js":"const $=s=>document.querySelector(s);const set=(id,v)=>$('#'+id).textContent=v;const t=k=>viz.t(k);function cells(id,values){$('#'+id).replaceChildren(...values.map(v=>{const e=document.createElement('div');e.className='cell '+(v.cls||'');e.textContent=v.text;e.title=v.title||v.text;return e;}));}function pressed(mode){document.querySelectorAll('[data-mode]').forEach(e=>e.setAttribute('aria-pressed',String(e.dataset.mode===mode)));}const esc=s=>String(s).replace(/[&<>\"']/g,c=>({'&':'&','<':'<','>':'>','\"':'"',\"'\":'''}[c]));const fmt=s=>'{'+[...s].sort().join(', ')+'}';function graph(id,nodes,edges){const el=$('#'+id);el.setAttribute('viewBox','0 0 360 200');el.innerHTML='<defs><marker id=\"arrow-'+id+'\" viewBox=\"0 0 10 10\" refX=\"8\" refY=\"5\" markerWidth=\"5\" markerHeight=\"5\" orient=\"auto-start-reverse\"><path d=\"M 0 0 L 10 5 L 0 10 z\" fill=\"var(--muted)\"/></marker></defs>'+edges.map(e=>{const a=nodes[e[0]],b=nodes[e[1]],dx=b.x-a.x,dy=b.y-a.y,d=Math.hypot(dx,dy)||1;if(edges.some(r=>r[0]===e[1]&&r[1]===e[0])){const nx=-dy/d,ny=dx/d;return '<path class=\"edge '+(e[2]?'on':'')+'\" d=\"M '+(a.x+dx/d*22+nx*12)+' '+(a.y+dy/d*22+ny*12)+' Q '+((a.x+b.x)/2+nx*38)+' '+((a.y+b.y)/2+ny*38)+' '+(b.x-dx/d*25+nx*12)+' '+(b.y-dy/d*25+ny*12)+'\" marker-end=\"url(#arrow-'+id+')\"/>';}return '<line class=\"edge '+(e[2]?'on':'')+'\" x1=\"'+(a.x+dx/d*25)+'\" y1=\"'+(a.y+dy/d*25)+'\" x2=\"'+(b.x-dx/d*28)+'\" y2=\"'+(b.y-dy/d*28)+'\" marker-end=\"url(#arrow-'+id+')\"/>';}).join('')+nodes.map(n=>'<g><circle class=\"node '+(n.cls||'')+'\" cx=\"'+n.x+'\" cy=\"'+n.y+'\" r=\"25\"/><text x=\"'+n.x+'\" y=\"'+(n.y+4)+'\" text-anchor=\"middle\">'+esc(n.name)+'</text>'+(n.sub?'<text x=\"'+n.x+'\" y=\"'+(n.y+43)+'\" text-anchor=\"middle\">'+esc(n.sub)+'</text>':'')+'</g>').join('');}function table(id,rows){$('#'+id).classList.add('table');$('#'+id).replaceChildren(...rows.map(r=>{const row=document.createElement('div');row.className='tr';row.style.gridTemplateColumns='repeat('+r.length+',minmax(0,1fr))';r.forEach(x=>{const c=document.createElement('div');if(typeof x==='object'){c.textContent=x.text;c.className=x.cls||'';}else c.textContent=x;row.append(c)});return row;}));}function band(id,parts,total){$('#'+id).replaceChildren(...parts.map((p,i)=>{const e=document.createElement('i');e.style.width=(100*p.value/total)+'%';e.className=i%2?'gold':'on';e.title=p.name+': '+p.value;return e}));}function intervals(id,items,total){$('#'+id).replaceChildren(...items.map(p=>{const row=document.createElement('div');row.className='interval';const bar=document.createElement('i');bar.style.left=(100*p.start/total)+'%';bar.style.width=(100*(p.end-p.start)/total)+'%';bar.className=p.cls||'';const label=document.createElement('span');label.textContent=p.label;row.append(bar,label);return row;}));}function plot(id,series,xr,yr,opts={}){const e=$('#'+id),X=x=>42+(x-xr[0])/(xr[1]-xr[0])*298,Y=y=>175-(y-yr[0])/(yr[1]-yr[0])*148;let out='<defs><clipPath id=\"clip-'+id+'\"><rect x=\"42\" y=\"27\" width=\"298\" height=\"148\"/></clipPath></defs>';for(let i=0;i<3;i++){let x=xr[0]+i*(xr[1]-xr[0])/2,y=yr[0]+i*(yr[1]-yr[0])/2;out+='<path class=\"gridline\" d=\"M '+X(x)+' 27V175M42 '+Y(y)+'H340\"/><text x=\"'+X(x)+'\" y=\"194\" text-anchor=\"middle\">'+esc(opts.xfmt?opts.xfmt(x):Number(x.toFixed(2)))+'</text><text x=\"36\" y=\"'+(Y(y)+4)+'\" text-anchor=\"end\">'+esc(opts.yfmt?opts.yfmt(y):Number(y.toFixed(2)))+'</text>';}out+='<g clip-path=\"url(#clip-'+id+')\">';for(const s of series){out+='<path fill=\"none\" stroke=\"'+(s.color||'var(--accent)')+'\" stroke-width=\"2.5\" '+(s.dash?'stroke-dasharray=\"5 4\"':'')+' d=\"'+s.pts.map((p,i)=>(i?'L':'M')+X(p[0]).toFixed(2)+','+Y(p[1]).toFixed(2)).join(' ')+'\"/>';}for(const p of opts.points||[])out+='<circle cx=\"'+X(p[0])+'\" cy=\"'+Y(p[1])+'\" r=\"4\" fill=\"var(--second)\" stroke=\"var(--paper)\" stroke-width=\"1.5\"/>';if(opts.cursor!==undefined)out+='<path d=\"M'+X(opts.cursor)+' 27V175\" stroke=\"var(--muted)\" stroke-dasharray=\"3 3\"/>';e.innerHTML=out+'</g>';e.dataset.curves=JSON.stringify(series.map(s=>s.pts));}const curve=(fn,lo,hi,n=120)=>Array.from({length:n+1},(_,i)=>{const x=lo+(hi-lo)*i/n;return[x,fn(x)]});function inputs(draw){document.querySelectorAll('input,select').forEach(e=>{e.addEventListener('input',draw);e.addEventListener('change',draw)});draw()}const n=id=>+$('#'+id).value;const show=(id,v,d=3)=>set(id,Number(v.toFixed(d)));function bars(id,values,max=1){$('#'+id).className='bars';$('#'+id).innerHTML=values.map(v=>'<div class=\"bar-row\"><span>'+esc(v.label)+'</span><div class=\"bar-track\"><i class=\"'+(v.cls||'')+'\" style=\"width:'+Math.max(0,Math.min(100,v.value/max*100))+'%\"></i></div><output>'+esc(v.text??(v.value*100).toFixed(1)+'%')+'</output></div>').join('')}function softmax(z,T=1){let max=Math.max(...z),w=z.map(x=>Math.exp((x-max)/T)),sum=w.reduce((a,b)=>a+b,0);return w.map(x=>x/sum)}let mode='medium';function draw(){pressed(mode);let N={small:20,medium:100,large:1000}[mode],k=Math.round(n('rate')/100*N),p=k/N,z=1.96,d=1+z*z/N,c=(p+z*z/(2*N))/d,h=z*Math.sqrt(p*(1-p)/N+z*z/(4*N*N))/d,lo=Math.max(0,c-h),hi=Math.min(1,c+h);set('rate-value',n('rate')+'%');intervals('band',[{label:'0% → 100%',start:0,end:1},{label:(lo*100).toFixed(1)+'% — '+(hi*100).toFixed(1)+'%',start:lo,end:hi,cls:'gold'}],1);set('interval',(lo*100).toFixed(1)+'–'+(hi*100).toFixed(1)+'%');set('count',k+' / '+N);set('math','p̂ = '+p.toFixed(2)+'; n = '+N+'\\ncenter = '+c.toFixed(4)+'\\nhalf-width = '+h.toFixed(4));set('explanation',t('explain'));document.body.dataset.interval=JSON.stringify([lo,hi]);}document.querySelectorAll('[data-mode]').forEach(e=>e.onclick=()=>{mode=e.dataset.mode;draw()});inputs(draw);","audio":false,"strings":{"count":"Successes / trials","explain":"For 95/100, the interval is about 88.8%–97.8%. It is not a 95% probability statement about the fixed parameter in this computed interval; 95% refers to the method’s repeated-sampling coverage target.","heading":"How strong is a 95% success rate?","interval":"95% Wilson interval","intro":"Keep the success fraction and vary sample size; inspect the Wilson interval.","large":"1000","medium":"100","rate":"Observed success (%)","reset":"Reset","scope":"IID Bernoulli trials; approximate two-sided 95% Wilson interval with z=1.96. Correlated tasks, selection bias and repeated looks require separate treatment.","scopeLabel":"Model scope and assumptions","size":"Sample size","small":"20"}}
Post-release Closed Loop
Before entering regression for an incident, first save versions and necessary traces sufficient to explain the failure, and appropriately handle personal information and sensitive data. Directly copying real data into long-term evaluation sets may expand exposure scope; synthetic cases preserving the same failure conditions can be constructed.
A useful regression is not just "asking the same question again," but also recording expected results and failure locations. For example, cross-tenant ticket incidents can be split into executor rejection tests and end-to-end injection tests: the former verifies permission boundaries, the latter observes whether the model continues to correctly complete the original task. Run directly related tests after modification, then supplement regressions based on dependency impacts, and re-confirm necessary evidence belongs to the current combination before release.
Checklists will be adjusted as the system changes. Delete invalid, duplicate checks and retain reasons; add checks that cover real gaps; the maintenance goal is to make evidence better predict production behavior, rather than letting the number of checkboxes increase endlessly.