---
title: Release validation, canaries and recovery
url: https://doc.liz6.com/en/ai/04-evaluation-and-production/03-release-and-recovery
locale: en
area: ai
tags:
- Models & agents
- Evaluation and Production
date: 2026-06-30
modified: 2026-09-10
description: A release combines code, model, prompts, tools, policy and indexes. Assemble evidence for that exact combination before choosing a canary, wider rollout or stop. This article integrates validation; detailed formulas and execution rules remain in their dedicated articles.
---

# Release validation, canaries and recovery

A release combines code, model, prompts, tools, policy and indexes. Assemble evidence for that exact combination before choosing a canary, wider rollout or stop. This article integrates validation; detailed formulas and execution rules remain in their dedicated articles.

## Fixed Candidate Versions and Task Boundaries

First, clearly define what the system accomplishes and how completion is confirmed. Query-based assistants may require answers with verifiable sources; execution-based assistants also need to check actual business states. A completed model response, HTTP 200, or tools not throwing exceptions cannot independently prove task success. Manual processing is also a distinct result branch; its proportion and cost should be tracked and cannot be silently counted as automatic success.

Candidate versions must at least allow locating application code, model identifiers and parameters, prompt templates, tool schemas, and permission policies. When using retrieval, document document snapshots, chunking and embedding configurations, and index versions; when using persistent state, document state formats and migration rules. Do not just assign a version number to the prompt while letting other dependencies change silently after verification. For services that only support floating model aliases, record the actually obtainable version information and verification time, and acknowledge that changes by the provider reduce reproducibility.

<svg viewBox="0 0 760 390.3666687011719" xmlns="http://www.w3.org/2000/svg" role="img" aria-label="What goes live is a set of versions; the basis for release must belong to this set of versions" style="max-width:100%;height:auto" font-family="Source Han Sans CN,Microsoft YaHei,sans-serif">
<defs><marker id="ai-release-evidence-arrow" markerWidth="8" markerHeight="8" refX="7" refY="4" orient="auto"><path d="M0,0 L8,4 L0,8 Z" fill="#64748b"></path></marker></defs>
<rect width="760" height="390.3666687011719" rx="12" fill="#f8fafc"></rect>


<g transform="translate(0 0)"><rect x="30" y="85" width="190" height="80" rx="8" fill="#e0e7ff" stroke="#c7d2fe"></rect><text x="125.0" y="122.0" font-size="15" fill="#1e293b" text-anchor="middle" font-weight="600">Candidate Versions</text><text x="125.0" y="144.0" font-size="12" fill="#475569" text-anchor="middle" font-weight="400">Models, Prompts, Tools, Indexes</text><rect x="285" y="85" width="190" height="80" rx="8" fill="#e0e7ff" stroke="#c7d2fe"></rect><text x="380.0" y="122.0" font-size="15" fill="#1e293b" text-anchor="middle" font-weight="600">Validation Evidence</text><text x="380.0" y="144.0" font-size="12" fill="#475569" text-anchor="middle" font-weight="400">Quality, authority, recovery, load</text><rect x="540" y="85" width="190" height="80" rx="8" fill="#e0e7ff" stroke="#c7d2fe"></rect><text x="635.0" y="122.0" font-size="15" fill="#1e293b" text-anchor="middle" font-weight="600">Limited Rollout</text><text x="635.0" y="144.0" font-size="12" fill="#475569" text-anchor="middle" font-weight="400">Observe or Restore Version</text><line x1="220" y1="125" x2="280" y2="125" stroke="#64748b" stroke-width="1.8" marker-end="url(#ai-release-evidence-arrow)"></line><line x1="475" y1="125" x2="535" y2="125" stroke="#64748b" stroke-width="1.8" marker-end="url(#ai-release-evidence-arrow)"></line><rect x="180" y="225" width="400" height="50" rx="8" fill="#fef3c7" stroke="#c7d2fe"></rect><line x1="635" y1="170" x2="635" y2="250" stroke="#64748b" stroke-width="1.8"></line><line x1="635" y1="250" x2="585" y2="250" stroke="#64748b" stroke-width="1.8" marker-end="url(#ai-release-evidence-arrow)"></line><line x1="175" y1="250" x2="125" y2="250" stroke="#64748b" stroke-width="1.8"></line><line x1="125" y1="250" x2="125" y2="170" stroke="#64748b" stroke-width="1.8" marker-end="url(#ai-release-evidence-arrow)"></line></g><text x="24" y="29" font-size="19" fill="#0f172a" text-anchor="start" font-weight="700"><tspan x="24" dy="0">What goes live is a set of versions; the basis for release must belong to this </tspan><tspan x="24" dy="25.650000000000002">set of versions</tspan></text><text x="24" y="310" font-size="15" fill="#1e293b" text-anchor="start" font-weight="600"><tspan x="24" dy="0">Configuration changes → Recheck affected evidence</tspan></text><text x="24" y="347.20897102355957" font-size="13" fill="#475569" text-anchor="start" font-weight="400"><tspan x="24" dy="0">Passing offline does not mean no online risks; rolling back versions cannot revoke business actions that have already </tspan><tspan x="24" dy="17.55">occurred.</tspan></text>
</svg>

Related explanations: [Agent Loop](/ai/03-agent-systems/01-agent-loop.md), [Multi-Agent Orchestration](/ai/03-agent-systems/06-multi-agent-orchestration.md).

## Pre-release Checklist

Record the responsible person, candidate version, verification time, evidence location, and conclusion for each item. Record not executed and failed items separately, but neither can masquerade as passed; for not applicable items, state the reason, and do not simply delete the checkboxes.

### Quality and Model Configuration

| Check Item | Evidence to Retain | Signals That Cannot Replace It |
|---|---|---|
| [ ] Define task success and failure | Output requirements, business post-conditions, rejection/manual transfer rules | Model self-statement "completed" |
| [ ] Cover representative tasks and boundaries | Fixed evaluation set, grouped results by difficulty, language, tool type, etc. | Single demonstration or overall average score |
| [ ] Retain independent verification set | Split between tuning set and verification set, repeated trial strategy | Repeatedly tuning to full score on the same batch of questions |
| [ ] Calibrate scoring methods | Program check coverage, manual review, judge misjudgment examples | "Used program scoring so it must be objective" |
| [ ] Verify model capabilities and parameters | Parameters actually supported by the selected endpoint, verification of stop/reject/truncation branches | Setting high effort uniformly for all models |
| [ ] Verify structure and business constraints | Schema validation and negative examples for quantity, status, resource ownership, etc. | JSON can be parsed |

The "what to observe" and "who scores" in evaluation are two dimensions: business results can be checked by programs or evaluated by humans or model judges; do not treat outcome-based evaluation as a lower-level method placed below program scoring. One success does not prove stable repeated execution; the number of trials per task and aggregation methods should be stated. [Anthropic: Agent Evaluation](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents)

Sampling, inference budget, and output limits serve different purposes. Do not mechanically translate temperature into effort, nor assume all new models have removed sampling parameters. Related principles can be found in [Tokens and Sampling](/ai/01-model-foundations/01-tokens-and-sampling.md), [Reasoning and Thinking](/ai/01-model-foundations/04-reasoning-and-verification.md), [Evaluation and Observability](/ai/04-evaluation-and-production/01-evaluation-and-observability.md), [Structured Output](/ai/02-context-and-interfaces/01-prompt-and-output-contracts.md).

### Context, Knowledge, and State

| Check Item | Evidence to Retain | Typical Failure Branches |
|---|---|---|
| [ ] Reserve budget for complete requests | Total budget for tool definitions, retrieval materials, history, and output reserves | Single document fits, but complete request exceeds limit |
| [ ] Verify retrieval and answer layers | Candidate recall, ranking, citation correspondence, and behavior when no evidence exists | Answer is fluent but citations do not support the conclusion |
| [ ] Check version and permission changes | Index and cache tests after document updates, deletions, and permission revocations | Original text revoked, but old snippets still returned |
| [ ] Verify session recovery | Recovery records for completed steps, pending results, and external object IDs | After recovery, unknown results are treated as unexecuted |
| [ ] Verify state compatibility | Old state reading, new state writing, and rollback compatibility scope | Rolled-back application cannot read new checkpoints |

Taking "prohibit cross-tenant reading" as an example, do not just check if the model finally outputs someone else's content; also confirm that retrieval and tool execution phases did not bring this content into the context. Caches and long-term memory are also access paths. Cache hits do not prove the answer used the latest data; cache hit rate targets need to be judged together with freshness requirements.

See [Context Engineering](/ai/02-context-and-interfaces/02-context-engineering.md), [RAG](/ai/02-context-and-interfaces/03-retrieval-and-evidence.md), [Memory and State](/ai/03-agent-systems/04-memory-and-recovery.md). Weights, KV, and hardware budgets in model selection can be found in [Model Architecture](/ai/01-model-foundations/03-inference-and-kv-cache.md).

### Execution Permissions, Reliability, and Observability

| Check Item | Evidence to Retain | Typical Failure Branches |
|---|---|---|
| [ ] Executor independently verifies authorization | Rejection tests for cross-user/project, forged subjects, and extra parameters | Model passes `approved` and is allowed through |
| [ ] Limit general tool capabilities | Isolation verification for file boundaries, command parameters, outbound network, etc. | Command name is legal but can write files arbitrarily |
| [ ] Check injection and sensitive data flow | Synthetic injection cases, log masking, output-side processing | Tool materials elevated to trusted instructions |
| [ ] Distinguish timeouts from business failures | Queries for unknown write results, deduplication, and recovery records | Re-executing actions that already succeeded after timeout |
| [ ] Constrain retries and queuing | Total deadline, max attempts, bounded queues, and overload strategies | SDK and outer-layer retries multiply backend calls |
| [ ] Observe complete tasks | Correlated IDs, key step durations, task results, and total costs | Only counting model response time and single token usage |
| [ ] Verify failure branches | Rate limiting, tool exceptions, stream breaks, truncation, partial batch failures | Stream request establishment success counts as completion |

The names and statistical methods of observability fields vary by service. Based on interface documentation, clarify which usage fields include or exclude cache, inference, and other usage; do not simply add all numbers together. Cache reads being zero for a long time may come from requests not meeting conditions, traffic characteristics, or configuration issues; do not judge the cause of failure based solely on this value.

Existing task authorizations should be passed by the system along the operation chain; actions requiring additional confirmation should be handled according to specific policies; successive manual clicks do not replace resource permissions and idempotency mechanisms. See [Security and Protection](/ai/03-agent-systems/05-authority-and-tool-boundaries.md), [Cost, Performance, and Reliability](/ai/04-evaluation-and-production/02-cost-performance-and-reliability.md). Boundaries for access protocols and tool discovery can be found in [MCP and Skills](/ai/03-agent-systems/02-tools-and-mcp.md).

## Evidence Status and Version Matching

Four states can be used: `pass` (meeting agreed conditions), `fail` (not meeting), `not_run` (not executed), `not_applicable` (reasonably and confirmed by responsible process as not applicable). The existence of a report does not mean it passed, and a passed report does not mean it belongs to the current candidate. For necessary checks, missing, failed, not executed, or version mismatch should all prevent automatic release.

The local teaching example below only demonstrates this rule and will not deploy anything. Abstract combinations of multiple actual dependency versions into `candidate`; real systems also need to verify evidence source, content, timeliness, and scope of application, and cannot trust arbitrary callers' self-reported pass.

```python
from copy import deepcopy

REQUIRED = {"quality", "authorization", "recovery"}

def blockers(candidate, evidence):
    problems = []
    for name in sorted(REQUIRED):
        item = evidence.get(name)
        if not isinstance(item, dict):
            problems.append((name, "missing"))
        elif item.get("candidate") != candidate:
            problems.append((name, "stale"))
        elif item.get("status") != "pass":
            problems.append((name, "not_passed"))
        elif not isinstance(item.get("report"), str) or not item["report"].strip():
            problems.append((name, "no_report"))
    return problems

reports = {name: {"candidate": "release-17", "status": "pass",
                  "report": f"reports/{name}-17.json"} for name in REQUIRED}
assert blockers("release-17", reports) == []
assert len(blockers("release-18", reports)) == 3
for state in ["fail", "not_run", "not_applicable"]:
    changed = deepcopy(reports)
    changed["quality"]["status"] = state
    assert blockers("release-17", changed) == [("quality", "not_passed")]
changed = deepcopy(reports)
del changed["authorization"]
assert blockers("release-17", changed) == [("authorization", "missing")]
changed = deepcopy(reports)
changed["recovery"]["report"] = ""
assert blockers("release-17", changed) == [("recovery", "no_report")]
print("Match version for release; old versions, failures, missing items, and missing reports all prevent release")
```

Necessary items here do not allow skipping with "not applicable". Conditional items can be judged as not applicable by release rules in clear scenarios, for example, query services with no business writes do not need refund deduplication checks, but still need their own timeout recovery checks. Candidate versions cannot reduce the list of necessary items themselves.

### Which candidate do green checks cover?

Checklist status should bind to the candidate, input set and check environment, not persist as a timeless boolean. This model expands revision only, separating a past pass from currently usable evidence.

**Which candidate do green checks cover?**

Release gates need valid evidence for the current candidate; previous green results do not automatically carry over.


## Canarying and Recovery

After passing offline acceptance, observe the new version with limited real tasks and compare with comparable old version traffic. Canarying requires pre-defined metrics, observation conditions, and stop rules; setting traffic ratio to 1% does not automatically yield sufficient evidence. Low-traffic businesses may go a long time without encountering important failure scenarios. [Google SRE: Canarying Releases](https://sre.google/workbook/canarying-releases/)

| Stage | Key Confirmation | Basis for Continuing or Stopping |
|---|---|---|
| Shadow Verification (when applicable) | Output and resource consumption under same input | Write tools should be isolated or simulated, not repeat real side effects |
| Small-scale Canary | Quality groups, manual takeover, task costs, tail latency | Meet preset observation conditions and no stop rules triggered |
| Gradual Expansion | Quotas, queuing, cache, and downstream capacity | Still within budget and service goals after scale growth |
| Rollback to Stable Version | New task routing, in-flight tasks, and state compatibility | Confirm recovery results and handle already occurred side effects separately |

For example, 95 successes out of 100 tasks does not mean the real success rate is proven to be at least 95%. Under independent, identically distributed simplifying assumptions, the two-sided 95% Wilson interval lower bound for 95/100 is approximately 88.8%; real tasks also have distribution changes and correlations. Whether to expand should be combined with sample size, key groups, and business tolerance; do not use a single point estimate to replace judgment.

The rollback plan needs to answer three different questions: how new requests return to old configurations, how in-flight tasks stop or complete, and how to verify effects like refunds/sending emails that have already occurred. Rolling back the prompt does not revoke refunds; forcibly switching model or tool versions for in-progress tasks may also break their state contracts. Prioritize binding tasks to candidate versions, and clearly define upgrade, drain, and recovery strategies.

### How strong is a 95% success rate?

Release evidence also needs enough sample information. More samples narrow the interval at a fixed observed proportion, but repeated similar tasks do not establish independence or replace group coverage. The interval is not a universal release threshold. [NIST: Wilson interval](https://www.itl.nist.gov/div898/handbook/prc/section2/prc241.htm)

$$
\frac{\hat p+\frac{z^2}{2n}\ \pm\ z\sqrt{\frac{\hat p(1-\hat p)}n+\frac{z^2}{4n^2}}}{1+\frac{z^2}{n}}
$$

**How strong is a 95% success rate?**

The same observed rate has a wider interval with fewer samples; zero failures do not prove zero failure probability.


## Post-release Closed Loop

Before entering regression for an incident, first save versions and necessary traces sufficient to explain the failure, and appropriately handle personal information and sensitive data. Directly copying real data into long-term evaluation sets may expand exposure scope; synthetic cases preserving the same failure conditions can be constructed.

A useful regression is not just "asking the same question again," but also recording expected results and failure locations. For example, cross-tenant ticket incidents can be split into executor rejection tests and end-to-end injection tests: the former verifies permission boundaries, the latter observes whether the model continues to correctly complete the original task. Run directly related tests after modification, then supplement regressions based on dependency impacts, and re-confirm necessary evidence belongs to the current combination before release.

Checklists will be adjusted as the system changes. Delete invalid, duplicate checks and retain reasons; add checks that cover real gaps; the maintenance goal is to make evidence better predict production behavior, rather than letting the number of checkboxes increase endlessly.
