Release validation, canaries and recovery

On this page

A release combines code, model, prompts, tools, policy and indexes. Assemble evidence for that exact combination before choosing a canary, wider rollout or stop. This article integrates validation; detailed formulas and execution rules remain in their dedicated articles.

Fixed Candidate Versions and Task Boundaries

First, clearly define what the system accomplishes and how completion is confirmed. Query-based assistants may require answers with verifiable sources; execution-based assistants also need to check actual business states. A completed model response, HTTP 200, or tools not throwing exceptions cannot independently prove task success. Manual processing is also a distinct result branch; its proportion and cost should be tracked and cannot be silently counted as automatic success.

Candidate versions must at least allow locating application code, model identifiers and parameters, prompt templates, tool schemas, and permission policies. When using retrieval, document document snapshots, chunking and embedding configurations, and index versions; when using persistent state, document state formats and migration rules. Do not just assign a version number to the prompt while letting other dependencies change silently after verification. For services that only support floating model aliases, record the actually obtainable version information and verification time, and acknowledge that changes by the provider reduce reproducibility.

Candidate VersionsModels, Prompts, Tools, IndexesValidation EvidenceQuality, authority, recovery, loadLimited RolloutObserve or Restore VersionWhat goes live is a set of versions; the basis for release must belong to this set of versionsConfiguration changes → Recheck affected evidencePassing offline does not mean no online risks; rolling back versions cannot revoke business actions that have already occurred.

Related explanations: Agent Loop, Multi-Agent Orchestration.

Pre-release Checklist

Record the responsible person, candidate version, verification time, evidence location, and conclusion for each item. Record not executed and failed items separately, but neither can masquerade as passed; for not applicable items, state the reason, and do not simply delete the checkboxes.

Quality and Model Configuration

Check ItemEvidence to RetainSignals That Cannot Replace It
[ ] Define task success and failureOutput requirements, business post-conditions, rejection/manual transfer rulesModel self-statement "completed"
[ ] Cover representative tasks and boundariesFixed evaluation set, grouped results by difficulty, language, tool type, etc.Single demonstration or overall average score
[ ] Retain independent verification setSplit between tuning set and verification set, repeated trial strategyRepeatedly tuning to full score on the same batch of questions
[ ] Calibrate scoring methodsProgram check coverage, manual review, judge misjudgment examples"Used program scoring so it must be objective"
[ ] Verify model capabilities and parametersParameters actually supported by the selected endpoint, verification of stop/reject/truncation branchesSetting high effort uniformly for all models
[ ] Verify structure and business constraintsSchema validation and negative examples for quantity, status, resource ownership, etc.JSON can be parsed

The "what to observe" and "who scores" in evaluation are two dimensions: business results can be checked by programs or evaluated by humans or model judges; do not treat outcome-based evaluation as a lower-level method placed below program scoring. One success does not prove stable repeated execution; the number of trials per task and aggregation methods should be stated. Anthropic: Agent Evaluation

Sampling, inference budget, and output limits serve different purposes. Do not mechanically translate temperature into effort, nor assume all new models have removed sampling parameters. Related principles can be found in Tokens and Sampling, Reasoning and Thinking, Evaluation and Observability, Structured Output.

Context, Knowledge, and State

Check ItemEvidence to RetainTypical Failure Branches
[ ] Reserve budget for complete requestsTotal budget for tool definitions, retrieval materials, history, and output reservesSingle document fits, but complete request exceeds limit
[ ] Verify retrieval and answer layersCandidate recall, ranking, citation correspondence, and behavior when no evidence existsAnswer is fluent but citations do not support the conclusion
[ ] Check version and permission changesIndex and cache tests after document updates, deletions, and permission revocationsOriginal text revoked, but old snippets still returned
[ ] Verify session recoveryRecovery records for completed steps, pending results, and external object IDsAfter recovery, unknown results are treated as unexecuted
[ ] Verify state compatibilityOld state reading, new state writing, and rollback compatibility scopeRolled-back application cannot read new checkpoints

Taking "prohibit cross-tenant reading" as an example, do not just check if the model finally outputs someone else's content; also confirm that retrieval and tool execution phases did not bring this content into the context. Caches and long-term memory are also access paths. Cache hits do not prove the answer used the latest data; cache hit rate targets need to be judged together with freshness requirements.

See Context Engineering, RAG, Memory and State. Weights, KV, and hardware budgets in model selection can be found in Model Architecture.

Execution Permissions, Reliability, and Observability

Check ItemEvidence to RetainTypical Failure Branches
[ ] Executor independently verifies authorizationRejection tests for cross-user/project, forged subjects, and extra parametersModel passes approved and is allowed through
[ ] Limit general tool capabilitiesIsolation verification for file boundaries, command parameters, outbound network, etc.Command name is legal but can write files arbitrarily
[ ] Check injection and sensitive data flowSynthetic injection cases, log masking, output-side processingTool materials elevated to trusted instructions
[ ] Distinguish timeouts from business failuresQueries for unknown write results, deduplication, and recovery recordsRe-executing actions that already succeeded after timeout
[ ] Constrain retries and queuingTotal deadline, max attempts, bounded queues, and overload strategiesSDK and outer-layer retries multiply backend calls
[ ] Observe complete tasksCorrelated IDs, key step durations, task results, and total costsOnly counting model response time and single token usage
[ ] Verify failure branchesRate limiting, tool exceptions, stream breaks, truncation, partial batch failuresStream request establishment success counts as completion

The names and statistical methods of observability fields vary by service. Based on interface documentation, clarify which usage fields include or exclude cache, inference, and other usage; do not simply add all numbers together. Cache reads being zero for a long time may come from requests not meeting conditions, traffic characteristics, or configuration issues; do not judge the cause of failure based solely on this value.

Existing task authorizations should be passed by the system along the operation chain; actions requiring additional confirmation should be handled according to specific policies; successive manual clicks do not replace resource permissions and idempotency mechanisms. See Security and Protection, Cost, Performance, and Reliability. Boundaries for access protocols and tool discovery can be found in MCP and Skills.

Evidence Status and Version Matching

Four states can be used: pass (meeting agreed conditions), fail (not meeting), not_run (not executed), not_applicable (reasonably and confirmed by responsible process as not applicable). The existence of a report does not mean it passed, and a passed report does not mean it belongs to the current candidate. For necessary checks, missing, failed, not executed, or version mismatch should all prevent automatic release.

The local teaching example below only demonstrates this rule and will not deploy anything. Abstract combinations of multiple actual dependency versions into candidate; real systems also need to verify evidence source, content, timeliness, and scope of application, and cannot trust arbitrary callers' self-reported pass.

from copy import deepcopy

REQUIRED = {"quality", "authorization", "recovery"}

def blockers(candidate, evidence):
    problems = []
    for name in sorted(REQUIRED):
        item = evidence.get(name)
        if not isinstance(item, dict):
            problems.append((name, "missing"))
        elif item.get("candidate") != candidate:
            problems.append((name, "stale"))
        elif item.get("status") != "pass":
            problems.append((name, "not_passed"))
        elif not isinstance(item.get("report"), str) or not item["report"].strip():
            problems.append((name, "no_report"))
    return problems

reports = {name: {"candidate": "release-17", "status": "pass",
                  "report": f"reports/{name}-17.json"} for name in REQUIRED}
assert blockers("release-17", reports) == []
assert len(blockers("release-18", reports)) == 3
for state in ["fail", "not_run", "not_applicable"]:
    changed = deepcopy(reports)
    changed["quality"]["status"] = state
    assert blockers("release-17", changed) == [("quality", "not_passed")]
changed = deepcopy(reports)
del changed["authorization"]
assert blockers("release-17", changed) == [("authorization", "missing")]
changed = deepcopy(reports)
changed["recovery"]["report"] = ""
assert blockers("release-17", changed) == [("recovery", "no_report")]
print("Match version for release; old versions, failures, missing items, and missing reports all prevent release")

Necessary items here do not allow skipping with "not applicable". Conditional items can be judged as not applicable by release rules in clear scenarios, for example, query services with no business writes do not need refund deduplication checks, but still need their own timeout recovery checks. Candidate versions cannot reduce the list of necessary items themselves.

Which candidate do green checks cover?

Checklist status should bind to the candidate, input set and check environment, not persist as a timeless boolean. This model expands revision only, separating a past pass from currently usable evidence.

Preparing the visual
Which candidate do green checks cover?

Release gates need valid evidence for the current candidate; previous green results do not automatically carry over.

Canarying and Recovery

After passing offline acceptance, observe the new version with limited real tasks and compare with comparable old version traffic. Canarying requires pre-defined metrics, observation conditions, and stop rules; setting traffic ratio to 1% does not automatically yield sufficient evidence. Low-traffic businesses may go a long time without encountering important failure scenarios. Google SRE: Canarying Releases

StageKey ConfirmationBasis for Continuing or Stopping
Shadow Verification (when applicable)Output and resource consumption under same inputWrite tools should be isolated or simulated, not repeat real side effects
Small-scale CanaryQuality groups, manual takeover, task costs, tail latencyMeet preset observation conditions and no stop rules triggered
Gradual ExpansionQuotas, queuing, cache, and downstream capacityStill within budget and service goals after scale growth
Rollback to Stable VersionNew task routing, in-flight tasks, and state compatibilityConfirm recovery results and handle already occurred side effects separately

For example, 95 successes out of 100 tasks does not mean the real success rate is proven to be at least 95%. Under independent, identically distributed simplifying assumptions, the two-sided 95% Wilson interval lower bound for 95/100 is approximately 88.8%; real tasks also have distribution changes and correlations. Whether to expand should be combined with sample size, key groups, and business tolerance; do not use a single point estimate to replace judgment.

The rollback plan needs to answer three different questions: how new requests return to old configurations, how in-flight tasks stop or complete, and how to verify effects like refunds/sending emails that have already occurred. Rolling back the prompt does not revoke refunds; forcibly switching model or tool versions for in-progress tasks may also break their state contracts. Prioritize binding tasks to candidate versions, and clearly define upgrade, drain, and recovery strategies.

How strong is a 95% success rate?

Release evidence also needs enough sample information. More samples narrow the interval at a fixed observed proportion, but repeated similar tasks do not establish independence or replace group coverage. The interval is not a universal release threshold. NIST: Wilson interval

Preparing the visual
How strong is a 95% success rate?

The same observed rate has a wider interval with fewer samples; zero failures do not prove zero failure probability.

Post-release Closed Loop

Before entering regression for an incident, first save versions and necessary traces sufficient to explain the failure, and appropriately handle personal information and sensitive data. Directly copying real data into long-term evaluation sets may expand exposure scope; synthetic cases preserving the same failure conditions can be constructed.

A useful regression is not just "asking the same question again," but also recording expected results and failure locations. For example, cross-tenant ticket incidents can be split into executor rejection tests and end-to-end injection tests: the former verifies permission boundaries, the latter observes whether the model continues to correctly complete the original task. Run directly related tests after modification, then supplement regressions based on dependency impacts, and re-confirm necessary evidence belongs to the current combination before release.

Checklists will be adjusted as the system changes. Delete invalid, duplicate checks and retain reasons; add checks that cover real gaps; the maintenance goal is to make evidence better predict production behavior, rather than letting the number of checkboxes increase endlessly.