Memory, state and recovery

On this page

Multi-step execution creates decisions, artifact versions, test records and operations with unknown outcomes. Separate information purpose from persistence, build checkpoints and conditional updates, and use recovery tests to see whether stored state is sufficient to continue.

Separate Lifecycle and Usage

Model parameters carry capabilities and knowledge acquired during training; application state carries the inputs, decisions, files, and execution results for this specific task. The current context is the material actually accessible during this inference, which is usually only a subset of the saved state. You cannot infer that "since an interface requires explicit history transmission, all model knowledge comes from history," nor can you treat "summaries" as inherently ephemeral data.

Task CheckpointGoals, progress, call logsLong-term MemoryPreferences, rules, sourced factsEvidence & ArtifactsFile versions, test reportsRetrieve, verify, assemble contextPersistence and current visibility are two different axesLoad only what is currently needed and authorizedSummaries can also be persisted; full history might also be written to disk; format names do not determine lifecycle.
Information FormPrimary UseRetained Across Restarts?
Conversation and tool historyReproduce interactions, resume protocolsDepends on whether it falls into persistent backend
Compressed summariesReduce context, quick recoveryCan be saved; not determined by summary format
Task checkpointsRestore execution position and to-dosRequires explicit persistence configuration
Long-term memoryReuse preferences or conventions across tasksRequires validity period, source, and access scope
Authoritative business dataFacts like inventory, orders, permissionsManaged by corresponding business systems

"Short-term memory" sometimes refers to state belonging only to a single thread, not necessarily RAM. For example, LangGraph distinguishes between thread checkpoints and cross-thread stores, but the in-memory implementation itself does not survive process restarts; you need to choose an appropriate persistent backend. LangGraph Persistence

If you resend the growing history every round, the input per round increases with the history, and the cumulative input for the entire multi-round task may grow even faster. Assuming h tokens are added per round, the count of resent new history for n rounds is h×n(n+1)/2; prefix caching can change computation or cost, but it cannot remove this history's use of context space. See Context Engineering for window management.

What Checkpoints Need to Save

Checkpoints should allow the recovery program to answer: What is the current goal? Which actions are definitely completed? Which action results are unknown? What needs to be checked before the next step? Saving only the final natural language summary often loses these boundaries.

Suppose an Agent is fixing an importer, has submitted a patch and run unit tests, but hasn't done a restart recovery test. You could use the following application customization record:

{
  "task_id": "importer-fix-42",
  "revision": 7,
  "state": "running",
  "goal": "Fix duplicate imports, keep existing interfaces",
  "artifacts": [{"path": "src/importer.py", "revision": "patch-3"}],
  "verified": [{"check": "unit", "artifact_revision": "patch-3",
                "report": "artifacts/unit-patch-3.txt"}],
  "pending": ["restart-recovery-test"],
  "unknown_operations": [],
  "next_action": "Verify patch-3 against test report, then execute recovery test"
}

revision identifies the version of this state record; artifact_revision identifies the version of the artifact the test targeted; the two are not the same counter. Actual systems can use commit IDs, content hashes, or their own version identifiers, but they must be able to locate the corresponding content.

Upon recovery, first read the checkpoint, then verify if the artifacts exist and match the version. If the file has been modified to patch-4, old test results only prove patch-3 and cannot directly continue marking "all verification complete." Finally, handle unknown_operations: when a remote write operation has been issued but the response is lost, query its true status first; do not blindly replay. See Agent Loop for side-effect recovery boundaries.

Checkpoints and external operations are usually not in the same transaction. Recording completion before execution creates a window where the record leads the fact; executing before recording creates a window where the fact leads the record. You need to handle this with mechanisms like business idempotency keys, result queries, or shared transactions, rather than relying on "auto-saving chat logs" to eliminate duplicate executions.

From Candidate Facts to Valid Memories

Save Information That Changes Future Actions

Information suitable for long-term storage is stable and reusable, such as "project release artifacts must include a compatibility report." Thousands of lines of complete output from a single tool are usually better saved as raw evidence, while memory retains only the conclusion and location. Both can coexist, avoiding repeated loading of large history segments every time.

A memory entry should at least be able to explain the content, scope, and source:

FieldExampleProblem Solved
ContentProject P's default report language is ChineseHow to change behavior next time
ScopeUser U, Project PDoes not apply to other users or projects
SourceUser explicitly requested this time, message IDDistinguish user decisions from model speculation
Statusactive, superseded, unverifiedDoes not treat old decisions as current constraints
Time or VersionUpdate time, applicable project versionDetermine if review is needed
Original LocationRequirement doc or record locationCan re-read in case of dispute

"One fact per file" facilitates manual management but is not a universal requirement. When data volume is large, concurrency is needed, or complex filtering is required, database entries may be more suitable; the key is that entries can be independently located, updated, and revoked, and sources and scopes are preserved during retrieval.

Memory Writes Can Also Fail

The model saying "users always prefer very short answers" might just be a preference inferred from a single task. Its speculative nature should be preserved; speculation should not be upgraded to permanent user authorization. When new requirements conflict with old memories, first judge the scope and time; explicit new decisions should replace old entries, rather than letting two conflicting records coexist.

Repeated summarization gradually loses qualifiers. For example, the original text is "test environments allow skipping approval," but the secondary summary becomes "allow skipping approval," changing the permission meaning. For such constraints, preserve the original location and qualified scope; upon retrieving the memory, still verify against the current task and authorization.

Claude's memory tool is a client-side execution interface: the model requests file operations, the application maps logical memory paths to its own storage, and returns the result. Declaring a tool does not mean persistence is implemented, nor does it mean the service automatically manages tenants and permissions. Memory tool

How do two workers overwrite state?

Lost updates need no storage corruption: two individually plausible writes suffice. CAS includes the version used for the decision; after rejection recompute, rather than relabeling the old write with a new version.

Preparing the visual
How do two workers overwrite state?

CAS compares expected versions and rejects stale overwrites; rereading still requires recomputing the update.

Versioned Updates and Merging

Why "Read Then Overwrite" Loses Decisions

A and B both read version v7. A adds "report includes source," and B adds "record time range." If both overwrite the same file entirely, B's later write might erase A's modifications. Guaranteeing atomic visibility for a single write does not solve the problem of modifying based on an old version.

Storage Version v7Editor ABoth modify based on v7Editor BBoth modify based on v7Submit first: v7 → v8Submit later: Expects v7, mismatchTwo editors read the same version, detecting overwrite conflicts upon submissionB must re-read v8 and decide how to merge; cannot automatically treat conflicts as successful overwrites.

One approach is conditional updates: carry the read version upon submission, and the storage only writes and increments if the version still matches. After a conflict occurs, re-read and determine if the two changes can be merged; if they involve mutually exclusive decisions, resolve them according to business rules, rather than mechanically concatenating them.

The following demonstrates conditional updates using a temporary SQLite database, and verifies the saved results by reopening after closing the connection. The stale versions of the two editors are simulated via sequential submissions; this is not a concurrency stress or power-failure test.

import sqlite3
import tempfile
from pathlib import Path

with tempfile.TemporaryDirectory() as directory:
    path = Path(directory) / "memory.sqlite"
    db = sqlite3.connect(path)
    db.execute("""CREATE TABLE memory (
        owner TEXT, key TEXT, version INTEGER NOT NULL, value TEXT,
        PRIMARY KEY(owner, key))""")
    db.execute("INSERT INTO memory VALUES (?, ?, ?, ?)",
               ("user-1", "report-rule", 7, "Record time range"))
    db.commit()

    def update(expected, value):
        with db:
            cursor = db.execute("""UPDATE memory
                SET value=?, version=version+1
                WHERE owner=? AND key=? AND version=?""",
                (value, "user-1", "report-rule", expected))
            return cursor.rowcount == 1

    assert update(7, "Record time range, and include source")
    assert not update(7, "Stale version overwrite")
    db.close()

    reopened = sqlite3.connect(path)
    row = reopened.execute(
        "SELECT version, value FROM memory WHERE owner=? AND key=?",
        ("user-1", "report-rule")).fetchone()
    assert row == (8, "Record time range, and include source")
    reopened.close()
    print(row)

SQLite's UPDATE only modifies records satisfying the WHERE clause; if no records match, it is not considered a SQL error; therefore, you must check the number of affected rows. SQLite UPDATE In the example, owner is fixed on the trusted application side; real services must derive accessible scopes from authenticated identities and cannot let model parameters specify someone else's owner.

Conditional updates prevent silent overwrites but do not guarantee that the merged content is true. Stricter systems also need version history, auditing, and deletion markers; if restoring backups or rebuilding indexes, you must avoid already revoked entries being re-invoked by old copies.

Recovery Testing and Isolation Boundaries

Test Persistence, Retrieval, and Correct Usage Separately

CheckMethodWhat Failure Means
Save successfulRead after closing backend connection or restarting processData might only be in RAM, or commit incomplete
Retrieval correctRetrieve using relevant and irrelevant tasks respectivelyIssues with index, scope, or sorting
Usage correctProvide old memory and new explicit requirementsOutdated records might override current needs
Concurrency safeTwo writers update based on the same versionMight silently overwrite
Task recoveryInterrupt before tool, after tool, after result saveMight cause duplicate side effects or false completion reports
Deletion effectiveRetrieve again after deletion and check derived indexesOld vectors, caches, or copies still active

Successful saving and reading are just the first hurdle. Acceptance testing for "can continue after closing and reopening" should also check if it repeatedly explores negated solutions, retains user constraints, and if to-dos match actual artifacts. Engineering experience with long-running Agents also emphasizes progress logging, environment checks, and verifiable incremental results. Effective harnesses for long-running agents

File Directories Cannot Replace Full Permission Models

Self-managed file memories need to constrain root directories, resolve paths, handle symbolic links, and manage write races. Checking resolve() before open() has a time gap; in directories subject to concurrent modification, a TOCTOU (Time-of-Check to Time-of-Use) race might occur; you cannot describe a single string or path check as a complete file isolation solution. Execution identity, directory permissions, and appropriate restricted file operation mechanisms need to jointly bear the boundary.

Memories are not suitable for storing key bodies; credentials are provided by dedicated credential management components; document links and resource handles in entries must also have permissions checked upon use. Filter by user, project, and authorization scope before retrieval; do not fetch the entire database and then ask the model "don't look at other users' data." These requirements work in conjunction with the tool boundaries in Security and Protection.

Restored state: is evidence still valid?

Recovery must reestablish trustworthy facts. Unchanged filenames and summaries do not mean tests cover current content; binding reports to artifact versions or hashes reveals stale evidence.

Preparing the visual
Restored state: is evidence still valid?

Checkpoints retain references and known facts; old tests do not validate new artifacts, and unknown effects require reconciliation.

Continue with:Authority, trust and tool boundaries。