---
title: Agent loops and executors
url: https://doc.liz6.com/en/ai/03-agent-systems/01-agent-loop
locale: en
area: ai
tags:
- Models & agents
- Agent Execution Systems
date: 2026-06-30
modified: 2026-09-10
description: Tool use turns proposals into verifiable execution records. Build a minimal loop around a stock query, then add result correlation, stopping states, dependencies and unknown write outcomes. Distinguish a proposed action, an executed action and an accepted task result.
---

# Agent loops and executors

Tool use turns proposals into verifiable execution records. Build a minimal loop around a stock query, then add result correlation, stopping states, dependencies and unknown write outcomes. Distinguish a proposed action, an executed action and an accepted task result.

## From Fixed Flows to Feedback-Driven Decisions

Fixed workflows have their main steps pre-arranged by the program, such as reading an order, extracting fields, validating, and storing; the model may participate in a specific step, but the main path is controlled by code. An Agent, on the other hand, allows the model to dynamically decide the next step based on observations, such as locating the relevant file after a test failure and then choosing a new fix action. The two can be combined: a deterministic outer flow can also contain restricted Agent sub-tasks. [Building effective agents](https://www.anthropic.com/engineering/building-effective-agents)

ReAct research organizes reasoning and action alternately, allowing the model to update subsequent processing based on environmental feedback. [ReAct](https://arxiv.org/abs/2210.03629) This does not define products by "whether there is a chat interface": chat systems can also call tools, and Agents might complete simple requests in a single answer.

Model reasoning generates representations and outputs, while external actions are implemented by an execution environment. Client-side tools are executed by the application; server-side tools may be executed by the provider. Therefore, permission and audit boundaries are distributed across actual execution components. One cannot assume that all operations happen locally on the host, nor treat model-generated command text as already executed.

<svg viewBox="0 0 760 385" xmlns="http://www.w3.org/2000/svg" role="img" aria-label="Model proposes action, executor verifies and executes, result enters next decision" style="max-width:100%;height:auto" font-family="Source Han Sans CN,Microsoft YaHei,sans-serif">
<defs><marker id="agent-execution-arrow" markerWidth="8" markerHeight="8" refX="7" refY="4" orient="auto"><path d="M0,0 L8,4 L0,8 Z" fill="#64748b"></path></marker></defs>
<rect width="760" height="385" rx="12" fill="#f8fafc"></rect>


<g transform="translate(0 0)"><rect x="30" y="85" width="200" height="75" rx="8" fill="#e0e7ff" stroke="#c7d2fe"></rect><text x="130.0" y="119.5" font-size="15" fill="#1e293b" text-anchor="middle" font-weight="600">Model Response</text><text x="130.0" y="141.5" font-size="12" fill="#475569" text-anchor="middle" font-weight="400">Text or structured call</text><rect x="280" y="85" width="200" height="75" rx="8" fill="#e0e7ff" stroke="#c7d2fe"></rect><text x="380.0" y="119.5" font-size="15" fill="#1e293b" text-anchor="middle" font-weight="600">Pre-execution Verification</text><text x="380.0" y="141.5" font-size="12" fill="#475569" text-anchor="middle" font-weight="400">Name, args, authority, budget</text><rect x="530" y="85" width="200" height="75" rx="8" fill="#e0e7ff" stroke="#c7d2fe"></rect><text x="630.0" y="119.5" font-size="15" fill="#1e293b" text-anchor="middle" font-weight="600">Tool Execution</text><text x="630.0" y="141.5" font-size="12" fill="#475569" text-anchor="middle" font-weight="400">Local or remote service</text><rect x="280" y="230" width="200" height="75" rx="8" fill="#e0e7ff" stroke="#c7d2fe"></rect><text x="380.0" y="264.5" font-size="15" fill="#1e293b" text-anchor="middle" font-weight="600">Result and Task Status</text><text x="380.0" y="286.5" font-size="12" fill="#475569" text-anchor="middle" font-weight="400">Call ID, return value, evidence</text><line x1="230" y1="122" x2="275" y2="122" stroke="#64748b" stroke-width="1.8" marker-end="url(#agent-execution-arrow)"></line><line x1="480" y1="122" x2="525" y2="122" stroke="#64748b" stroke-width="1.8" marker-end="url(#agent-execution-arrow)"></line><line x1="630" y1="165" x2="630" y2="267" stroke="#64748b" stroke-width="1.8" marker-end="url(#agent-execution-arrow)"></line><line x1="630" y1="267" x2="485" y2="267" stroke="#64748b" stroke-width="1.8" marker-end="url(#agent-execution-arrow)"></line><line x1="280" y1="267" x2="130" y2="267" stroke="#64748b" stroke-width="1.8" marker-end="url(#agent-execution-arrow)"></line><line x1="130" y1="267" x2="130" y2="165" stroke="#64748b" stroke-width="1.8" marker-end="url(#agent-execution-arrow)"></line></g><text x="24" y="29" font-size="19" fill="#0f172a" text-anchor="start" font-weight="700"><tspan x="24" dy="0">Model proposes action, executor verifies and executes, result enters next </tspan><tspan x="24" dy="25.650000000000002">decision</tspan></text><text x="24" y="338" font-size="13" fill="#475569" text-anchor="start" font-weight="400"><tspan x="24" dy="0">Model response end, a single tool return, and business goal completion are three different completion boundaries.</tspan></text>
</svg>

### Completion claims and evidence

Completion is a claim to verify. The model permits claiming without execution or observation, while acceptance independently checks evidence. A read before mutation cannot prove the later state.

**Completion claims and evidence**

Execution, observation and a completion claim are distinct; acceptance must inspect evidence.


## The Complete Path of a Single Tool Call

### Tool Definition is Both Model Interface and Execution Contract

The name helps the model identify the action, the description explains the purpose and limitations, and the schema defines the parameter structure. Taking a read-only inventory query as an example, the description should explain what is being queried (sellable quantity), which identifier the product uses, and the time or version meaning of the result; the phrase "get inventory" cannot express these boundaries.

```json
{
  "name": "lookup_stock",
  "description": "Query current sellable quantity by product SKU; read-only, does not reserve inventory.",
  "input_schema": {
    "type": "object",
    "properties": {"sku": {"type": "string"}},
    "required": ["sku"],
    "additionalProperties": false
  }
}
```

This is a Claude-style tool definition; other protocols may have different field names. The schema constrains the shape but does not automatically validate identity, whether the product belongs to the current tenant, or whether the operator has permission to view. Structured output cannot replace execution-side validation: the executor must only dispatch registered tools and re-verify inputs and permissions. [Tool use with Claude](https://platform.claude.com/docs/en/agents-and-tools/tool-use/overview)

A single query can be tracked with the following records:

| Stage | Example Record | Meaning to Preserve |
|---|---|---|
| Model proposes action | Call ID c17, lookup_stock, sku=A1 | This is a request, not yet executed |
| Pre-execution verification | Tool exists, parameters valid, subject has read permission | Scope allowed for execution |
| Tool actually returns | available=3, inventory version v83 | Which read result this is |
| Observation returned | Call ID c17 corresponds to the above result | Prevent result mismatch |
| Model continues decision | Answer sellable quantity or propose next step | Query does not equal reserved inventory |

Execution results should ideally include clear status, necessary data, and source location. Errors should also distinguish between product not found, invalid parameters, insufficient permissions, rate limiting, or unknown results; a generic "failed, retry" return will induce meaningless loops.

### Message Protocol Must Be Complete

Claude client tools associate `tool_use` with matching `tool_result`; a single response can contain multiple calls. Applications should preserve the complete assistant response and organize all corresponding results according to the protocol. Thinking, opaque states, and non-text content cannot be discarded by simplified logic that "only saves text."

Error format feedback may affect subsequent model behavior, but it does not "quietly train model weights" in the request. Protocol requirements and speculative behavioral explanations should be separated; specific tool message arrangement rules are subject to the [Tool Interface Documentation](https://platform.claude.com/docs/en/agents-and-tools/tool-use/overview).

## Response Stops and Task Status

The model no longer requesting tools only indicates that this model response has stopped. It may have correctly completed the task, or it may have missed verification, need user-supplied information, be truncated, or incorrectly claim success. Applications need to independently record task status, such as running, waiting, succeeded, failed, budget_exhausted, and decide when it is succeeded based on task acceptance criteria.

The following table shows common branches in the Claude interface; for a complete enumeration and subsequent request formats, see [Stop reasons and fallback](https://platform.claude.com/docs/en/build-with-claude/handling-stop-reasons).

| Stop Reason | Meaning | Handling Direction |
|---|---|---|
| `tool_use` | Requesting client tool | Complete verification, execution, and result return |
| `end_turn` | Model ends this round | Check task acceptance criteria and wrap up |
| `max_tokens` | Output reached limit | Check completeness; do not execute truncated parameters |
| `pause_turn` | Server-side long tool process paused | Resume according to interface requirements; application still responsible for loop limits |
| `refusal` | Model refuses current generation | Record reason, terminate or provide alternative according to business policy |

Streaming does not lift `max_tokens`. If tool parameter JSON is cut off halfway, do not guess remaining fields and execute; if a server-side tool is still running, do not treat pause as a signal for the client to re-execute the same side effect.

Taking "fix import duplicate submissions" as an example, acceptance criteria might be: patch exists, normal tests pass, restart recovery tests pass. If the model outputs "fixed" but only ran normal tests, the task is still incomplete. Acceptance standards should be clear at the start and associated with actual artifact versions.

## Dependencies, Replays, and Side Effects

### Parallel Eligibility Comes from Dependencies

Reading two independent documents simultaneously is usually parallelizable; "query inventory then reserve based on result" has data dependencies and cannot be parallelized by guessing parameters first. Even if two actions have independent inputs, if they modify the same file or resource simultaneously, conflicts may occur.

The executor needs to consider data dependencies, shared resources, service concurrency limits, and permissions in parallel. Parallelization saves overlapping wait times, with costs including peak load and result merging; one cannot unconditionally execute all concurrently just because the model proposes multiple calls in the same round.

### Call ID is Not the Business Idempotency Key

The call ID is used to associate the response back to the proposal. When the model retries, it may generate a new call ID, but express the same business action; in this case, deduplicating solely by call ID may still result in duplicate order creation or notification sending. Business operation keys should be generated and saved by the application according to clear transaction semantics, and verified by idempotent services to ensure the same key and request are consistent.

<svg viewBox="0 0 760 370.2190856933594" xmlns="http://www.w3.org/2000/svg" role="img" aria-label="Key question after timeout: Did the action fail, or did the result fail to return?" style="max-width:100%;height:auto" font-family="Source Han Sans CN,Microsoft YaHei,sans-serif">
<defs><marker id="agent-unknown-outcome-arrow" markerWidth="8" markerHeight="8" refX="7" refY="4" orient="auto"><path d="M0,0 L8,4 L0,8 Z" fill="#64748b"></path></marker></defs>
<rect width="760" height="370.2190856933594" rx="12" fill="#f8fafc"></rect>


<g transform="translate(0 14.06145715713501)"><rect x="25" y="62" width="150" height="48" rx="8" fill="#e0e7ff" stroke="#c7d2fe"></rect><text x="100.0" y="91.0" font-size="15" fill="#1e293b" text-anchor="middle" font-weight="600">Executor</text><rect x="315" y="62" width="150" height="48" rx="8" fill="#e0e7ff" stroke="#c7d2fe"></rect><text x="390.0" y="91.0" font-size="15" fill="#1e293b" text-anchor="middle" font-weight="600">Order Service</text><rect x="585" y="62" width="150" height="48" rx="8" fill="#e0e7ff" stroke="#c7d2fe"></rect><text x="660.0" y="91.0" font-size="15" fill="#1e293b" text-anchor="middle" font-weight="600">Database</text><line x1="100" y1="114" x2="100" y2="280" stroke="#64748b" stroke-width="1.8" stroke-dasharray="5 4"></line><line x1="390" y1="114" x2="390" y2="280" stroke="#64748b" stroke-width="1.8" stroke-dasharray="5 4"></line><line x1="660" y1="114" x2="660" y2="280" stroke="#64748b" stroke-width="1.8" stroke-dasharray="5 4"></line><line x1="100" y1="150" x2="390" y2="150" stroke="#64748b" stroke-width="1.8" marker-end="url(#agent-unknown-outcome-arrow)"></line><text x="245.0" y="140" font-size="13" fill="#334155" text-anchor="middle" font-weight="400">Request: Operation Key K42</text><line x1="390" y1="200" x2="660" y2="200" stroke="#64748b" stroke-width="1.8" marker-end="url(#agent-unknown-outcome-arrow)"></line><text x="525.0" y="190" font-size="13" fill="#334155" text-anchor="middle" font-weight="400">Submission Completed</text><line x1="390" y1="250" x2="230" y2="250" stroke="#64748b" stroke-width="1.8" stroke-dasharray="5 4"></line><text x="245" y="223" font-size="13" fill="#b45309" text-anchor="middle" font-weight="400">Response Lost</text><text x="208" y="256" font-size="22" fill="#b91c1c" text-anchor="start" font-weight="400">×</text></g><text x="24" y="29" font-size="19" fill="#0f172a" text-anchor="start" font-weight="700"><tspan x="24" dy="0">Key question after timeout: Did the action fail, or did the result fail to </tspan><tspan x="24" dy="25.650000000000002">return?</tspan></text><text x="24" y="327.061457157135" font-size="13" fill="#475569" text-anchor="start" font-weight="400"><tspan x="24" dy="0">Do not treat timeout directly as unexecuted; query by business operation key or retry according to service </tspan><tspan x="24" dy="17.55">idempotency protocol.</tspan></text>
</svg>

| Failure Status | Can Retry Directly? | Reasonable Handling |
|---|---|---|
| Read-only query timeout | Usually yes, but data may have changed | Limited retries and preserve read timestamp |
| Parameter validation failure | Retrying as-is is meaningless | Return specific field errors that can be corrected |
| No permission | Do not bypass by rephrasing | Adjust within authorized scope or report blockage |
| Write operation timeout, unknown result | Cannot assume unexecuted | Query result by operation key or use service idempotency protocol |
| Completed action, return lost | Redoing may cause duplicate side effects | Replay saved results and verify business status |

Cancellation is also not rollback: stopping model generation or canceling local wait does not guarantee that remotely started actions are revoked. The executor should distinguish between "confirmed unexecuted," "completed," and "unknown result," and eliminate unknown states first during recovery.

### After timeout, how many writes?

Separate network and business outcomes: no response does not mean no write. Server truth is visible here while the caller has an unknown result; querying or retrying the same key confirms it. Deduplication needs a server guarantee.

**After timeout, how many writes?**

Call IDs identify attempts; operation keys identify business operations. Server deduplication prevents repeating the same operation key.


## A Runnable Local Executor Example

The following simulates a read-only inventory query using preset model responses. It verifies tool whitelists, parameters, call result association, and request limits; it does not call the model API nor execute real business writes. Production systems also need persistent state, identity and permissions, timeouts, complete protocol adaptation, and auditing.

```python
STOCK = {"A1": 3}

def execute(call):
    if call.get("name") != "lookup_stock":
        return {"ok": False, "error": "unknown_tool"}
    args = call.get("input")
    if not isinstance(args, dict) or set(args) != {"sku"}:
        return {"ok": False, "error": "invalid_arguments"}
    if not isinstance(args["sku"], str):
        return {"ok": False, "error": "invalid_sku"}
    if args["sku"] not in STOCK:
        return {"ok": False, "error": "not_found"}
    return {"ok": True, "available": STOCK[args["sku"]]}

def run(model, max_requests=3):
    observations, seen = [], set()
    for _ in range(max_requests):
        response = model(list(observations))
        if response["kind"] == "final":
            # Model end does not equal business acceptance success.
            return {"state": "model_ended", "text": response["text"],
                    "observations": observations}
        if response["kind"] != "tool":
            return {"state": "invalid_response"}
        call = response["call"]
        call_id = call.get("id")
        if not isinstance(call_id, str) or not call_id or call_id in seen:
            return {"state": "invalid_call_id"}
        seen.add(call_id)
        observations.append({"call_id": call_id, "result": execute(call)})
    return {"state": "budget_exhausted", "observations": observations}

def scripted(observations):
    if not observations:
        return {"kind": "tool", "call": {
            "id": "c17", "name": "lookup_stock", "input": {"sku": "A1"}}}
    assert observations[0] == {
        "call_id": "c17", "result": {"ok": True, "available": 3}}
    return {"kind": "final", "text": "Current query found sellable quantity of 3, not yet reserved."}

result = run(scripted)
assert result["state"] == "model_ended"
assert execute({"name": "delete_stock"})["error"] == "unknown_tool"
assert execute({"name": "lookup_stock", "input": {"sku": 1}})["error"] == "invalid_sku"
assert run(scripted, max_requests=1)["state"] == "budget_exhausted"
print(result["text"])
```

The example exhausts the request budget after one tool call but before obtaining the final response, explicitly returning `budget_exhausted`. This is more accurate than mapping any loop exit to "success." Real applications should also persist completed calls to avoid losing side effect records after process restarts.

## Tool Boundaries and Loop Convergence

Setting "run at most ten rounds" as the final insurance is not enough. One should observe whether each round gains new evidence, modifies artifacts, or eliminates errors; when repeatedly getting the same failure with no change in conditions, repeating actions usually cannot make progress. Instead of infinitely resending, return more specific errors, change the retrieval scope, or report missing conditions.

Tool outputs need to be size-limited and retain readable positions. When truncating logs, clearly mark the truncation interval, total amount, and file location, so the model does not mistakenly believe it has received complete evidence. For context management, see [Context Engineering](/ai/02-context-and-interfaces/02-context-engineering.md).

Specialized tools expose parameters and business semantics to the executor, facilitating validation, authorization, and recording. Shell tools provide more general capabilities, but the command allowlist alone is insufficient to limit everything the program can do; actual permissions still need to be controlled via execution identity, file and network access scope, isolated environments, and resource limits. Specific strategies should match already authorized tasks, and not let every reversible operation degrade into repeated confirmation. See [Security and Protection](/ai/03-agent-systems/05-authority-and-tool-boundaries.md) for details.

## Supplement: Environment Simulators and Real Execution

Agent training and evaluation can use environment simulators: the policy proposes actions, the simulator generates observations, and then the policy continues. Language world models attempt to learn these environmental responses; for example, the [Qwen-AgentWorld Model Card](https://huggingface.co/Qwen/Qwen-AgentWorld-35B-A3B) describes simulation capabilities for multi-class interactive environments. This belongs to a different execution path from actually modifying files or calling business services.

Simulators may generate results that seem reasonable but are state-inconsistent. During evaluation, do not just look at whether a single observation looks real, but also at the causal consistency of continuous actions, resource constraints, and whether long-term state is maintained; high scores in a simulated environment cannot directly prove success rates in real tool execution. Suitability as a simulator or Agent should be measured separately; one cannot assert that a certain role is naturally feasible or infeasible based on the number of activated parameters.

Subsequent [MCP and Skills](/ai/03-agent-systems/02-tools-and-mcp.md) discusses how capabilities are integrated and provided on demand, and [Memory and State](/ai/03-agent-systems/04-memory-and-recovery.md) discusses what needs to be saved for long tasks and fault recovery.

Continue with：[Tool interfaces and MCP](/ai/03-agent-systems/02-tools-and-mcp)。
