Agent loops and executors

On this page

Tool use turns proposals into verifiable execution records. Build a minimal loop around a stock query, then add result correlation, stopping states, dependencies and unknown write outcomes. Distinguish a proposed action, an executed action and an accepted task result.

From Fixed Flows to Feedback-Driven Decisions

Fixed workflows have their main steps pre-arranged by the program, such as reading an order, extracting fields, validating, and storing; the model may participate in a specific step, but the main path is controlled by code. An Agent, on the other hand, allows the model to dynamically decide the next step based on observations, such as locating the relevant file after a test failure and then choosing a new fix action. The two can be combined: a deterministic outer flow can also contain restricted Agent sub-tasks. Building effective agents

ReAct research organizes reasoning and action alternately, allowing the model to update subsequent processing based on environmental feedback. ReAct This does not define products by "whether there is a chat interface": chat systems can also call tools, and Agents might complete simple requests in a single answer.

Model reasoning generates representations and outputs, while external actions are implemented by an execution environment. Client-side tools are executed by the application; server-side tools may be executed by the provider. Therefore, permission and audit boundaries are distributed across actual execution components. One cannot assume that all operations happen locally on the host, nor treat model-generated command text as already executed.

Model ResponseText or structured callPre-execution VerificationName, args, authority, budgetTool ExecutionLocal or remote serviceResult and Task StatusCall ID, return value, evidenceModel proposes action, executor verifies and executes, result enters next decisionModel response end, a single tool return, and business goal completion are three different completion boundaries.

Completion claims and evidence

Completion is a claim to verify. The model permits claiming without execution or observation, while acceptance independently checks evidence. A read before mutation cannot prove the later state.

Preparing the visual
Completion claims and evidence

Execution, observation and a completion claim are distinct; acceptance must inspect evidence.

The Complete Path of a Single Tool Call

Tool Definition is Both Model Interface and Execution Contract

The name helps the model identify the action, the description explains the purpose and limitations, and the schema defines the parameter structure. Taking a read-only inventory query as an example, the description should explain what is being queried (sellable quantity), which identifier the product uses, and the time or version meaning of the result; the phrase "get inventory" cannot express these boundaries.

{
  "name": "lookup_stock",
  "description": "Query current sellable quantity by product SKU; read-only, does not reserve inventory.",
  "input_schema": {
    "type": "object",
    "properties": {"sku": {"type": "string"}},
    "required": ["sku"],
    "additionalProperties": false
  }
}

This is a Claude-style tool definition; other protocols may have different field names. The schema constrains the shape but does not automatically validate identity, whether the product belongs to the current tenant, or whether the operator has permission to view. Structured output cannot replace execution-side validation: the executor must only dispatch registered tools and re-verify inputs and permissions. Tool use with Claude

A single query can be tracked with the following records:

StageExample RecordMeaning to Preserve
Model proposes actionCall ID c17, lookup_stock, sku=A1This is a request, not yet executed
Pre-execution verificationTool exists, parameters valid, subject has read permissionScope allowed for execution
Tool actually returnsavailable=3, inventory version v83Which read result this is
Observation returnedCall ID c17 corresponds to the above resultPrevent result mismatch
Model continues decisionAnswer sellable quantity or propose next stepQuery does not equal reserved inventory

Execution results should ideally include clear status, necessary data, and source location. Errors should also distinguish between product not found, invalid parameters, insufficient permissions, rate limiting, or unknown results; a generic "failed, retry" return will induce meaningless loops.

Message Protocol Must Be Complete

Claude client tools associate tool_use with matching tool_result; a single response can contain multiple calls. Applications should preserve the complete assistant response and organize all corresponding results according to the protocol. Thinking, opaque states, and non-text content cannot be discarded by simplified logic that "only saves text."

Error format feedback may affect subsequent model behavior, but it does not "quietly train model weights" in the request. Protocol requirements and speculative behavioral explanations should be separated; specific tool message arrangement rules are subject to the Tool Interface Documentation.

Response Stops and Task Status

The model no longer requesting tools only indicates that this model response has stopped. It may have correctly completed the task, or it may have missed verification, need user-supplied information, be truncated, or incorrectly claim success. Applications need to independently record task status, such as running, waiting, succeeded, failed, budget_exhausted, and decide when it is succeeded based on task acceptance criteria.

The following table shows common branches in the Claude interface; for a complete enumeration and subsequent request formats, see Stop reasons and fallback.

Stop ReasonMeaningHandling Direction
tool_useRequesting client toolComplete verification, execution, and result return
end_turnModel ends this roundCheck task acceptance criteria and wrap up
max_tokensOutput reached limitCheck completeness; do not execute truncated parameters
pause_turnServer-side long tool process pausedResume according to interface requirements; application still responsible for loop limits
refusalModel refuses current generationRecord reason, terminate or provide alternative according to business policy

Streaming does not lift max_tokens. If tool parameter JSON is cut off halfway, do not guess remaining fields and execute; if a server-side tool is still running, do not treat pause as a signal for the client to re-execute the same side effect.

Taking "fix import duplicate submissions" as an example, acceptance criteria might be: patch exists, normal tests pass, restart recovery tests pass. If the model outputs "fixed" but only ran normal tests, the task is still incomplete. Acceptance standards should be clear at the start and associated with actual artifact versions.

Dependencies, Replays, and Side Effects

Parallel Eligibility Comes from Dependencies

Reading two independent documents simultaneously is usually parallelizable; "query inventory then reserve based on result" has data dependencies and cannot be parallelized by guessing parameters first. Even if two actions have independent inputs, if they modify the same file or resource simultaneously, conflicts may occur.

The executor needs to consider data dependencies, shared resources, service concurrency limits, and permissions in parallel. Parallelization saves overlapping wait times, with costs including peak load and result merging; one cannot unconditionally execute all concurrently just because the model proposes multiple calls in the same round.

Call ID is Not the Business Idempotency Key

The call ID is used to associate the response back to the proposal. When the model retries, it may generate a new call ID, but express the same business action; in this case, deduplicating solely by call ID may still result in duplicate order creation or notification sending. Business operation keys should be generated and saved by the application according to clear transaction semantics, and verified by idempotent services to ensure the same key and request are consistent.

ExecutorOrder ServiceDatabaseRequest: Operation Key K42Submission CompletedResponse Lost×Key question after timeout: Did the action fail, or did the result fail to return?Do not treat timeout directly as unexecuted; query by business operation key or retry according to service idempotency protocol.
Failure StatusCan Retry Directly?Reasonable Handling
Read-only query timeoutUsually yes, but data may have changedLimited retries and preserve read timestamp
Parameter validation failureRetrying as-is is meaninglessReturn specific field errors that can be corrected
No permissionDo not bypass by rephrasingAdjust within authorized scope or report blockage
Write operation timeout, unknown resultCannot assume unexecutedQuery result by operation key or use service idempotency protocol
Completed action, return lostRedoing may cause duplicate side effectsReplay saved results and verify business status

Cancellation is also not rollback: stopping model generation or canceling local wait does not guarantee that remotely started actions are revoked. The executor should distinguish between "confirmed unexecuted," "completed," and "unknown result," and eliminate unknown states first during recovery.

After timeout, how many writes?

Separate network and business outcomes: no response does not mean no write. Server truth is visible here while the caller has an unknown result; querying or retrying the same key confirms it. Deduplication needs a server guarantee.

Preparing the visual
After timeout, how many writes?

Call IDs identify attempts; operation keys identify business operations. Server deduplication prevents repeating the same operation key.

A Runnable Local Executor Example

The following simulates a read-only inventory query using preset model responses. It verifies tool whitelists, parameters, call result association, and request limits; it does not call the model API nor execute real business writes. Production systems also need persistent state, identity and permissions, timeouts, complete protocol adaptation, and auditing.

STOCK = {"A1": 3}

def execute(call):
    if call.get("name") != "lookup_stock":
        return {"ok": False, "error": "unknown_tool"}
    args = call.get("input")
    if not isinstance(args, dict) or set(args) != {"sku"}:
        return {"ok": False, "error": "invalid_arguments"}
    if not isinstance(args["sku"], str):
        return {"ok": False, "error": "invalid_sku"}
    if args["sku"] not in STOCK:
        return {"ok": False, "error": "not_found"}
    return {"ok": True, "available": STOCK[args["sku"]]}

def run(model, max_requests=3):
    observations, seen = [], set()
    for _ in range(max_requests):
        response = model(list(observations))
        if response["kind"] == "final":
            # Model end does not equal business acceptance success.
            return {"state": "model_ended", "text": response["text"],
                    "observations": observations}
        if response["kind"] != "tool":
            return {"state": "invalid_response"}
        call = response["call"]
        call_id = call.get("id")
        if not isinstance(call_id, str) or not call_id or call_id in seen:
            return {"state": "invalid_call_id"}
        seen.add(call_id)
        observations.append({"call_id": call_id, "result": execute(call)})
    return {"state": "budget_exhausted", "observations": observations}

def scripted(observations):
    if not observations:
        return {"kind": "tool", "call": {
            "id": "c17", "name": "lookup_stock", "input": {"sku": "A1"}}}
    assert observations[0] == {
        "call_id": "c17", "result": {"ok": True, "available": 3}}
    return {"kind": "final", "text": "Current query found sellable quantity of 3, not yet reserved."}

result = run(scripted)
assert result["state"] == "model_ended"
assert execute({"name": "delete_stock"})["error"] == "unknown_tool"
assert execute({"name": "lookup_stock", "input": {"sku": 1}})["error"] == "invalid_sku"
assert run(scripted, max_requests=1)["state"] == "budget_exhausted"
print(result["text"])

The example exhausts the request budget after one tool call but before obtaining the final response, explicitly returning budget_exhausted. This is more accurate than mapping any loop exit to "success." Real applications should also persist completed calls to avoid losing side effect records after process restarts.

Tool Boundaries and Loop Convergence

Setting "run at most ten rounds" as the final insurance is not enough. One should observe whether each round gains new evidence, modifies artifacts, or eliminates errors; when repeatedly getting the same failure with no change in conditions, repeating actions usually cannot make progress. Instead of infinitely resending, return more specific errors, change the retrieval scope, or report missing conditions.

Tool outputs need to be size-limited and retain readable positions. When truncating logs, clearly mark the truncation interval, total amount, and file location, so the model does not mistakenly believe it has received complete evidence. For context management, see Context Engineering.

Specialized tools expose parameters and business semantics to the executor, facilitating validation, authorization, and recording. Shell tools provide more general capabilities, but the command allowlist alone is insufficient to limit everything the program can do; actual permissions still need to be controlled via execution identity, file and network access scope, isolated environments, and resource limits. Specific strategies should match already authorized tasks, and not let every reversible operation degrade into repeated confirmation. See Security and Protection for details.

Supplement: Environment Simulators and Real Execution

Agent training and evaluation can use environment simulators: the policy proposes actions, the simulator generates observations, and then the policy continues. Language world models attempt to learn these environmental responses; for example, the Qwen-AgentWorld Model Card describes simulation capabilities for multi-class interactive environments. This belongs to a different execution path from actually modifying files or calling business services.

Simulators may generate results that seem reasonable but are state-inconsistent. During evaluation, do not just look at whether a single observation looks real, but also at the causal consistency of continuous actions, resource constraints, and whether long-term state is maintained; high scores in a simulated environment cannot directly prove success rates in real tool execution. Suitability as a simulator or Agent should be measured separately; one cannot assert that a certain role is naturally feasible or infeasible based on the number of activated parameters.

Subsequent MCP and Skills discusses how capabilities are integrated and provided on demand, and Memory and State discusses what needs to be saved for long tasks and fault recovery.

Continue with:Tool interfaces and MCP。