Tool use turns proposals into verifiable execution records. Build a minimal loop around a stock query, then add result correlation, stopping states, dependencies and unknown write outcomes. Distinguish a proposed action, an executed action and an accepted task result.
From Fixed Flows to Feedback-Driven Decisions
Fixed workflows have their main steps pre-arranged by the program, such as reading an order, extracting fields, validating, and storing; the model may participate in a specific step, but the main path is controlled by code. An Agent, on the other hand, allows the model to dynamically decide the next step based on observations, such as locating the relevant file after a test failure and then choosing a new fix action. The two can be combined: a deterministic outer flow can also contain restricted Agent sub-tasks. Building effective agents
ReAct research organizes reasoning and action alternately, allowing the model to update subsequent processing based on environmental feedback. ReAct This does not define products by "whether there is a chat interface": chat systems can also call tools, and Agents might complete simple requests in a single answer.
Model reasoning generates representations and outputs, while external actions are implemented by an execution environment. Client-side tools are executed by the application; server-side tools may be executed by the provider. Therefore, permission and audit boundaries are distributed across actual execution components. One cannot assume that all operations happen locally on the host, nor treat model-generated command text as already executed.
Completion claims and evidence
Completion is a claim to verify. The model permits claiming without execution or observation, while acceptance independently checks evidence. A read before mutation cannot prove the later state.
Preparing the visual
Completion claims and evidence
Execution, observation and a completion claim are distinct; acceptance must inspect evidence.
Tool Definition is Both Model Interface and Execution Contract
The name helps the model identify the action, the description explains the purpose and limitations, and the schema defines the parameter structure. Taking a read-only inventory query as an example, the description should explain what is being queried (sellable quantity), which identifier the product uses, and the time or version meaning of the result; the phrase "get inventory" cannot express these boundaries.
{"name":"lookup_stock","description":"Query current sellable quantity by product SKU; read-only, does not reserve inventory.","input_schema":{"type":"object","properties":{"sku":{"type":"string"}},"required":["sku"],"additionalProperties":false}}
This is a Claude-style tool definition; other protocols may have different field names. The schema constrains the shape but does not automatically validate identity, whether the product belongs to the current tenant, or whether the operator has permission to view. Structured output cannot replace execution-side validation: the executor must only dispatch registered tools and re-verify inputs and permissions. Tool use with Claude
A single query can be tracked with the following records:
Stage
Example Record
Meaning to Preserve
Model proposes action
Call ID c17, lookup_stock, sku=A1
This is a request, not yet executed
Pre-execution verification
Tool exists, parameters valid, subject has read permission
Scope allowed for execution
Tool actually returns
available=3, inventory version v83
Which read result this is
Observation returned
Call ID c17 corresponds to the above result
Prevent result mismatch
Model continues decision
Answer sellable quantity or propose next step
Query does not equal reserved inventory
Execution results should ideally include clear status, necessary data, and source location. Errors should also distinguish between product not found, invalid parameters, insufficient permissions, rate limiting, or unknown results; a generic "failed, retry" return will induce meaningless loops.
Message Protocol Must Be Complete
Claude client tools associate tool_use with matching tool_result; a single response can contain multiple calls. Applications should preserve the complete assistant response and organize all corresponding results according to the protocol. Thinking, opaque states, and non-text content cannot be discarded by simplified logic that "only saves text."
Error format feedback may affect subsequent model behavior, but it does not "quietly train model weights" in the request. Protocol requirements and speculative behavioral explanations should be separated; specific tool message arrangement rules are subject to the Tool Interface Documentation.
Response Stops and Task Status
The model no longer requesting tools only indicates that this model response has stopped. It may have correctly completed the task, or it may have missed verification, need user-supplied information, be truncated, or incorrectly claim success. Applications need to independently record task status, such as running, waiting, succeeded, failed, budget_exhausted, and decide when it is succeeded based on task acceptance criteria.
The following table shows common branches in the Claude interface; for a complete enumeration and subsequent request formats, see Stop reasons and fallback.
Stop Reason
Meaning
Handling Direction
tool_use
Requesting client tool
Complete verification, execution, and result return
end_turn
Model ends this round
Check task acceptance criteria and wrap up
max_tokens
Output reached limit
Check completeness; do not execute truncated parameters
pause_turn
Server-side long tool process paused
Resume according to interface requirements; application still responsible for loop limits
refusal
Model refuses current generation
Record reason, terminate or provide alternative according to business policy
Streaming does not lift max_tokens. If tool parameter JSON is cut off halfway, do not guess remaining fields and execute; if a server-side tool is still running, do not treat pause as a signal for the client to re-execute the same side effect.
Taking "fix import duplicate submissions" as an example, acceptance criteria might be: patch exists, normal tests pass, restart recovery tests pass. If the model outputs "fixed" but only ran normal tests, the task is still incomplete. Acceptance standards should be clear at the start and associated with actual artifact versions.
Dependencies, Replays, and Side Effects
Parallel Eligibility Comes from Dependencies
Reading two independent documents simultaneously is usually parallelizable; "query inventory then reserve based on result" has data dependencies and cannot be parallelized by guessing parameters first. Even if two actions have independent inputs, if they modify the same file or resource simultaneously, conflicts may occur.
The executor needs to consider data dependencies, shared resources, service concurrency limits, and permissions in parallel. Parallelization saves overlapping wait times, with costs including peak load and result merging; one cannot unconditionally execute all concurrently just because the model proposes multiple calls in the same round.
Call ID is Not the Business Idempotency Key
The call ID is used to associate the response back to the proposal. When the model retries, it may generate a new call ID, but express the same business action; in this case, deduplicating solely by call ID may still result in duplicate order creation or notification sending. Business operation keys should be generated and saved by the application according to clear transaction semantics, and verified by idempotent services to ensure the same key and request are consistent.
Failure Status
Can Retry Directly?
Reasonable Handling
Read-only query timeout
Usually yes, but data may have changed
Limited retries and preserve read timestamp
Parameter validation failure
Retrying as-is is meaningless
Return specific field errors that can be corrected
No permission
Do not bypass by rephrasing
Adjust within authorized scope or report blockage
Write operation timeout, unknown result
Cannot assume unexecuted
Query result by operation key or use service idempotency protocol
Completed action, return lost
Redoing may cause duplicate side effects
Replay saved results and verify business status
Cancellation is also not rollback: stopping model generation or canceling local wait does not guarantee that remotely started actions are revoked. The executor should distinguish between "confirmed unexecuted," "completed," and "unknown result," and eliminate unknown states first during recovery.
After timeout, how many writes?
Separate network and business outcomes: no response does not mean no write. Server truth is visible here while the caller has an unknown result; querying or retrying the same key confirms it. Deduplication needs a server guarantee.
Preparing the visual
After timeout, how many writes?
Call IDs identify attempts; operation keys identify business operations. Server deduplication prevents repeating the same operation key.
{"id":"ai-unknown-write","title":"After timeout, how many writes?","summary":"Call IDs identify attempts; operation keys identify business operations. Server deduplication prevents repeating the same operation key.","height":1000,"html":"<h2 data-i18n=\"heading\"></h2><p class=\"intro\" data-i18n=\"intro\"></p><div class=\"actions\"><button id=\"send\" data-i18n=\"send\"></button><button id=\"same\" data-i18n=\"same\"></button><button id=\"fresh\" data-i18n=\"fresh\"></button><button id=\"query\" data-i18n=\"query\"></button><button id=\"cancel\" data-i18n=\"cancel\"></button><button id=\"reset\" data-i18n=\"reset\"></button></div><div class=\"metrics\"><div><span data-i18n=\"effects\"></span><strong id=\"effects\"></strong></div><div><span data-i18n=\"known\"></span><strong id=\"known\"></strong></div></div><section class=\"panel\"><h3 data-i18n=\"history\"></h3><div id=\"history\"></div></section><div class=\"copy\"><p id=\"explanation\" role=\"status\"></p><p class=\"reserve\" aria-hidden=\"true\" data-i18n=\"explain\"></p></div><details class=\"scope\"><summary data-i18n=\"scopeLabel\"></summary><p data-i18n=\"scope\"></p></details>","css":".intro{margin:8px 0 18px;color:var(--muted);font-size:14px}.control{display:grid;grid-template-columns:minmax(0,1fr) auto;gap:6px 12px;align-items:center;margin:14px 0;font-size:13px}.control output{color:var(--accent);text-align:right;font-variant-numeric:tabular-nums;min-width:4em}.control input{grid-column:1/-1;width:100%;margin:0}.presets{display:flex;gap:4px;align-items:stretch}.presets button{flex:1;min-width:0;overflow-wrap:anywhere;font-size:13px;min-height:44px}.actions{display:grid;grid-template-columns:repeat(2,minmax(0,1fr));gap:8px;margin:16px 0}.actions button{min-width:0;min-height:44px;font-size:13px;overflow-wrap:anywhere}.copy{display:grid;font-size:14px;margin:16px 0}.copy>*{grid-area:1/1;margin:0;overflow-wrap:anywhere}.reserve{visibility:hidden}.scope{margin-top:14px;color:var(--muted);font-size:12px}.scope summary{padding:8px 0;cursor:pointer}.scope p{margin-top:10px;overflow-wrap:anywhere}.metrics{display:grid;grid-template-columns:repeat(2,minmax(0,1fr));gap:14px;margin:18px 0}.metrics>div{border-left:2px solid var(--accent);padding:0 6px 0 10px;min-width:0}.metrics span{display:block;min-height:3.6em;color:var(--muted);font-size:12px;overflow-wrap:anywhere}.metrics strong{font-size:21px;display:block;min-height:2em;font-weight:550;overflow-wrap:anywhere;font-variant-numeric:tabular-nums}.check{display:flex;align-items:center;gap:9px;font-size:13px;margin:14px 0}.check input{width:18px;height:18px;flex-shrink:0;accent-color:var(--accent)}select,input[type=text],input[type=number]{background:var(--paper);color:var(--ink);font:inherit;font-size:14px;padding:10px;border:1px solid var(--rule);border-radius:7px;max-width:100%;min-width:0}select{width:100%;min-height:44px}.panel{padding:14px;border:1px solid var(--rule);border-radius:9px;margin:16px 0;min-width:0}.panel>header{font-size:13px;font-weight:600;margin-bottom:12px;color:var(--muted)}.panel>div{overflow-wrap:anywhere}.track-row{margin:16px 0}.track-row>span{display:block;font-size:12px;color:var(--muted);margin-bottom:8px}.track{display:flex;gap:5px;min-width:0}.cell{flex:1;min-width:0;min-height:56px;border:1px solid var(--rule);border-radius:5px;background:var(--surface);display:flex;align-items:center;justify-content:center;text-align:center;padding:6px 3px;font-size:12px;overflow-wrap:anywhere}.cell.on{background:var(--accent-soft);border-color:var(--accent)}.cell.gold{background:var(--second-soft);border-color:var(--second)}.cell.empty{border-style:dashed;color:var(--muted)}.cell.error{border-color:var(--second);text-decoration:line-through}.table{display:grid;gap:5px;font-size:12px}.tr{display:grid;gap:5px}.tr>div{background:var(--surface);border-radius:4px;padding:9px 6px;min-width:0;min-height:5em;overflow-wrap:anywhere;display:flex;align-items:center}.tr>.gold{background:var(--second-soft)}.tr>.on{background:var(--accent-soft)}.formula,.source{font:13px/1.8 ui-monospace,monospace;white-space:pre-wrap;overflow-wrap:anywhere;margin:16px 0}.formula{border-top:1px solid var(--rule);padding-top:12px;min-height:7.2em}.source{min-height:8em;background:var(--surface);padding:12px;border-radius:8px}.intervals{margin:16px 0}.interval{position:relative;height:38px;background:var(--surface);margin:6px 0;border-radius:4px;overflow:hidden}.interval i{position:absolute;height:100%;background:var(--accent-soft);border-left:2px solid var(--accent)}.interval i.gold{background:var(--second-soft);border-color:var(--second)}.interval span{position:relative;z-index:1;font:12px/38px ui-monospace,monospace;padding-left:7px}.bars{display:grid;gap:10px;margin:16px 0}.bar-row{display:grid;grid-template-columns:44px minmax(0,1fr) 62px;gap:8px;align-items:center;font:12px ui-monospace,monospace}.bar-track{height:24px;background:var(--surface);border-radius:4px;overflow:hidden}.bar-track i{display:block;height:100%;background:var(--accent);transition:width .2s}.bar-track i.gold{background:var(--second)}.bar-track i.empty{opacity:.2}.bar-row output{text-align:right}.diagram{display:block;width:100%;height:auto;margin:18px 0}.diagram text{font:13px ui-monospace,monospace;fill:var(--ink)}.node{fill:var(--surface);stroke:var(--rule)}.node.on{fill:var(--accent-soft);stroke:var(--accent)}.node.gold{fill:var(--second-soft);stroke:var(--second)}.edge{stroke:var(--rule);stroke-width:2;fill:none}.edge.on{stroke:var(--accent)}.matrix{display:grid;grid-template-columns:repeat(3,minmax(0,1fr));gap:6px;margin:16px 0}.matrix button{min-width:0;min-height:44px;font:13px ui-monospace,monospace}.matrix button[aria-pressed=true]{background:var(--accent-soft);border-color:var(--accent)}.matrix button.masked{opacity:.35}.legend{font-size:12px;color:var(--muted);min-height:3.6em;overflow-wrap:anywhere}.numeric{font-variant-numeric:tabular-nums}.scope p{line-break:strict}@media(max-width:480px){.presets button,.actions button{font-size:12px}.panel{padding:12px}.metrics{gap:10px}.metrics strong{font-size:19px}.bar-row{grid-template-columns:34px minmax(0,1fr) 58px;gap:6px}}@media(prefers-reduced-motion:reduce){*,*::before,*::after{transition:none!important;animation:none!important}}\n\n.panel>h3{font-size:13px;font-weight:600;margin:0 0 12px;color:var(--muted)}\n\n.bar-row{grid-template-columns:44px minmax(0,1fr) 82px}.bar-row output{white-space:nowrap}@media(max-width:480px){.bar-row{grid-template-columns:34px minmax(0,1fr) 76px}}\n\n.metrics strong{font-size:16px;min-height:4em}","js":"const $=s=>document.querySelector(s);const set=(id,v)=>$('#'+id).textContent=v;const t=k=>viz.t(k);function cells(id,values){$('#'+id).replaceChildren(...values.map(v=>{const e=document.createElement('div');e.className='cell '+(v.cls||'');e.textContent=v.text;e.title=v.title||v.text;return e;}));}function pressed(mode){document.querySelectorAll('[data-mode]').forEach(e=>e.setAttribute('aria-pressed',String(e.dataset.mode===mode)));}const esc=s=>String(s).replace(/[&<>\"']/g,c=>({'&':'&','<':'<','>':'>','\"':'"',\"'\":'''}[c]));const fmt=s=>'{'+[...s].sort().join(', ')+'}';function graph(id,nodes,edges){const el=$('#'+id);el.setAttribute('viewBox','0 0 360 200');el.innerHTML='<defs><marker id=\"arrow-'+id+'\" viewBox=\"0 0 10 10\" refX=\"8\" refY=\"5\" markerWidth=\"5\" markerHeight=\"5\" orient=\"auto-start-reverse\"><path d=\"M 0 0 L 10 5 L 0 10 z\" fill=\"var(--muted)\"/></marker></defs>'+edges.map(e=>{const a=nodes[e[0]],b=nodes[e[1]],dx=b.x-a.x,dy=b.y-a.y,d=Math.hypot(dx,dy)||1;if(edges.some(r=>r[0]===e[1]&&r[1]===e[0])){const nx=-dy/d,ny=dx/d;return '<path class=\"edge '+(e[2]?'on':'')+'\" d=\"M '+(a.x+dx/d*22+nx*12)+' '+(a.y+dy/d*22+ny*12)+' Q '+((a.x+b.x)/2+nx*38)+' '+((a.y+b.y)/2+ny*38)+' '+(b.x-dx/d*25+nx*12)+' '+(b.y-dy/d*25+ny*12)+'\" marker-end=\"url(#arrow-'+id+')\"/>';}return '<line class=\"edge '+(e[2]?'on':'')+'\" x1=\"'+(a.x+dx/d*25)+'\" y1=\"'+(a.y+dy/d*25)+'\" x2=\"'+(b.x-dx/d*28)+'\" y2=\"'+(b.y-dy/d*28)+'\" marker-end=\"url(#arrow-'+id+')\"/>';}).join('')+nodes.map(n=>'<g><circle class=\"node '+(n.cls||'')+'\" cx=\"'+n.x+'\" cy=\"'+n.y+'\" r=\"25\"/><text x=\"'+n.x+'\" y=\"'+(n.y+4)+'\" text-anchor=\"middle\">'+esc(n.name)+'</text>'+(n.sub?'<text x=\"'+n.x+'\" y=\"'+(n.y+43)+'\" text-anchor=\"middle\">'+esc(n.sub)+'</text>':'')+'</g>').join('');}function table(id,rows){$('#'+id).classList.add('table');$('#'+id).replaceChildren(...rows.map(r=>{const row=document.createElement('div');row.className='tr';row.style.gridTemplateColumns='repeat('+r.length+',minmax(0,1fr))';r.forEach(x=>{const c=document.createElement('div');if(typeof x==='object'){c.textContent=x.text;c.className=x.cls||'';}else c.textContent=x;row.append(c)});return row;}));}function band(id,parts,total){$('#'+id).replaceChildren(...parts.map((p,i)=>{const e=document.createElement('i');e.style.width=(100*p.value/total)+'%';e.className=i%2?'gold':'on';e.title=p.name+': '+p.value;return e}));}function intervals(id,items,total){$('#'+id).replaceChildren(...items.map(p=>{const row=document.createElement('div');row.className='interval';const bar=document.createElement('i');bar.style.left=(100*p.start/total)+'%';bar.style.width=(100*(p.end-p.start)/total)+'%';bar.className=p.cls||'';const label=document.createElement('span');label.textContent=p.label;row.append(bar,label);return row;}));}function plot(id,series,xr,yr,opts={}){const e=$('#'+id),X=x=>42+(x-xr[0])/(xr[1]-xr[0])*298,Y=y=>175-(y-yr[0])/(yr[1]-yr[0])*148;let out='<defs><clipPath id=\"clip-'+id+'\"><rect x=\"42\" y=\"27\" width=\"298\" height=\"148\"/></clipPath></defs>';for(let i=0;i<3;i++){let x=xr[0]+i*(xr[1]-xr[0])/2,y=yr[0]+i*(yr[1]-yr[0])/2;out+='<path class=\"gridline\" d=\"M '+X(x)+' 27V175M42 '+Y(y)+'H340\"/><text x=\"'+X(x)+'\" y=\"194\" text-anchor=\"middle\">'+esc(opts.xfmt?opts.xfmt(x):Number(x.toFixed(2)))+'</text><text x=\"36\" y=\"'+(Y(y)+4)+'\" text-anchor=\"end\">'+esc(opts.yfmt?opts.yfmt(y):Number(y.toFixed(2)))+'</text>';}out+='<g clip-path=\"url(#clip-'+id+')\">';for(const s of series){out+='<path fill=\"none\" stroke=\"'+(s.color||'var(--accent)')+'\" stroke-width=\"2.5\" '+(s.dash?'stroke-dasharray=\"5 4\"':'')+' d=\"'+s.pts.map((p,i)=>(i?'L':'M')+X(p[0]).toFixed(2)+','+Y(p[1]).toFixed(2)).join(' ')+'\"/>';}for(const p of opts.points||[])out+='<circle cx=\"'+X(p[0])+'\" cy=\"'+Y(p[1])+'\" r=\"4\" fill=\"var(--second)\" stroke=\"var(--paper)\" stroke-width=\"1.5\"/>';if(opts.cursor!==undefined)out+='<path d=\"M'+X(opts.cursor)+' 27V175\" stroke=\"var(--muted)\" stroke-dasharray=\"3 3\"/>';e.innerHTML=out+'</g>';e.dataset.curves=JSON.stringify(series.map(s=>s.pts));}const curve=(fn,lo,hi,n=120)=>Array.from({length:n+1},(_,i)=>{const x=lo+(hi-lo)*i/n;return[x,fn(x)]});function inputs(draw){document.querySelectorAll('input,select').forEach(e=>{e.addEventListener('input',draw);e.addEventListener('change',draw)});draw()}const n=id=>+$('#'+id).value;const show=(id,v,d=3)=>set(id,Number(v.toFixed(d)));function bars(id,values,max=1){$('#'+id).className='bars';$('#'+id).innerHTML=values.map(v=>'<div class=\"bar-row\"><span>'+esc(v.label)+'</span><div class=\"bar-track\"><i class=\"'+(v.cls||'')+'\" style=\"width:'+Math.max(0,Math.min(100,v.value/max*100))+'%\"></i></div><output>'+esc(v.text??(v.value*100).toFixed(1)+'%')+'</output></div>').join('')}function softmax(z,T=1){let max=Math.max(...z),w=z.map(x=>Math.exp((x-max)/T)),sum=w.reduce((a,b)=>a+b,0);return w.map(x=>x/sum)}let keys=new Set,history=[],known='none';function draw(){set('effects',keys.size);set('known',t(known));table('history',Array.from({length:4},(_,i)=>history[i]||['—','—','—']));$('#send').disabled=history.length>0;for(const id of['same','fresh'])$('#'+id).disabled=!history.length||history.length>=4;for(const id of['query','cancel'])$('#'+id).disabled=!history.length;set('explanation',t('explain'));document.body.dataset.effects=keys.size;document.body.dataset.known=known;}function send(key,lost=false){let duplicate=keys.has(key);keys.add(key);history.push(['call-'+(history.length+1),key,lost?'?':duplicate?'↩':'✓']);if(key==='op-A')known=lost?'unknown':'done';draw()}$('#send').onclick=()=>send('op-A',true);$('#same').onclick=()=>send('op-A');$('#fresh').onclick=()=>send('op-'+(history.length+1));$('#query').onclick=()=>{known=keys.has('op-A')?'done':'none';draw()};$('#cancel').onclick=draw;$('#reset').onclick=()=>{keys.clear();history=[];known='none';draw()};draw();","audio":false,"strings":{"cancel":"Cancel local wait","done":"Commit confirmed","effects":"Server writes","explain":"The first timeout follows a server commit. Reusing the key finds its deduplication record; a new key creates another operation. Canceling the wait cannot undo the commit.","fresh":"Retry new key","heading":"After timeout, how many writes?","history":"Call / operation key / result","intro":"Send a write whose response is lost, then retry or query.","known":"Known original status","none":"Not sent","query":"Query original operation","reset":"Reset","same":"Retry same key","scope":"An in-memory server deduplicates keys; each new key increments once, up to four calls. The first response is lost. Real systems also need retention and atomicity.","scopeLabel":"Model scope and assumptions","send":"Write, lose response","unknown":"Unknown outcome"}}
A Runnable Local Executor Example
The following simulates a read-only inventory query using preset model responses. It verifies tool whitelists, parameters, call result association, and request limits; it does not call the model API nor execute real business writes. Production systems also need persistent state, identity and permissions, timeouts, complete protocol adaptation, and auditing.
STOCK={"A1":3}defexecute(call):ifcall.get("name")!="lookup_stock":return{"ok":False,"error":"unknown_tool"}args=call.get("input")ifnotisinstance(args,dict)orset(args)!={"sku"}:return{"ok":False,"error":"invalid_arguments"}ifnotisinstance(args["sku"],str):return{"ok":False,"error":"invalid_sku"}ifargs["sku"]notinSTOCK:return{"ok":False,"error":"not_found"}return{"ok":True,"available":STOCK[args["sku"]]}defrun(model,max_requests=3):observations, seen=[], set()for_inrange(max_requests):response=model(list(observations))ifresponse["kind"]=="final":# Model end does not equal business acceptance success.
return{"state":"model_ended","text":response["text"],"observations":observations}ifresponse["kind"]!="tool":return{"state":"invalid_response"}call=response["call"]call_id=call.get("id")ifnotisinstance(call_id,str)ornotcall_idorcall_idinseen:return{"state":"invalid_call_id"}seen.add(call_id)observations.append({"call_id":call_id,"result":execute(call)})return{"state":"budget_exhausted","observations":observations}defscripted(observations):ifnotobservations:return{"kind":"tool","call":{"id":"c17","name":"lookup_stock","input":{"sku":"A1"}}}assertobservations[0]=={"call_id":"c17","result":{"ok":True,"available":3}}return{"kind":"final","text":"Current query found sellable quantity of 3, not yet reserved."}result=run(scripted)assertresult["state"]=="model_ended"assertexecute({"name":"delete_stock"})["error"]=="unknown_tool"assertexecute({"name":"lookup_stock","input":{"sku":1}})["error"]=="invalid_sku"assertrun(scripted,max_requests=1)["state"]=="budget_exhausted"print(result["text"])
The example exhausts the request budget after one tool call but before obtaining the final response, explicitly returning budget_exhausted. This is more accurate than mapping any loop exit to "success." Real applications should also persist completed calls to avoid losing side effect records after process restarts.
Tool Boundaries and Loop Convergence
Setting "run at most ten rounds" as the final insurance is not enough. One should observe whether each round gains new evidence, modifies artifacts, or eliminates errors; when repeatedly getting the same failure with no change in conditions, repeating actions usually cannot make progress. Instead of infinitely resending, return more specific errors, change the retrieval scope, or report missing conditions.
Tool outputs need to be size-limited and retain readable positions. When truncating logs, clearly mark the truncation interval, total amount, and file location, so the model does not mistakenly believe it has received complete evidence. For context management, see Context Engineering.
Specialized tools expose parameters and business semantics to the executor, facilitating validation, authorization, and recording. Shell tools provide more general capabilities, but the command allowlist alone is insufficient to limit everything the program can do; actual permissions still need to be controlled via execution identity, file and network access scope, isolated environments, and resource limits. Specific strategies should match already authorized tasks, and not let every reversible operation degrade into repeated confirmation. See Security and Protection for details.
Supplement: Environment Simulators and Real Execution
Agent training and evaluation can use environment simulators: the policy proposes actions, the simulator generates observations, and then the policy continues. Language world models attempt to learn these environmental responses; for example, the Qwen-AgentWorld Model Card describes simulation capabilities for multi-class interactive environments. This belongs to a different execution path from actually modifying files or calling business services.
Simulators may generate results that seem reasonable but are state-inconsistent. During evaluation, do not just look at whether a single observation looks real, but also at the causal consistency of continuous actions, resource constraints, and whether long-term state is maintained; high scores in a simulated environment cannot directly prove success rates in real tool execution. Suitability as a simulator or Agent should be measured separately; one cannot assert that a certain role is naturally feasible or infeasible based on the number of activated parameters.
Subsequent MCP and Skills discusses how capabilities are integrated and provided on demand, and Memory and State discusses what needs to be saved for long tasks and fault recovery.