---
title: Multi-agent orchestration and integration
url: https://doc.liz6.com/en/ai/03-agent-systems/06-multi-agent-orchestration
locale: en
area: ai
tags:
- Models & agents
- Agent Execution Systems
date: 2026-06-30
modified: 2026-09-10
description: Establish a measurable single-agent baseline before splitting into independent reasoning loops. Compute parallel limits from dependencies, define inputs, artifacts and commit ownership, and handle late results and recovery. Judge the change by complete-task quality, latency and cost.
---

# Multi-agent orchestration and integration

Establish a measurable single-agent baseline before splitting into independent reasoning loops. Compute parallel limits from dependencies, define inputs, artifacts and commit ownership, and handle late results and recovery. Judge the change by complete-task quality, latency and cost.

## From Workflows to Multi-Agent

A single Agent can also call independent tools in parallel; therefore, "needing to read multiple files simultaneously" does not necessarily require multiple reasoning loops. Multi-Agent is better suited for sub-tasks that require longer individual analysis, different information scopes, or independent workflows. Once each Worker forms its own context, returning compressed results and evidence locations to the Coordinator can reduce the amount of raw material the main context has to bear.

Multiple model calls do not automatically equal Multi-Agent. Fixed extraction, validation, and formatting pipelines can be implemented by standard workflows; dynamically decomposing tasks and allowing branches to proceed independently based on tool feedback is closer to the collaborative Agents discussed here. [Building effective agents](https://www.anthropic.com/engineering/building-effective-agents)

| Organization Method | Who Decides Steps | Main Benefits and Costs |
|---|---|---|
| Chaining and Routing | Pre-defined program paths or classification rules | Easy to reproduce, but requires process expansion for unknown branches |
| Parallel Sharding | Pre-defined independent scopes | Shortens overlapping work, still requires merging |
| Dynamic Coordinator and Workers | Model proposes decomposition, scheduler executes | Adapts to unknown tasks, coordination and verification are more complex |
| Generation and Review Loop | Fixed process plus model feedback | Improves candidates, constrained by review quality and iteration limits |

Multi-Agent dialogue frameworks demonstrate the ability to combine roles, tools, and communication patterns, but providing a dialogue mechanism does not automatically ensure task correctness. When deciding which structure to adopt, refer to the [AutoGen paper](https://arxiv.org/abs/2308.08155). First, identify the specific limitations of the current single-Agent approach, then verify whether adding collaboration improves it.

## Map Dependencies First, Then Calculate Benefits

Assume a task consists of Preparation A, Implementation B, Independent Compatibility Check C, and Summary D. Both B and C require A's interface contract; D must receive results from both. If B and C can indeed proceed independently, the following parallel structure holds:

<svg viewBox="0 0 760 371.061457157135" xmlns="http://www.w3.org/2000/svg" role="img" aria-label="Parallelism is determined by dependencies: Preparation, Independent Branches, Summary Verification" style="max-width:100%;height:auto" font-family="Source Han Sans CN,Microsoft YaHei,sans-serif">
<defs><marker id="agent-dependencies-arrow" markerWidth="8" markerHeight="8" refX="7" refY="4" orient="auto"><path d="M0,0 L8,4 L0,8 Z" fill="#64748b"></path></marker></defs>
<rect width="760" height="371.061457157135" rx="12" fill="#f8fafc"></rect>


<g transform="translate(0 11.06145715713501)"><rect x="25" y="130" width="145" height="75" rx="8" fill="#e0e7ff" stroke="#c7d2fe"></rect><text x="97.5" y="164.5" font-size="15" fill="#1e293b" text-anchor="middle" font-weight="600">Prepare A: 2 min</text><text x="97.5" y="186.5" font-size="12" fill="#475569" text-anchor="middle" font-weight="400">Fix goals and interfaces</text><rect x="270" y="65" width="205" height="75" rx="8" fill="#e0e7ff" stroke="#c7d2fe"></rect><text x="372.5" y="99.5" font-size="15" fill="#1e293b" text-anchor="middle" font-weight="600">Branch B: 4 min</text><text x="372.5" y="121.5" font-size="12" fill="#475569" text-anchor="middle" font-weight="400">Migration Implementation</text><line x1="170" y1="167" x2="265" y2="102" stroke="#64748b" stroke-width="1.8" marker-end="url(#agent-dependencies-arrow)"></line><line x1="475" y1="102" x2="560" y2="167" stroke="#64748b" stroke-width="1.8" marker-end="url(#agent-dependencies-arrow)"></line><rect x="270" y="205" width="205" height="75" rx="8" fill="#e0e7ff" stroke="#c7d2fe"></rect><text x="372.5" y="239.5" font-size="15" fill="#1e293b" text-anchor="middle" font-weight="600">Branch C: 6 min</text><text x="372.5" y="261.5" font-size="12" fill="#475569" text-anchor="middle" font-weight="400">Independent Compatibility Check</text><line x1="170" y1="167" x2="265" y2="242" stroke="#64748b" stroke-width="1.8" marker-end="url(#agent-dependencies-arrow)"></line><line x1="475" y1="242" x2="560" y2="167" stroke="#64748b" stroke-width="1.8" marker-end="url(#agent-dependencies-arrow)"></line><rect x="565" y="130" width="170" height="75" rx="8" fill="#e0e7ff" stroke="#c7d2fe"></rect><text x="650.0" y="164.5" font-size="15" fill="#1e293b" text-anchor="middle" font-weight="600">Summary D: 3 min</text><text x="650.0" y="186.5" font-size="12" fill="#475569" text-anchor="middle" font-weight="400">Verify and Integrate</text></g><text x="24" y="29" font-size="19" fill="#0f172a" text-anchor="start" font-weight="700"><tspan x="24" dy="0">Parallelism is determined by dependencies: Preparation, Independent </tspan><tspan x="24" dy="25.650000000000002">Branches, Summary Verification</tspan></text><text x="24" y="324.061457157135" font-size="13" fill="#475569" text-anchor="start" font-weight="400"><tspan x="24" dy="0">Ideal total duration: 2 + max(4, 6) + 3 = 11 minutes; does not yet include delegation, queuing, and rework.</tspan></text>
</svg>

Serial time is 2+4+6+3=15 minutes; ideal parallel time is 2+max(4,6)+3=11 minutes, saving 4 minutes. If delegation, queuing, and merging conflicts take an additional 5 minutes, the total time becomes 16 minutes. Launching two Workers does not mean the task is twice as fast.

If C checks patches that B has not yet produced, C depends on B, and this diagram needs to be changed to serial; alternatively, C can first independently organize the check standards, and then verify after the patches are completed. **Changing the task definition can create some parallel space, but you cannot use model guesses to eliminate real data dependencies.**

```python
duration = {"A": 2, "B": 4, "C": 6, "D": 3}
deps = {"A": [], "B": ["A"], "C": ["A"], "D": ["B", "C"]}
finish = {}
# This dictionary lists nodes in topological order of dependencies, not a general scheduler.
for node, minutes in duration.items():
    finish[node] = max((finish[x] for x in deps[node]), default=0) + minutes
assert finish["D"] == 11
assert sum(duration.values()) == 15
print("Ideal parallel minutes:", finish["D"])
print("Plus 5 minutes coordination overhead:", finish["D"] + 5)
```

Costs must also include the main Agent, all Workers, repeated reads, tool calls, and failed attempts. Using cheaper models for suitable sub-tasks may save money, but multiple independent windows will also redundantly load constraints and materials; whether prefix caching hits depends on the service, model, and input. It is not guaranteed that "opening sub-Agents is always cheaper than switching models in the main loop."

Engineering reports from actual Multi-Agent research systems discuss context isolation, independent research, and higher token consumption, while also pointing out coordination limitations for tightly dependent tasks. The benefits of such systems should be understood as specific cases, not as a universal acceleration factor for all tasks. [Anthropic Multi-Agent Research System](https://www.anthropic.com/engineering/multi-agent-research-system)

### Why does parallelism not scale by headcount?

Draw dependencies before estimating parallel benefits. Dependencies determine start times and merge consumes time too. Total work is 15 while the default critical path is 9: resource use and completion latency are different quantities.

$$
T_{\mathrm{parallel}}=\max_{p\in\mathrm{paths}}\sum_{i\in p}t_i+T_{\mathrm{coordination}}
$$

**Why does parallelism not scale by headcount?**

Parallel latency follows the longest dependency path plus coordination overhead, not total work divided by workers.


## Delegation is Delivering a Task Contract

"You are responsible for checking compatibility" is too vague: which version to check, which interfaces, whether modifications are allowed, and what evidence is required are all undefined. The Coordinator should deliver complete information necessary for the Worker to complete the task, while avoiding dumping the entire main session history.

| Contract Item | Example |
|---|---|
| Goal | Check if the migration plan breaks existing import formats |
| Inputs and Versions | schema-v3, sample set R12, plan draft-4 |
| Scope | Only analyze format compatibility, do not modify implementation |
| Available Capabilities and Permissions | Read samples, run local validation; do not publish artifacts |
| Output | Compatibility matrix, failing samples, evidence locations, uncovered items |
| Stop Condition | Complete the matrix or reach the allocated budget, clearly stating remaining gaps |
| Merge Interface | Write results to a task-specific directory and return identifiers |

<svg viewBox="0 0 760 340" xmlns="http://www.w3.org/2000/svg" role="img" aria-label="Sub-task results require evidence and versions; the Coordinator is responsible for accepting or rejecting" style="max-width:100%;height:auto" font-family="Source Han Sans CN,Microsoft YaHei,sans-serif">
<defs><marker id="agent-result-contract-arrow" markerWidth="8" markerHeight="8" refX="7" refY="4" orient="auto"><path d="M0,0 L8,4 L0,8 Z" fill="#64748b"></path></marker></defs>
<rect width="760" height="340" rx="12" fill="#f8fafc"></rect>


<g transform="translate(0 0)"><rect x="25" y="95" width="185" height="100" rx="8" fill="#e0e7ff" stroke="#c7d2fe"></rect><text x="117.5" y="142.0" font-size="15" fill="#1e293b" text-anchor="middle" font-weight="600">Task Contract</text><text x="117.5" y="164.0" font-size="12" fill="#475569" text-anchor="middle" font-weight="400">Goal, input, authority, output</text><rect x="275" y="95" width="210" height="100" rx="8" fill="#e0e7ff" stroke="#c7d2fe"></rect><text x="380.0" y="142.0" font-size="15" fill="#1e293b" text-anchor="middle" font-weight="600">Worker Artifact</text><text x="380.0" y="164.0" font-size="12" fill="#475569" text-anchor="middle" font-weight="400">Findings, changes, evidence, gaps</text><rect x="550" y="95" width="185" height="100" rx="8" fill="#e0e7ff" stroke="#c7d2fe"></rect><text x="642.5" y="142.0" font-size="15" fill="#1e293b" text-anchor="middle" font-weight="600">Coordinator Acceptance</text><text x="642.5" y="164.0" font-size="12" fill="#475569" text-anchor="middle" font-weight="400">Version, coverage, validation</text><line x1="210" y1="145" x2="270" y2="145" stroke="#64748b" stroke-width="1.8" marker-end="url(#agent-result-contract-arrow)"></line><line x1="485" y1="145" x2="545" y2="145" stroke="#64748b" stroke-width="1.8" marker-end="url(#agent-result-contract-arrow)"></line><line x1="640" y1="201" x2="640" y2="263" stroke="#64748b" stroke-width="1.8" marker-end="url(#agent-result-contract-arrow)"></line><line x1="640" y1="263" x2="380" y2="263" stroke="#64748b" stroke-width="1.8" marker-end="url(#agent-result-contract-arrow)"></line><line x1="380" y1="263" x2="380" y2="201" stroke="#64748b" stroke-width="1.8" marker-end="url(#agent-result-contract-arrow)"></line></g><text x="24" y="29" font-size="19" fill="#0f172a" text-anchor="start" font-weight="700"><tspan x="24" dy="0">Sub-task results require evidence and versions; the Coordinator is </tspan><tspan x="24" dy="25.650000000000002">responsible for accepting or rejecting</tspan></text><text x="24" y="297" font-size="14" fill="#334155" text-anchor="start" font-weight="400"><tspan x="24" dy="0">Provide specific gaps when rejecting; receiving a result does not mean accepting it</tspan></text>
</svg>

Whether tools, history, and file systems are shared depends on the runtime environment. Even with shared directories, Workers may not have read the files; even without shared history, necessary context can be restored through explicit inputs. Delegation messages should specify actually accessible resources, rather than assuming the other party "knows what was discussed earlier."

If the main task requires "only change Chinese first," this constraint must be passed to all Workers that might modify documents. The Main Agent's authorization is also not a reason to arbitrarily expand tool permissions: a read-only check task should receive the capabilities it needs, and out-of-scope findings should be returned to the Coordinator for decision.

## Artifact Ownership and Result Acceptance

### Define Boundaries for Writes

In a shared workspace, two Agents modifying the same file simultaneously may overwrite content; assigning different files does not necessarily mean independence, as they might jointly change an interface contract. You can choose to divide work by directory or artifact, use independent branches or working copies, and then have a clear Integrator merge them.

Isolated copies reduce direct overwrites but cannot avoid semantic conflicts. For example, if B renames a field to `item_id`, C's test might still assume the field is called `sku`; even if the patch texts can be automatically merged, integration will fail. Fixing interface versions in the contract and propagating changes to dependent parties first can reduce such wasted effort.

### Summaries for Location, Evidence for Confirmation

A Worker returning "check passed" is insufficient. Results should include the input version targeted, the method used, accessible artifacts, and uncovered items. The Coordinator first checks goal coverage, then verifies versions and evidence, and finally decides to accept, supplement, or re-divide.

| Return Situation | Coordinator Action |
|---|---|
| All conclusions have corresponding samples and check results | Accept and incorporate into overall evidence |
| Conclusions are correct but only cover part of the format | Mark as partially complete, assign specific gaps |
| Two Workers have conflicting conclusions | Return to original conditions and versions, do not directly vote by majority |
| Referenced files exist but have been modified subsequently | Check hash or commit, re-verify related conclusions |
| Only fluent summaries, no verifiable artifacts | Require additional evidence based on task risk |

Independent review helps identify omissions, but another Agent using the same model might share the same bias. "Separating author and reviewer" provides role and context differences, but does not equal statistical independence or proof of correctness. Combining explicit standards, independent testing, actual execution, and necessary human review creates verifiable boundaries.

### Why reject a successful worker result?

Archive all attempts, but integrate only results satisfying the current contract. Keep worker completion, result receipt and accepted integration as separate states, rather than overloading one success field.

**Why reject a successful worker result?**

Success is not current validity; integration must check task identity, current attempt, input version and evidence.


## Orchestration Requires Recovery Semantics

The Coordinator should save the task ID, attempt number, input version, status, and artifact location for each sub-task. Status should distinguish at least between queued, running, completed, accepted, failed, cancelled: Worker completion only means a result was produced; accepted means the Coordinator has confirmed it is available for overall delivery.

Assume Attempt 1 times out, and the Coordinator launches Attempt 2; later, the stale result from Attempt 1 arrives. If you only accept based on sub-task name, the old content will overwrite the new one. You need to verify late results against attempt numbers and input versions, and if necessary, keep them as references without automatically publishing them.

Cancellation does not guarantee that old Workers stop writing immediately. In shared resource scenarios, control commit permissions, for example, by having the Integrator verify the current valid attempt or version before writing to the authoritative artifact; otherwise, a Worker displayed as cancelled in the interface might still overwrite files in the background. Cross-executor state rules echo the conditional updates in [Memory and State](/ai/03-agent-systems/04-memory-and-recovery.md).

| Failure | Confirm First During Recovery |
|---|---|
| Worker disconnected | Are there completed artifacts, and is execution still possible? |
| Coordinator restart | Which tasks are accepted, and which results are unknown? |
| Sub-task restarted | Will two attempts write to the same resource? |
| Failure during merge | Which parts are persisted, and can they be recovered or rolled back? |
| Lost response to external side effects | Was the business operation actually completed? Do not just look at thread status. |

## How to Judge if Multi-Agent Actually Provides Benefits

Compare single-Agent, single-Agent parallel tools, and Multi-Agent using the same task set, keeping tool permissions, input evidence, and acceptance criteria as consistent as possible. In addition to task success rate and total latency, record the Main Agent's coordination time, proportion of duplicate work, number of conflicts, failure retries, and total costs.

If Workers are fast but the Main Agent spends a long time clarifying and merging, the bottleneck is task decomposition or result contracts; if two Workers always query the same materials, the scopes might overlap; if integration always fails after merging, check shared interfaces and integration tests. If single-Agent batch tools achieve the same effect, there is no need to increase maintenance costs just for the sake of having more roles.

More detailed metrics and per-task tracking are expanded in [Evaluation and Observability](/ai/04-evaluation-and-production/01-evaluation-and-observability.md). The unit of evaluation should be the user's complete task, not just counting how many Agents were started or how many messages were generated.

Continue with：[Evaluation, trials and observability](/ai/04-evaluation-and-production/01-evaluation-and-observability)。
