- AI
- AI agents
- Claude
- Architecture
- Anthropic
AI subagents: delegating between models without losing quality
Djtal explains how to split an AI agent's work across several models: the orchestrator and subagents pattern, what Claude 5.5 changes and three key settings.

In Djtal’s view, an AI agent can hand each subtask to a different model without losing quality, provided the orchestrator checks every report before it signs off the deliverable. The orchestrator, running on a more capable model, splits the work and writes the briefs, then checks and assembles the reports. Subagents, running on smaller models, handle tightly scoped tasks. With Claude Haiku 5.5, released on 7 October 2026, the bottom tier of the range now costs a hundredth of the top model’s output price. Its score on real-world professional tasks has more than doubled. Three settings, detailed below, decide whether the set-up pays for itself.
How does delegation between models work?
Anthropic describes this pattern as orchestrator-workers. A central model splits the task, delegates it to other models and combines their results. A simpler variant, routing, classifies each request and sends it to the right model: everyday questions go to a small model, hard cases to a more capable one.
Every exchange follows the same round trip.
- The brief is passed down. Drawing on its own research system, Anthropic lists what a brief must contain at minimum: an objective, an output format, the tools and sources to use, and clear task boundaries. With a vague brief, subagents repeat the same searches and leave gaps.
- The subagent starts from a fresh context. It knows nothing of the orchestrator’s conversation and receives only what the brief passes on.
- The report comes back, condensed. The orchestrator receives the essentials and is the only one to hold the full context.
- The orchestrator checks. It assesses each report, sends back a targeted correction when a result deviates from the brief, then writes the synthesis.
How do you keep quality up?
Quality rests on two structural safeguards. On the way out, the brief sets the scope: objective, format, sources, limits. On the way back, checking comes before signing. The model that signs off the deliverable re-reads what the subagents produce. Anthropic calls this loop evaluator-optimizer, in which one model produces and another evaluates. This structure, above all, sets the level of quality.
The delegated task must stay within the capabilities of the model that receives it. Anthropic positions Haiku 5.5 for narrowly scoped work such as summaries, context compaction, classification, database queries and subagent tasks. The company says it ‘pairs well with Opus 5.5 and Sonnet 5.5 as a subagent on coding work’. The limit is stated too. On complex agentic work, Sonnet 5.5 and Opus 5.5 remain better. In autonomous terminal work, Haiku 5.5 scores 39.2% where Sonnet 5.5 reaches 70.6%. Before delegating to the small model, then, check that the task is one it can handle.
The combination can beat the single model. When Anthropic measured its research system in June 2025, an Opus 4 orchestrator directing Sonnet 4 subagents outperformed Opus 4 working alone by 90.2% on its internal evaluation. The gain came from working in parallel on questions that split well.
The structure supports governance too. Each subtask is documented by its brief and its report, so a manager can later trace who produced what, on which model and from which instructions. Auditing delegated work then means reading these documents step by step.
What does the Claude 5.5 range change?
| Model | Role | Input | Output | Cache read |
|---|---|---|---|---|
| Fable 5.1 | orchestration, synthesis | $10 | $50 | $0.25 |
| Opus 5.5 | expert review, judgement | $4 | $20 | $0.20 |
| Sonnet 5.5 | high-volume production | $2 | $10 | $0.10 |
| Haiku 5.5 | checks, extraction, summaries | $0.10 | $0.50 | $0.01 |
Anthropic API prices per million tokens as at 8 October 2026. Fable 5.1 sits above the 5.5 range. Haiku 5.5 prices apply to prompts of up to 100,000 tokens and are five times higher beyond that.
The bottom tier has moved up a class. On GDPval-AA, a test of real-world professional tasks scored in Elo points, Haiku 5.5 reaches 1,620, compared with 735 for Haiku 4.5 and 1,840 for Sonnet 5.5. Its price falls by 90% for prompts of up to 100,000 tokens. It is also the first Haiku with an adjustable effort setting, so the same model answers fast on a sorting job and reasons for longer on a check.
The middle tier has caught up with the top. Sonnet 5.5, released on 28 September, finishes within two points of Opus 5.5 on GDPval-AA and beats it in terminal work (70.6% to 66.4%) at half the price. That lets you move high-volume production down a tier.
The orchestrator re-reads more cheaply. An orchestrator re-reads its whole context on every request, and that line soon dominates the bill of a long-running agent. A cache read costs 2.5% of the input price on Fable 5.1 and 5% on Opus 5.5 and Sonnet 5.5, compared with 10% on other models, Haiku 5.5 included.
Two models compete for the orchestrator seat. Anthropic presents Fable 5.1 as its most capable generally available model and describes it in this very role: planning the work, calling tools and recovering when a step fails. The nine benchmarks published with the Opus 5.5 announcement all favour Fable 5.1, while Opus 5.5 is priced 2.5 times lower. Anthropic itself notes that in practice the gap is narrower than those scores suggest. The mission decides. Fable 5.1 gets the seat when the mission runs for hours and the model has to recover from failures on its own. Opus 5.5 takes it when budget comes first.
Which settings make or break the set-up?
Declare the model for each subagent. In Claude Code, a subagent with no declared model inherits the main conversation’s model. A fork, which copies the current session, always inherits the session’s model. An orchestrator set to the top of the range therefore pulls all its subagents up to the same price. Set the model parameter at every launch.
Keep the bottom tier’s context short. Haiku 5.5 costs five times more beyond 100,000 prompt tokens. The brief should pass the subagent only the extract it needs. A subagent’s context window, incidentally, depends on its own model.
Measure before you scale up. According to Anthropic, a multi-agent system uses about 15 times more tokens than a chat. Token volume alone explains 80% of the performance differences on its research benchmark. A pilot batch, with consumption recorded per model, shows whether the split pays off before you roll it out.
When should you keep a single model?
In 2025, Anthropic judged delegation poorly suited to tasks where all the agents must share the same context or depend heavily on one another. Most development tasks fell into that category, since they are harder to split than research. With Haiku 5.5, Anthropic now also recommends delegation for code, on tightly scoped subtasks. The criterion has not changed: delegation pays off on work that splits well, such as page-by-page translations, parallel searches, serial checks or extraction across batches of documents.
Frequently asked questions
What is an AI subagent? An AI subagent is an agent that another agent launches for a specific task. It works in its own context window, with its own tools and model, then reports back to the agent that launched it.
Which Claude model should I choose for a subagent? Haiku 5.5 for narrow, repeated tasks (sorting, extraction, summarising, checking), Sonnet 5.5 for high-volume production, Opus 5.5 for expert review and cases that call for judgement. The orchestrator runs on a top-of-the-range model: Fable 5.1 for long, autonomous missions, Opus 5.5 when budget comes first.
Is delegation always cheaper? The price per token falls, but the volume rises: a multi-agent system uses about 15 times more tokens than a chat, which is why the saving should be measured on a pilot batch.
Sources
- Anthropic, Introducing Claude Haiku 5.5, 7 October 2026.
- Anthropic, Pricing, accessed 8 October 2026.
- Anthropic, Introducing Claude Sonnet 5.5, 28 September 2026.
- Anthropic, Introducing Claude Opus 5.5, 22 September 2026.
- Anthropic, Claude Fable 5.1, accessed 8 October 2026.
- Anthropic, How we built our multi-agent research system, 13 June 2025.
- Anthropic, Building effective agents, 19 December 2024.
- Claude Code, Subagents, accessed 8 October 2026.
A topic worth exploring for your business?
Let's spend 30 minutes on your context and find where AI can help your business.
Speak to one of our experts