Scaling Claude Code Agents with Sub-Agent Delegation

Scaling Claude Code agents without scaling the bill: how sub-agent delegation works, which work belongs on a cheaper model, and what it costs to get wrong.

Scaling Claude Code Agents with Sub-Agent Delegation | Kaxo

TL;DR: Scaling Claude Code agents is a cost problem before it is an architecture problem. Sub-agent delegation is how a setup grows without its bill growing with it. A parent agent keeps the work that needs judgement; bounded, repetitive work goes to sub-agents on a cheaper model tier. We run our own agent fleet this way, and tiering cut the cost of routine work by roughly an order of magnitude with no loss of quality.


What Is Sub-Agent Architecture? (Quick Answer)

Sub-agent architecture means a main Claude Code agent delegates specific tasks to other Claude Code agents, each with its own context window and its own model. The main agent holds the thread of the work. The sub-agents do pieces of it and report back.

Two things make it worth the trouble. The obvious one is cost: a task that needs no judgement does not need an expensive model, and on a busy fleet that describes most tasks. The less obvious one is context. A sub-agent’s work happens in its own window, so the output of a long file scan never lands in the main agent’s context at all. The main agent gets the answer instead of the evidence, and it can keep working on the actual problem for much longer.

It suits work that: recurs, can be described precisely enough that the result is checkable, and needs no creative judgement.


Contents


Why Scaling Claude Code Agents Gets Expensive

A single agent doing everything is the natural way to start and it stops working for two reasons at once.

The first is cost. An agent that handles strategy also ends up fetching data, parsing files and generating the same report every week, all on whatever model you chose for the hardest thing it does. You are paying a premium rate for work that has no judgement in it.

The second is context. Every file an agent reads and every command it runs leaves its output in the context window. Routine work produces most of that volume while contributing least to the decision at hand, and once the window fills, the agent starts losing the constraints it was given at the beginning. The symptom looks like the model getting worse. The cause is that the useful part of the conversation has been crowded out by the evidence.

Both problems have the same shape: undifferentiated work on an undifferentiated agent.

The Solution in 30 Seconds

Separate the work that needs judgement from the work that only needs doing.

The parent agent keeps the first kind: deciding what to do, synthesising results, making the call. Bounded tasks go to sub-agents, which run on a cheaper model, work in their own context, and hand back a result rather than a transcript.

No framework is involved. This is Claude Code’s own agent invocation, driven by a clear description of the task.

What Are Sub-Agents?

A sub-agent is a Claude Code agent invoked for one task and then finished with. It has no memory of previous invocations and no stake in the wider project. It receives a task, does it, and returns a result.

That disposability is the whole point. Because a sub-agent carries no state and its job is narrow, you can describe what a correct result looks like, which means you can run it on a cheaper model and still know whether it worked. An agent that persists and makes judgement calls cannot be checked that cheaply, which is why it stays on a capable model.

The Economics of Sub-Agents

Tiering saves roughly the proportion of your work that is routine, and the arithmetic is simple enough to do in your head. Model tiers differ in price by a large multiple. If most of the tasks on your fleet are routine, and on a working fleet most of them are, then moving that majority down a tier changes the total bill by far more than any prompt optimisation will.

What that is worth in practice: on our own fleet, tiering routine work cut the cost of that work by roughly an order of magnitude, and the quality of the results did not drop. We are not quoting an industry benchmark here. It is what we measured on the system we run ourselves.

The more interesting consequence is what it does to your decisions. Once a narrow, specialised agent costs very little to run, you stop rationing them. You can build one for a task that only comes up weekly, because the incremental cost of having it is close to nothing. The constraint moves from “can we afford another agent” to “is this task actually well enough defined to delegate”, which is a much better question to be arguing about.

What this does not fix: a premium model doing premium work still costs what it costs. Tiering makes the routine cheap. It does not make the hard thinking cheap, and a system whose cost is dominated by genuine synthesis will not see the same effect.

Can Claude Code Agents Call Other Agents?

Yes, and without any orchestration framework.

The parent agent describes a task clearly enough that its result can be checked, invokes a sub-agent to do it, and carries on with the result. The sub-agent executes in its own context and returns structured output. That is the entire mechanism. No LangChain, no CrewAI, no message bus.

The hard part is not the wiring. It is writing a task description precise enough that you would notice a wrong answer, which is a writing problem rather than an engineering one.

What Belongs on a Cheaper Model

The judgement is about the task, not the model. Ask whether a correct result is distinguishable from an incorrect one without redoing the work yourself.

It usually is for fetching structured data from an API, parsing or transforming files, producing a report in a fixed format, and scanning for a pattern across many files. These have a right answer you can recognise on sight.

It usually is not for synthesis across sources, anything depending on context from earlier in a session, design decisions, and one-off work where defining the task costs more than doing it.

Between those lies a band of tasks needing moderate judgement, and those belong on a mid-tier model. If you are still deciding whether to build this yourself at all, the build-vs-buy comparison for agent tooling covers that choice. The error that costs real money is not misjudging one task. It is never revisiting the split, so everything stays on a premium model by default because that was the safe choice on day one.

Common Problems Solved by Sub-Agent Architecture

“My Claude Code costs are too high.” Almost always this means one capable model is doing everything, including the work with no judgement in it. The fix is tiering, and the saving is roughly proportional to how much of your volume is routine, which is usually most of it.

“I hit rate limits constantly.” A single agent queues all its work through one place. Independent tasks dispatched to separate sub-agents proceed in parallel rather than waiting behind each other, which both relieves the limit and finishes sooner.

“My agent’s context gets bloated.” This is the problem sub-agents solve best and it is the one people notice last. A task whose output runs to hundreds of lines belongs in another context window. Delegate it and the parent receives the conclusion rather than the raw material, which is what keeps a long session coherent.

Where Sub-Agent Delegation Goes Wrong

Three ways: an underspecified task, the overhead of delegating at all, and a cheap model that is only mostly right. Worth knowing before you commit to the pattern.

A vague task specification is worse than no delegation. A sub-agent given an underspecified task returns something plausible, and plausible-but-wrong is expensive precisely because it does not announce itself. If you cannot say what a correct result looks like, the task is not ready to delegate.

Delegation has overhead. Describing a task, invoking an agent and checking what comes back is not free. For anything genuinely trivial, doing it yourself is cheaper than arranging for it to be done.

A cheap wrong answer costs more than an expensive right one. The tiering decision is about reliability first and price second. If a cheaper model does the task correctly, the saving is real. If it does it correctly most of the time, you have bought yourself an intermittent fault, and those cost more to find than the model ever saved.

Key Takeaways

  • Separate judgement from execution first. The split is what makes delegation worth anything; the cost saving follows from it.
  • Tiering is the main cost lever on a fleet whose volume is mostly routine work, and on our own fleet it reduced the cost of that work by roughly an order of magnitude.
  • Context isolation matters as much as cost. Keeping bulky output out of the parent’s window is what lets a long session stay coherent.
  • Delegate only what you can check. If you cannot recognise a wrong result, a cheaper model is not saving you anything.
  • Cheap agents change what you build. When a narrow agent costs almost nothing to run, specialising becomes the obvious move rather than a luxury.
  • Revisit the split. The expensive failure is leaving everything on a premium model because nobody went back to look.

Working With Kaxo on This

Ready to scale your AI agent infrastructure without scaling costs?

Kaxo Technologies builds production agent systems for Canadian SMBs. We run our own agent fleet on the patterns described here, and we measure what it costs us.

Our services:

  • Agent architecture consulting
  • Sub-agent fleet implementation and optimization
  • Cost optimization audits for existing agent systems
  • Claude Code training and best practices workshops

Service Areas: Kawartha Lakes | Peterborough | Durham Region

Expertise: AI automation, agentic workflows, Claude Code infrastructure


Need someone to build agents like this?

We design, build, and deploy custom AI agents on your infrastructure. Production-grade reliability, full code ownership, no vendor lock-in. See our AI Agent Development service for the operational details, or book a discovery call.


Soli Deo Gloria

Frequently Asked Questions

What are sub-agents in Claude Code?

Sub-agents are specialised Claude Code agents that a parent agent invokes for a single bounded task. The parent keeps the work that needs judgement on a capable model, and hands routine, well-specified work to an agent running a cheaper one. The point is not that sub-agents are smarter. It is that most of what an agent does all day does not need the expensive model.

What's the difference between Claude Code agents and sub-agents?

An agent is project-level. It persists, it holds state between sessions, and it owns a scope. A sub-agent is task-level and ephemeral: it is handed a task, it returns a result, and it keeps nothing. That difference is why a sub-agent can run on a cheap model without much risk, and why an agent usually cannot.

How can I reduce Claude Code API costs?

Move routine work off your most capable model. On a real fleet the large majority of tasks are data fetching, parsing and structured reporting, and those are the ones that do not need judgement. Routing that work to a cheaper tier is where the saving comes from, and on our own fleet it has been close to an order of magnitude. Prompt and context discipline matters too, but tiering is the bigger lever.

Can Claude Code agents call other agents?

Yes, and it needs no framework. A parent agent delegates by describing the task clearly enough that the result can be checked, and the sub-agent returns a structured result. No LangChain, no CrewAI, no orchestration layer.

How do I scale Claude Code from one agent to many?

Not by adding agents. Start by separating the work that needs judgement from the work that only needs doing, because that split is what makes a second agent worth having. Shared context belongs in one place every agent can read, so a new agent inherits conventions instead of restating them. Capabilities that more than a couple of agents need should be built once, not per agent.

When should I use a sub-agent vs doing it myself?

Delegate when the task recurs, the inputs and outputs can be stated plainly, and you can tell from the result whether it worked. Keep it in the parent when the task depends on context from earlier work, when it needs a judgement call, or when it is a one-off. The question a longer checklist reduces to is simply whether you could tell a correct result from an incorrect one without redoing the work.

What's the cheapest way to run Claude Code agents?

Put bounded work on the cheapest model that completes it reliably, keep synthesis and design on a capable one, and review the split periodically rather than once. The expensive mistake is not picking the wrong tier for a task. It is leaving every task on a premium model because nobody went back to look.

About the Author

The Kaxo Team leads AI infrastructure development and autonomous agent deployment for Canadian businesses. Specializes in self-hosted AI security, multi-agent orchestration, and production automation systems. Based in Ontario, Canada.

Written by
The Kaxo Team
Last Updated: October 3, 2026
Back to Insights