Most teams hit the same wall with their first production AI agent. It works well on the job it was built for, so the business asks it to do more: answer billing questions, take returns, qualify sales leads. Each new responsibility makes the prompt longer, the tool list bigger, and the behaviour less predictable, until the agent is mediocre at everything it touches.
A multi-agent system is the standard way out. Instead of one agent stretched across every task, you build several narrow specialists and a coordinator that decides which one handles each request, runs them in sequence or in parallel, and assembles the result. Done well, each piece stays simple enough to test. Done badly, you've swapped one confused agent for five, plus a routing layer nobody fully understands.
This guide covers the coordinator-and-specialist architecture, how to decide between multi-agent and single-agent designs, how routing works, sequential versus parallel execution, managing shared state, error handling and fallbacks, why LangGraph fits this pattern, how to test the whole system, and when the extra complexity simply isn't worth it.
Why Single Agents Have Limits
A well-scoped AI agent does one thing reliably: qualifies leads, handles customer support, processes documents. The narrower the scope, the more reliable the behaviour. We've watched that pattern hold up across every project we've shipped.
The limit shows up when a business process needs multiple capabilities working together. A customer contacts your company. The message might be a support issue, a sales opportunity, a billing question, or a product return — and the right response in each case is completely different. A single agent trying to handle all four ends up either unreliably broad, or so tightly constrained that it fails most of the time.
Multi-agent systems are how you get out of that bind. Instead of one agent trying to do everything, you have multiple specialised agents each doing one thing well, with a coordinating layer that routes work to the right place and assembles the results. The trade-off is more moving parts to maintain, which we'll get into.
The Core Architecture: Coordinator and Specialists
Every multi-agent system has two types of agents:
The coordinator (orchestrator). Receives incoming requests, figures out what type of task it is, dispatches to the appropriate specialist, and assembles the result. The coordinator doesn't do the work — it manages the workflow.
Specialist agents. Each one is optimised for a specific task: customer support, lead qualification, order processing, document retrieval. They receive structured inputs from the coordinator, do their task reliably, and return structured outputs.
This separation produces a system where:
- Each specialist can be optimised, tested, and improved independently
- New capabilities can be added by building a new specialist and updating the routing
- Failures in one specialist don't cascade to others
- The system scales by adding more specialists or running them in parallel
Benefits of Multi-Agent Systems
When the workflow genuinely needs it, splitting work across specialists pays off in ways that are hard to get from one large agent. Most of the gains come from keeping each piece small.
Shorter prompts and more predictable behaviour
Each specialist only needs the instructions, examples, and tools for its own task. A billing agent doesn't carry return-policy rules; a returns agent doesn't see payment tools. Smaller prompts leave less room for instructions to conflict, so behaviour stays closer to what was tested. When something goes wrong, the set of possible causes is much smaller too.
Independent testing and improvement
Because specialists take structured inputs and return structured outputs, you can build a test set for each one and improve it without touching the others. A change to the support agent's prompt only needs the support tests to pass, plus a routing check. Teams can work on different specialists in parallel without stepping on each other.
Contained failures
A fault in one specialist, such as a broken integration or a bad prompt change, affects one category of request. The coordinator can fall back, reroute, or trip a circuit breaker while the rest of the system keeps working. With a single agent, the same fault can degrade every type of request at once.
Tighter tool access
Different tasks need different permissions. A specialist that only answers questions can have read-only access, while the one processing refunds gets write access to the billing system and nothing else. Scoping tools per specialist reduces the damage a confused or manipulated agent can do, and makes access easier to audit.
Lower latency for independent work
When tasks don't depend on each other, the coordinator can run specialists in parallel and wait for all of them before synthesising the result. Searching several data sources at once is noticeably faster than doing it one step at a time, which matters for user-facing workflows where people are waiting on an answer.
Room to grow without a rebuild
New capabilities arrive as new specialists plus a routing update. The coordinator, state object, and existing specialists stay as they are. That makes it practical to add a new request type months after launch without rewriting what already works.
When to Build Multi-Agent vs Single Agent
Use a single agent when:
- The workflow is genuinely one type of task (all customer support, all lead qualification)
- The volume is manageable by one agent
- The edge cases are predictable and few
Use a multi-agent system when:
- A single entry point needs to handle genuinely different types of requests that require different capabilities
- A complex workflow has sequential steps where different expertise is needed at each step
- Different parts of the workflow have different reliability requirements or tool access needs
- You need parallel processing — multiple tasks happening simultaneously rather than sequentially
If you can solve it with a single agent and tighter prompts, do that first. Multi-agent is more powerful and more expensive to maintain — both true.
| Dimension | Single Agent | Multi-Agent System |
|---|---|---|
| Best for | One task type at high volume (e.g. all customer support) | Mixed request types needing different capabilities at a single entry point |
| Build time | 2–4 weeks for a production-ready agent | 4–12 weeks depending on number of specialists |
| Ongoing maintenance | Low — one codebase, one set of prompts to tune | Higher — each specialist needs independent monitoring and updates |
| Routing complexity | None | Significant — misrouting degrades the entire system |
| Failure blast radius | Contained — one agent fails, one workflow stops | Wider — coordinator or routing bugs affect all downstream specialists |
| Parallel execution | Not native — sequential tool calls only | Native in LangGraph — multiple specialists can run simultaneously |
| Scalability ceiling | Hits limits when request types genuinely diverge | Scales by adding specialists without rebuilding the core |
| When to choose | Default choice until a single agent clearly fails | When single-agent approaches have hit a ceiling and the workflow truly requires it |
Multi-Agent System Use Cases
The architecture earns its complexity in a fairly small set of situations. In each of them, the deciding factor is the same: the work splits into parts that need different knowledge, different tools, or a different order of execution. These are the patterns that come up most often in production work.
A shared customer inbox
A single inbox or chat entry point receives support issues, billing questions, sales enquiries, and return requests. One agent handling all four becomes unreliable, because each needs different knowledge and tools. A coordinator classifies each message and routes it to the matching specialist, with low-confidence messages going to a generalist or a human. Customers get a response from an agent built for their actual question, and each specialist can be tuned on its own traffic.
Document processing pipelines
Incoming documents need data extracted, checked against business rules, routed to a department, and followed by notifications. Each step needs different logic, and step two can't start until step one finishes. A sequential pipeline of extraction, validation, routing, and notification agents passes a shared state object from one to the next. Failures are caught at the step where they occur, which makes bad documents easier to trace and fix than in one monolithic agent.
Research and synthesis
Answering a research question often means checking recent news, internal records, and the knowledge base. Doing those lookups one after another is slow. Running a web search agent, a database agent, and a document agent in parallel, then passing their combined output to a synthesis agent, cuts the wait substantially. The final answer draws on all three sources, and a timeout in one branch still leaves usable results from the others.
Workflows with different permission levels
Some processes mix low-risk steps, like answering questions, with high-risk ones, like issuing refunds or changing account details. Putting both in one agent means giving it every permission. Splitting them lets the question-answering specialist run with read-only access while a separate, tightly tested specialist handles the sensitive action, often with a human approval step. The result is a workflow that is easier to audit and harder to misuse.
The Routing Layer
The coordinator's routing decision is the most critical component. Get routing wrong and the whole system feels broken even when every specialist is doing its job perfectly.
Classification-Based Routing
The coordinator uses an LLM to classify the incoming request into one of a defined set of categories. Each category maps to a specialist agent.
ROUTING_PROMPT = """
Classify this customer message into exactly one category:
- SUPPORT: Technical issues, product problems, how-to questions
- BILLING: Payment, invoices, subscription, charges
- SALES: Pricing, upgrades, new features, purchasing
- RETURNS: Refunds, returns, exchanges
- OTHER: Anything that doesn't fit the above
Message: {message}
Respond with only the category name.
"""
Classification routing works well when categories are distinct. It struggles when requests fall into multiple categories at once, or when users phrase things ambiguously — which they do, constantly.
Keyword and Rule-Based Routing
For high-confidence routing on specific triggers, rule-based routing is faster and more reliable than LLM classification. "Order #12345" routes to the order status agent. An email from a known partner domain routes to the relevant agent.
In practice, most production multi-agent systems we build use both: rule-based routing for high-confidence cases, LLM classification for everything else.
Confidence Thresholds
When the classifier is uncertain, the system should not route to whichever specialist won by 0.51 vs 0.49. Low-confidence classifications go to a generalised handler or escalate to a human rather than making a potentially wrong routing decision. This single rule prevents a lot of bad outcomes.
Sequential vs Parallel Execution
Sequential Pipelines
Some workflows are inherently sequential: step 2 depends on the output of step 1.
A document processing pipeline might work sequentially:
- Extraction agent: Extract structured data from the document
- Validation agent: Check the extracted data against business rules
- Routing agent: Determine which department the validated data should go to
- Notification agent: Send the appropriate notifications
Each agent receives the previous agent's output, processes it, and passes to the next. The coordinator manages the sequence and handles failures at each step.
Parallel Execution
When multiple tasks can run simultaneously without dependencies, parallel execution cuts latency dramatically.
A research workflow might run in parallel:
- Web search agent: Searches for recent news on the topic
- Database agent: Retrieves internal records related to the topic
- Document agent: Searches the knowledge base for relevant content
All three run at the same time. The coordinator waits for all of them, then passes their combined output to a synthesis agent that assembles the final response.
LangGraph handles parallel execution natively. Standard LangChain chains are sequential.
State Management
Multi-agent systems need careful state management. Each agent needs to know:
- What the original request was
- What previous agents have done
- What context is relevant to its task
- What it should return
Shared State Object
Pass a structured TypedDict state object through the system that each agent reads from and writes to:
class AgentState(TypedDict):
original_message: str
customer_id: str
classification: str
classification_confidence: float
support_result: Optional[dict]
billing_result: Optional[dict]
final_response: Optional[str]
escalation_required: bool
escalation_reason: Optional[str]
Every agent receives this state, does its work, and returns an updated version. The coordinator reads the state to make routing decisions.
Memory Across Turns
For conversational multi-agent systems, each turn needs access to the conversation history. Store conversation history separately from the task state:
- Task state: The structured data flowing through the current workflow
- Conversation memory: The full history of the conversation for context
Mixing the two is one of the most common architectural mistakes we see. It looks fine at small scale and gets impossible to debug at any meaningful volume.
Error Handling and Fallbacks
Production multi-agent systems fail in ways single agents don't. A specialist might fail. The coordinator might misclassify. A parallel branch might time out while others complete. Every multi-agent system needs explicit error handling — assuming it'll work is how you end up with bizarre customer experiences nobody can reproduce.
Specialist failure: If the routed specialist fails, fall back to a generalised handler and log the failure. Never surface a technical error to the end user.
Timeout handling: Parallel branches have individual timeouts. If one branch times out, complete with the results that are available and note the gap.
Misclassification recovery: If a specialist receives a query it can't handle, it returns a structured signal indicating misclassification. The coordinator reroutes or escalates.
Circuit breakers: If a specialist fails repeatedly, stop routing to it and alert the operations team. A broken specialist producing errors at scale is worse than no specialist at all — at least with no specialist, the coordinator escalates cleanly.
LangGraph: The Right Tool for Multi-Agent Coordination
LangGraph's graph-based execution model is designed for multi-agent coordination. Nodes are agents or processing steps. Edges are routing decisions. The graph executor handles parallel execution, state management, and conditional routing.
from langgraph.graph import StateGraph, END
workflow = StateGraph(AgentState)
# Add nodes
workflow.add_node("classifier", classify_request)
workflow.add_node("support_agent", handle_support)
workflow.add_node("billing_agent", handle_billing)
workflow.add_node("synthesiser", synthesise_response)
# Add routing
workflow.add_conditional_edges(
"classifier",
route_to_specialist,
{
"SUPPORT": "support_agent",
"BILLING": "billing_agent",
"OTHER": END,
}
)
workflow.add_edge("support_agent", "synthesiser")
workflow.add_edge("billing_agent", "synthesiser")
LangGraph handles the execution graph, parallel branches, state passing, and conditional routing — the scaffolding that would otherwise need significant custom engineering. It's not the only way to build multi-agent systems, but in our work it's been the path of least resistance for anything beyond two or three agents.
Testing Multi-Agent Systems
Multi-agent systems are harder to test than single agents because failures can happen at the routing layer, within any specialist, or at the assembly layer. Bugs in one place look like bugs in another.
Test each layer independently:
- Test the classifier on a diverse set of inputs to verify routing accuracy
- Test each specialist independently with the range of inputs it might see
- Test the full system end-to-end with integration test cases
- Test error handling by deliberately failing individual components
Document the expected routing decision for every test case. Routing changes are the most common source of regression in multi-agent systems — and the hardest to spot without explicit checks.
Common Multi-Agent System Mistakes
Most multi-agent projects that struggle do so for architectural reasons, not model quality. These mistakes account for a large share of the problems we're asked to fix.
Going multi-agent before a single agent has hit its ceiling
Multi-agent designs are appealing on a whiteboard, so teams adopt them for workflows a single, well-scoped agent could handle. They take on routing, state design, and several specialists to monitor without getting anything in return. The extra complexity slows every change and makes bugs harder to trace. If tighter prompts and better escalation would solve the problem, that is the cheaper route.
Routing on a coin flip
A classifier that picks the top category at 0.51 versus 0.49 will send a meaningful share of requests to the wrong specialist. The user gets a confident answer to a question they didn't ask, and the specialist looks broken even though it worked as designed. Without a confidence threshold and a fallback path, routing errors quietly undermine the whole system.
Mixing task state with conversation memory
Putting the full conversation history and the structured workflow data in the same object seems simpler early on. At volume it becomes very hard to see what each agent actually received, what it changed, and why the coordinator made a decision. Keeping the two separate is far easier at the start than untangling them in production.
Assuming the happy path
Specialists fail, branches time out, and classifiers get it wrong. Systems built without explicit handlers for each case produce strange experiences that are hard to reproduce: a half-finished answer, a raw error, or a request that disappears. Every failure mode needs a defined response before launch, not after the first incident.
Changing routing without regression tests
A small edit to the routing prompt or rules can shift a whole category of requests to a different specialist. If expected routing isn't recorded for each test case, the change looks harmless until users notice. Routing changes are the most common source of regression, and the easiest to prevent with explicit checks.
Multi-Agent System Best Practices
These practices keep a multi-agent system maintainable as it grows from three specialists to ten. Most of them cost little to adopt at the start and a great deal to retrofit once the system is in production and handling real traffic.
- Start with one agent and earn the split. Label a week of real requests by the capability each needs. Only introduce specialists when the labels form distinct groups that a single agent handles poorly.
- Combine rule-based and LLM routing. Use deterministic rules for high-confidence triggers like order numbers or known partner domains, and an LLM classifier for the rest. Rules are faster, cheaper, and easier to test.
- Set a confidence threshold with a safe fallback. Send uncertain classifications to a generalist handler or a human, and log them. Review those logs to find categories the router is missing.
- Define structured contracts between agents. Give each specialist a typed input and output schema, including a way to signal "this isn't mine." Structured outputs make the coordinator's job simpler and failures easier to detect.
- Keep task state and conversation memory separate. Use a shared state object for the current workflow and a separate store for conversation history, so each can be inspected and debugged on its own.
- Handle every failure mode explicitly. Write fallbacks for specialist errors, per-branch timeouts for parallel work, rerouting for misclassification, and circuit breakers that alert the team when a specialist keeps failing.
- Test each layer, then the whole. Maintain separate test sets for the classifier and each specialist, plus end-to-end cases with the expected routing recorded. Deliberately break components to check fallbacks.
- Monitor specialists individually. Track error rates, escalation rates, and latency per specialist, not just for the system as a whole, so a struggling agent is visible before it drags down overall quality. Review these per-specialist numbers on a regular schedule.
Related guides
- OpenAI Assistants API vs building a custom agent
- LangChain vs LlamaIndex: which framework to use
- AI agent testing and QA before users find the bugs
- Our AI agent development services
When Multi-Agent Systems Are Overkill
Not every complex use case needs a multi-agent system. Before reaching for the architectural complexity, verify:
- Is the routing genuinely necessary, or can a single agent handle the variety of inputs with better prompt engineering?
- Is the parallel execution performance gain actually worth the complexity?
- Does the team have the capacity to maintain multiple agents instead of one?
A well-designed single agent with clear scope and good escalation handling is more maintainable than a complex multi-agent system with unclear routing logic. We've talked clients out of multi-agent architectures more than once. Build multi-agent when the single-agent approach has clearly hit a ceiling — not before.
Talk to us about your architecture — we build and maintain multi-agent systems in production and can help you assess honestly whether you actually need one.
Frequently Asked Questions
What is a multi-agent system in AI?
A multi-agent system is an architecture where multiple specialised AI agents work together, each handling a specific type of task, with a coordinator managing which agent receives which request. Rather than one agent trying to handle everything, each specialist is optimised and tested independently. The coordinator routes incoming work to the right specialist and assembles the results into a final response.
When does a business actually need a multi-agent system?
You need a multi-agent system when a single entry point genuinely receives different types of requests that require different capabilities — for example, a customer inbox that handles support, billing, and sales enquiries. If your use case is one type of task done at high volume, a single well-scoped agent is almost always the better choice. Multi-agent adds maintenance cost; only pay that cost when a single agent has clearly hit a ceiling.
How much does it cost to build a multi-agent system?
Cost depends on the number of specialist agents, the complexity of the routing logic, and the integrations required. A three-agent system with a coordinator typically takes four to eight weeks to build properly, including testing and error handling. More complex systems with five or more specialists and parallel execution can run to three to four months. The ongoing maintenance cost is higher than a single agent because each specialist needs to be monitored and updated independently.
What is the difference between LangGraph and LangChain for multi-agent systems?
LangChain is a framework for building sequential chains of LLM calls and tool use. LangGraph extends it with a graph-based execution model that supports parallel branches, conditional routing, and stateful multi-agent workflows. For anything beyond two or three agents, LangGraph's model maps directly onto multi-agent architecture: nodes are agents, edges are routing decisions, and the graph executor handles state passing and parallel execution. Standard LangChain chains require custom engineering to achieve the same result.
How do you handle errors in a multi-agent system?
Each failure mode needs an explicit handler. If a specialist agent fails, the coordinator falls back to a generalised handler rather than surfacing an error to the end user. If a parallel branch times out, the system completes with the results that are available. If the classifier is uncertain about routing, the request goes to a human rather than being routed to the wrong specialist. Circuit breakers stop routing to a repeatedly-failing specialist and trigger an alert — a broken specialist producing errors at scale is worse than no specialist.
Can multi-agent systems run tasks in parallel?
Yes, and parallel execution is one of the main reasons to build a multi-agent system. When multiple tasks are independent — for example, searching different data sources simultaneously — a coordinator can fan out to several specialists at once and wait for all of them before assembling the response. This cuts end-to-end latency significantly compared to running the same tasks sequentially. LangGraph handles parallel branches natively; standard LangChain chains are sequential by default.
How do you test a multi-agent system before going live?
Test each layer independently: first the classifier on a diverse set of real inputs to check routing accuracy, then each specialist in isolation with the range of inputs it might receive, then the full system end-to-end with integration test cases. Also test error paths deliberately — fail individual specialists and verify the fallbacks behave correctly. Document the expected routing decision for every test case, because routing changes are the most common source of regression and the hardest to catch without explicit checks.
Conclusion
Multi-agent systems solve a specific problem: a single entry point that receives genuinely different kinds of work, each needing different tools, knowledge, and guardrails. Splitting that work across narrow specialists with a coordinator in front keeps each agent reliable and independently testable, and parallel execution can cut latency when tasks don't depend on each other.
The cost is operational. Every specialist is something to monitor and update, routing becomes the most fragile part of the system, and shared state has to be designed deliberately rather than left to accumulate in prompts. Error handling has to be explicit at every layer: fallbacks for failed specialists, partial results for timed-out branches, human review for uncertain routing, and circuit breakers for anything failing repeatedly.
The most important caveat is the one teams skip. A well-scoped single agent with good escalation is easier to run than a multi-agent system with fuzzy routing, so only move to multiple agents once one has clearly hit its ceiling.
A practical next step is to take a week of real requests from the entry point you're automating and label each with the capability it needs. If the labels cluster into three or four distinct groups, you have a case for specialists. If you want help designing or reviewing that architecture, our AI agent development team builds and runs these systems in production.
