Field notes · AI Agency

    Multi-Agent Systems Explained: Orchestration Patterns That Actually Work.

    Multi-agent systems split work across specialized AI agents. Here are the four orchestration patterns that work in production, the frameworks to build them, and how to plug them into real business systems via MCP.

    7 sections
    AI Agency
    10
    Multi-Agent Systems Explained: Orchestration Patterns That Actually Work

    Multi-agent systems split a task across several specialized AI agents instead of forcing one model to do everything. The four patterns that actually work in production are supervisor-worker, hierarchical, swarm, and debate. Frameworks like Anthropic's Agent SDK, CrewAI, and LangGraph give you the primitives. The real unlock is wiring those agents into systems that take action on your business, which is what MCP (Model Context Protocol) and platforms like ACA make possible.

    Short answer: A multi-agent system is two or more AI agents that coordinate to complete a task. The orchestration patterns that hold up in production are supervisor-worker (one agent delegates to specialists), hierarchical (nested supervisors with deeper trees), swarm (peer agents that hand off control), and debate (agents critique each other before answering). Pick the pattern that matches the structure of the task, not the one with the cleanest demo on Twitter.

    What Is a Multi-Agent System?

    A multi-agent system (MAS) is an architecture where two or more AI agents, each with its own role, tools, memory, and prompts, work together to complete a task. Each agent is a wrapped LLM call with a defined system prompt, a tool set, and an exit condition. The system is the wiring between them: who calls whom, who sees what context, who decides when the task is done.

    The naive version of "AI agent" is one giant prompt with twenty tools attached. That works until the task gets complex. Then the model gets confused about which tool to use, forgets context halfway through, or burns through tokens reasoning about every possible step. Splitting the work into smaller, focused agents fixes most of these problems by reducing the cognitive load on each model call.

    A multi-agent system is not the same as a workflow. A workflow is a fixed pipeline. A multi-agent system reasons about what to do next at runtime, including which agent to call and what to pass it.

    When Single Agents Break (And When to Go Multi-Agent)

    Most tasks do not need a multi-agent system. If you can describe the work in three steps and it uses fewer than five tools, build it as a single agent. You will move faster and debug easier.

    Go multi-agent when one of these is true:

    • The task has distinct skill domains. Research, writing, code generation, and quality review pull on different prompting strategies. Splitting them lets you tune each one in isolation.
    • The context window is the bottleneck. One agent cannot hold the full document, the research notes, the draft, and the style guide at once. Subagents with isolated context windows let you process more.
    • The task is unpredictable in shape. If you do not know in advance how many steps it takes, a supervisor agent can decide at runtime.
    • You need parallelism. Three subagents running simultaneously finish in roughly a third of the time of one agent doing the same work sequentially.

    The Four Orchestration Patterns That Actually Work

    Most production multi-agent systems are a variant of one of these four patterns. Pick deliberately.

    Supervisor-Worker

    One supervisor agent receives the task, decides which worker agent to call, sends them a focused prompt, and assembles the results. Workers do not talk to each other. They only talk to the supervisor.

    This is the pattern Anthropic's research multi-agent system uses internally, and the one most teams should start with. It is easy to reason about, easy to debug (the supervisor's transcript is the audit log), and easy to extend. Add a new specialty? Add a new worker and update the supervisor's prompt.

    Use supervisor-worker when the task decomposes cleanly into independent subtasks. A research agent that needs to search the web, summarize sources, and write a report is the canonical example. The supervisor splits the topic into research questions, fans out workers to investigate each one, then composes the final answer.

    Hierarchical

    A supervisor of supervisors. The top-level agent delegates to mid-level supervisors, which delegate to workers. You get hierarchical when the task is large enough that the top supervisor cannot keep all the subtasks in context.

    This is what you build when supervisor-worker hits its scaling ceiling. A common shape: a project manager agent oversees a research team supervisor and a writing team supervisor. Each of those supervises its own workers. The project manager only sees high-level status updates, not the raw research notes.

    The risk: communication overhead. Every layer adds latency and token cost. Do not go hierarchical until you actually need it.

    Swarm

    Peer agents that hand off control to each other based on the current state of the task. There is no central supervisor. Each agent has a system prompt describing when it should take over and when it should hand off, and a tool that lets it pass control to a named peer.

    OpenAI's Swarm reference implementation popularized this. It works well for conversational systems where the "role" of the assistant changes based on what the user is doing: a triage agent hands off to a billing agent, which hands off to a refund agent if needed, which hands back to billing when done.

    Swarm is harder to debug than supervisor-worker because control flow is decentralized. Use it when the task is genuinely conversational and the appropriate agent depends on dynamic conditions rather than a planned decomposition.

    Debate

    Two or more agents produce competing answers, then critique each other's output, then a judge agent (or a final round of the same agents) produces the answer. Variants include the proposer/critic pattern, multi-agent debate for math and reasoning tasks, and red-team/blue-team setups for safety review.

    Debate is expensive. You are running the same task two or three times. But it raises accuracy on hard reasoning problems and surfaces edge cases that a single agent rubber-stamps. Use it for high-stakes outputs (legal analysis, code that touches money, content that ships under your name) where the cost of wrong is much higher than the cost of extra tokens.

    Use supervisor-worker when: the task splits into independent specialties, you want a clean audit trail, and you are building your first multi-agent system.

    Use hierarchical when: supervisor-worker is hitting context limits and you genuinely have nested supervisory needs.

    Use swarm when: the task is conversational and the right agent depends on runtime state rather than a planned decomposition.

    Use debate when: output quality matters more than token cost and you need to catch errors a single agent would miss.

    Frameworks for Building Multi-Agent Systems

    You can implement any of these patterns from scratch with raw API calls. You probably should not. Three frameworks cover most of what teams ship today.

    Anthropic Agent SDK and Subagents

    Anthropic's Agent SDK (the same primitives that power Claude Code) supports subagents directly. A subagent is a separate Claude instance with its own system prompt, its own tools, and its own context window. The parent agent calls it like a tool. The subagent runs to completion, returns a summary, and its context disappears.

    This is the cleanest implementation of the supervisor-worker pattern available right now. The context isolation matters: a subagent that searches the web and reads ten pages returns a 200-word summary to the parent, not the full content of every page. Your parent agent's context window stays clean even as the system does heavy work underneath.

    If you are building inside Claude's ecosystem and want the lowest-friction path to a working multi-agent system, start here.

    CrewAI Crews

    CrewAI models multi-agent systems as crews of agents, each with a role, a goal, a backstory, and a set of tools. You define tasks and assign them to agents, then run the crew. CrewAI handles the orchestration loop, including sequential and hierarchical execution modes.

    The strength of CrewAI is opinionation. The abstractions (Agent, Task, Crew, Process) match how most people think about multi-agent systems. The downside is the same: if your system does not fit the role/goal/task mental model cleanly, you will fight the framework.

    CrewAI is a strong default for teams that want to ship something this week and do not need fine-grained control over the message graph.

    LangGraph

    LangGraph treats multi-agent systems as state machines. You define nodes (agents, tools, decision points), edges (which node runs next under which conditions), and shared state that flows through the graph. It is the most flexible of the three and the one closest to what you would build by hand.

    Use LangGraph when your system does not fit a canonical pattern, when you need precise control over routing logic, or when you need to persist state across long-running runs (human-in-the-loop, interruptions, retries). The learning curve is steeper, but the ceiling is higher.

    Most production multi-agent systems we have seen at scale end up on LangGraph or a hand-rolled equivalent. CrewAI gets you started; LangGraph keeps you there.

    Plugging Multi-Agent Systems Into ACA via MCP

    A multi-agent system that only writes text is a demo. A multi-agent system that takes action on your business is a product. The bridge between them is Model Context Protocol (MCP), Anthropic's open standard for letting agents call external systems through a uniform interface.

    MCP lets any MCP-compatible agent (Claude, Cursor, your custom LangGraph crew, an Anthropic Agent SDK subagent) call tools defined by an MCP server. The server can be anything: a database, a CRM, an outreach platform, an internal API. The agent does not need to know how the server works. It just sees a list of tools and uses them.

    ACA exposes the outbound stack as an MCP server. Your agents can read your CRM, launch a campaign, generate content, check inbox replies, or update lead scores, all through standard MCP calls. The agent decides what to do; ACA executes it. This is the missing layer for most AI agencies: you can build an impressive research agent in CrewAI, but until it can actually book the meeting or send the LinkedIn message, it is just a chatbot in a fancy wrapper.

    A working setup looks like this: a supervisor agent receives a lead-generation brief, dispatches a research subagent to enrich the target list, dispatches a copywriting subagent to draft sequence variants, then calls the ACA MCP server to import the leads and launch the campaign across LinkedIn, email, and WhatsApp. The agents reason. ACA acts. See the AI agency stack for how this fits into a full delivery system.

    What multi-agent systems are good at versus what they are not: in our experience deploying these for agencies, multi-agent systems materially outperform single agents on tasks that decompose cleanly (research, content production, multi-source synthesis). They underperform on tasks that are essentially one continuous reasoning chain, where the handoff cost between agents exceeds the benefit of specialization. Plan accordingly.

    Common Failure Modes

    The same problems show up in almost every multi-agent system that breaks in production:

    • Context drift between agents. The supervisor passes a summary to a worker, the worker fills in gaps with its own assumptions, and the final output is subtly off-brief. Fix this by passing structured context (JSON with explicit fields) rather than freeform text.
    • Infinite handoff loops. Agent A hands off to Agent B, which hands back to A, which hands back to B. Add a maximum hop count and a forced termination path in every system.
    • Token explosion. A four-agent debate on a long input can burn 50,000+ tokens for a single output. Set per-agent token budgets and short-circuit when they are exceeded.
    • Tool collisions. Two agents both decide to update the same record. Either serialize tool calls through the supervisor or design tools to be idempotent.
    • Hallucinated handoffs. An agent invents a peer agent that does not exist and tries to hand off to it. Constrain handoff targets at the framework level, not just in the prompt.

    Frequently Asked Questions

    What is the difference between a multi-agent system and a workflow?

    A workflow is a fixed sequence of steps defined at design time. A multi-agent system uses agents to decide what happens next at runtime, including which agent to call and what to pass it. Workflows are predictable and cheap. Multi-agent systems are flexible and expensive. Use workflows when the steps never change; use multi-agent systems when the path through the task depends on the input.

    Do I need multiple LLM providers to build a multi-agent system?

    No. Most multi-agent systems use one provider (usually Anthropic or OpenAI) for all agents and vary the system prompts, tool sets, and models (for example, a cheap model for routing and a stronger model for final synthesis). Mixing providers adds operational complexity without clear quality gains for most use cases.

    How many agents should a multi-agent system have?

    Start with two. Add a third only if the second agent's role is doing too many things. In practice, most production systems we see in agencies use three to five agents. Past seven agents, you are usually better off restructuring into a hierarchical pattern or collapsing roles back together.

    Can I run multi-agent systems with open-source models?

    Yes, with caveats. Llama 3.3 70B and Qwen 2.5 72B handle supervisor roles reasonably well. Smaller open models tend to struggle with the meta-reasoning required to decide which subagent to call. If you go open-source, expect to spend more on prompt engineering and to accept lower reliability on edge cases.

    How do I monitor a multi-agent system in production?

    Trace every agent call as a separate span in your observability tool (LangSmith, Langfuse, Helicone, or a custom OpenTelemetry setup). Log the full message history per run, not just the final output. Most production failures are debugged by reading the supervisor's transcript and finding the moment context got passed wrong.

    What is MCP and why does it matter for multi-agent systems?

    MCP (Model Context Protocol) is an open standard from Anthropic for connecting AI agents to external tools and data sources through a uniform interface. It matters because it lets your multi-agent system act on real business systems (CRMs, outreach platforms, databases) without writing custom integration code for each one. Platforms like ACA expose their full feature set as MCP servers, so your agents can launch campaigns, generate content, or update leads through standard tool calls.