Multi-agent orchestration is a software pattern that coordinates several AI agents to divide a complex task, pass state between them, and produce a unified result. The pattern emerged because a single LLM running in one context window cannot reliably plan, research, and execute a long task without losing track of itself. Splitting the work across specialized agents, each with its own prompt, tools, and memory, raises the ceiling on what an agentic system can finish without human babysitting. Three Python frameworks dominate the conversation in production code: LangGraph from LangChain, the role-based CrewAI library, and Microsoft's AutoGen (renamed AG2). OpenAI ships lighter-weight primitives through OpenAI Swarm and the OpenAI Agents SDK on top of the Responses API, while Anthropic stays out of the orchestration layer and exposes a tool-use API that any of the above can call. The right choice depends on whether the workload needs explicit control flow modeled as a DAG, conversational coordination borrowed from the Actor model, or a thin handoff abstraction layered on a single vendor's API.
What Multi-Agent Orchestration Is and Why One Agent Is Not Enough
Multi-agent orchestration is the practice of splitting a complex task across several specialized agents so that each handles the slice it is best suited for, with a coordinator routing outputs from one agent as inputs to the next. The motivation is structural rather than aesthetic. A single LLM call has one context window, one system prompt, and one toolset, so a task that spans research, code generation, and review either runs sequentially through one overloaded prompt or fans out across agents that each carry only what they need. Task decomposition is the design move: the orchestrator splits the goal into subgoals, each agent owns a subgoal, and the framework passes state between them. The result is closer to how human teams handle work, with a project lead routing pieces to specialists and merging outputs back. The supervisor pattern and the orchestrator-worker pattern are the two topologies most teams reach for first, and both predate LLMs in distributed-systems literature. For background on what an individual agent is and how it differs from a chatbot, see how autonomous AI agents work and what makes them different from chatbots.
- Agent: an LLM-powered process with a defined role, prompt, and toolset that can call functions and reason over their output.
- Orchestrator: the control layer that decides which agent runs next, passes state between agents, and terminates the workflow.
- Handoff: the moment one agent transfers control and accumulated state to another, ending its own turn.
- Task decomposition: splitting a complex goal into smaller subgoals each handled by a specialized agent.
- Supervisor pattern: a topology where a single agent routes to worker agents and consolidates their results.
LangGraph: Multi-Agent Orchestration with Explicit Graph Control Flow
Multi-agent orchestration in LangGraph is modeled as a directed graph where each node is an agent or tool-call step, and edges define the conditional routing logic that decides which node runs next. The execution engine draws on Pregel, the vertex-centric model Google originally published for large-scale graph computation, adapted to agent steps instead of graph vertices. LangChain documents LangGraph as a low-level orchestration framework for building stateful, long-running agent workflows, with first-class persistence and a graph API that compiles nodes and edges into an executable runtime (LangChain, LangGraph overview). The graph model is explicit: developers declare nodes, declare conditional edges, and the compiled graph executes deterministically against a shared state object. State persistence is a first-class primitive. LangGraph's graph API documentation describes checkpointing as the mechanism that snapshots state at super-step boundaries, enabling pause-and-resume, human-in-the-loop interrupts, and time-travel debugging (LangChain, LangGraph graph API). The checkpointer is pluggable, with SQLite for local development and Postgres or Redis as production backends. For retrieval-heavy graphs, the same state can carry vector-store results between nodes; this pattern is covered in how retrieval-augmented generation feeds context into LLM agents.
- Declare state: define a typed state object, usually a Python TypedDict, that every node reads from and writes to.
- Declare nodes: register each agent or tool step as a node function that takes state and returns a state update.
- Declare edges: wire nodes with conditional edges; the edge function inspects state and returns the next node name.
- Compile the graph: bind a checkpointer for persistence and produce an executable graph object.
- Stream execution: invoke the graph with an input and stream node-by-node updates back to the caller.
CrewAI and AutoGen: Role-Based and Conversational Multi-Agent Orchestration
Multi-agent orchestration in CrewAI assigns each agent a named role and goal, while AutoGen (now AG2) structures coordination as a back-and-forth conversation between agents rather than a predetermined graph. CrewAI raises the abstraction level above LangGraph: a developer declares a Researcher, a Writer, and a Reviewer, gives each a role description and a goal, and the framework runs them as a crew with sequential or hierarchical task assignment. The mental model is closer to a job description than a state machine. AutoGen takes a third path: agents converse, and the orchestration logic is the conversation's turn-taking policy, an approach that owes its lineage to the Actor model that Erlang popularized for concurrent message-passing systems. Microsoft Research originally released AutoGen, and the project continued community development under the AG2 name. Conversational orchestration suits open-ended tasks where the sequence of steps is not knowable up front; graph orchestration suits workflows where the steps are. Pydantic AI sits in a fourth category: type-safe agent definitions with structured outputs validated against Pydantic models, which the surrounding orchestrator can then route. For how these frameworks call out to APIs and tools, see how agentic workflows combine tool calls and API access.
- CrewAI role definition: each agent receives a role string, a goal, a backstory, and an optional list of tools.
- CrewAI task list: tasks are declared separately and assigned to agents; the crew runs them sequentially by default.
- AutoGen conversable agents: agents inherit from a base ConversableAgent class and exchange messages through a group chat manager.
- AutoGen group chat: a manager agent picks the next speaker based on a selection function or LLM-driven routing.
- Output contract: CrewAI returns the final task output; AutoGen returns the full message history for inspection.
OpenAI Swarm and Platform-Native Multi-Agent Orchestration Primitives
Multi-agent orchestration can also be assembled from platform-native primitives: OpenAI's Swarm experiment and the Responses API handoff model provide lightweight agent-to-agent routing without a dedicated graph runtime. OpenAI describes its Agents SDK and Responses API as a set of building blocks for developers to construct agentic applications, including handoffs between agents, built-in tools, and tracing for debugging and evaluation (OpenAI, new tools for building agents). OpenAI Swarm came earlier as an experimental routine-and-handoff library, and the patterns it established carried into the Agents SDK. The trade-off is depth against simplicity. A handoff is a single function call that swaps the active agent and its tools; there is no graph to compile and no separate state machine to maintain. Microsoft's Semantic Kernel covers similar ground on the .NET and Python side, with planners and skills that compose into agentic flows. Teams that need cross-language calls between agents often wrap their tools as gRPC services, then expose those services to whichever framework holds the orchestration layer. The platform-native path suits teams that have already committed to one vendor's API surface and want the minimum integration surface. For how tools themselves get wired to agents through an open protocol, see how the Model Context Protocol standardizes tool connections for AI agents.
- Swarm routines: a sequence of instructions and tool calls that an agent executes before optionally handing off.
- Responses API handoff: an explicit tool call that transfers the active role to a named target agent.
- Built-in tools: the Agents SDK ships web search, file search, and computer use as first-party tools that any agent can invoke.
- Tracing: every run is recorded as a trace with spans per tool call and handoff, viewable in the OpenAI dashboard.
- Semantic Kernel parallel: Microsoft's planner-and-skill model offers a comparable lightweight primitive set on the .NET stack.
Framework Comparison: Topology, State, Concurrency, and Observability
Multi-agent orchestration frameworks differ most sharply on four axes: the topology model (graph, role crew, or conversation), how state is stored and recovered, whether agents can run concurrently, and what observability tooling ships with the framework. LangGraph centers on explicit graph topology with built-in checkpointing for state persistence, per its graph API documentation (LangChain, LangGraph graph API). CrewAI organizes work around roles and goals with sequential execution as the default. AutoGen treats coordination as a multi-agent conversation, with concurrency emerging from its async message loop. OpenAI's Agents SDK keeps state inside the Responses API session and surfaces tracing through the platform dashboard (OpenAI, new tools for building agents). Beyond the LLM frameworks themselves, durable-execution engines such as Temporal, distributed actor runtimes such as Ray, and task queues such as Celery often sit underneath, providing the retries, scheduling, and idempotency that LLM-native runtimes do not yet match. The table below maps the trade-offs side by side; note that feature surfaces shift between releases, so the vendor docs are the authoritative reference for any production decision.
| Framework | Topology | State model | Concurrency | Observability |
|---|---|---|---|---|
| LangGraph | Directed graph of nodes and conditional edges | Typed state object, checkpointed per node | Fan-out edges run nodes in parallel | LangSmith tracing integration |
| CrewAI | Role-based crew with task list | Task context passed between agents | Sequential by default; parallel task assignment supported | Verbose logging and crew telemetry |
| AutoGen / AG2 | Conversational group chat | Message history is the state | Async turn-taking with concurrent agent activity | Message logs and run inspection |
| OpenAI Agents SDK | Handoff graph between named agents | Session state inside the Responses API | Sequential handoffs per session | Built-in trace viewer in the OpenAI dashboard |
When to Use Each Framework: A Decision Guide
Multi-agent orchestration framework choice should follow the complexity and determinism of your workflow rather than framework popularity. The right question is what shape the work takes. A workflow with a known set of steps, conditional branching, and a hard requirement for pause-and-resume points reaches for LangGraph because the graph topology and checkpointing line up with that shape. A workflow that maps cleanly onto roles, where the steps each role takes are obvious from the role description, fits CrewAI's role-based abstraction without the ceremony of declaring a graph. A workflow where the steps are not knowable up front, and where agents need to negotiate or critique each other's output, suits AutoGen's conversational model. A workflow that lives entirely inside one vendor's API and needs the smallest integration surface picks the OpenAI Agents SDK or Microsoft's Semantic Kernel. The decision rarely splits along framework quality; it splits along workflow shape and the engineering team's preference for explicit control versus emergent coordination.
- Pick LangGraph when: the workflow has known branches, needs durable state across long runs, and benefits from human-in-the-loop interrupts.
- Pick CrewAI when: the work decomposes cleanly into roles with goals, and a sequential or hierarchical handoff is enough.
- Pick AutoGen / AG2 when: the task is open-ended, agents must negotiate or critique each other, and conversational logs are valuable as audit trails.
- Pick OpenAI Agents SDK when: the production stack is single-vendor OpenAI and built-in tracing through the platform is acceptable as the observability layer.
- Pick Semantic Kernel when: the host application runs .NET and planner-and-skill composition matches the existing service architecture.
Anthropic Tool Use and Claude as a Multi-Agent Orchestration Participant
Multi-agent orchestration that includes Claude relies on Anthropic's tool-use API rather than a first-party orchestration runtime, which makes Claude a capable participant in any of the frameworks above without requiring a vendor-specific integration. Anthropic's tool-use documentation describes a model-driven loop where the developer defines tools as JSON schemas, the model decides which to call, and the developer executes the call and returns the result for the next turn (Anthropic, tool use overview). The orchestration layer sits outside Claude. For GUI-driven subagents, Anthropic exposes a computer-use tool that lets Claude see a screenshot and emit mouse and keyboard actions, suitable for browser automation or desktop workflows when wrapped by an orchestrator (Anthropic, computer use tool). The practical implication is that LangGraph, CrewAI, and AutoGen all integrate Claude through the Messages API, and a team standardizing on Claude does not lose access to any of the orchestration topologies above. For background on how individual agents perceive and act on their environment, see how autonomous AI agents perceive and act on their environment.
- Tool schema: tools are declared as JSON schemas; Claude returns a structured tool_use block when it decides to call one.
- Developer loop: the orchestrator executes the tool, returns a tool_result block, and Claude continues with the new context.
- Subagent pattern: a parent Claude agent can call a child Claude agent through a tool that wraps the Messages API.
- Computer use: Claude can drive a virtual desktop through the computer-use tool when the orchestrator captures screenshots and forwards actions.
- Framework neutrality: Anthropic ships no orchestrator, so LangGraph, CrewAI, and AutoGen all treat Claude as one more model behind their abstraction.
References
- LangChain, LangGraph overview
- LangChain, LangGraph graph API
- OpenAI, new tools for building agents
- Anthropic, tool use overview
- Anthropic, computer use tool
Further reading
Frequently Asked Questions
What is the difference between multi-agent orchestration and a single LLM with tools?
Multi-agent orchestration splits a task across several specialized agents, each with its own context window and toolset, rather than routing everything through one model. A single LLM with tools handles all subtasks sequentially in one context, which caps at that model's context limit and cannot parallelize independent subtasks.
Can multiple agents run concurrently in these frameworks?
LangGraph supports fan-out edges that dispatch independent nodes in parallel within a single graph execution. AutoGen supports concurrent agent turns through its async conversation model. CrewAI's default execution is sequential per crew, though parallel task assignments are supported in newer versions.
What is a handoff in multi-agent orchestration?
A handoff is the moment one agent passes control and its output state to the next agent in the pipeline. In OpenAI's Responses API model, handoffs are explicit tool calls that transfer the active agent role; in LangGraph, they are conditional edges between graph nodes.
Does Anthropic offer its own multi-agent orchestration framework?
Anthropic does not ship a first-party orchestration framework equivalent to LangGraph or CrewAI. Claude participates in multi-agent systems through the tool-use API, and developers orchestrate it using LangGraph, AutoGen, or a custom routing layer built on the Anthropic Messages API.









