Book a consultation

Multi-agent systems: choose the coordination cost before the pattern

A production guide to deterministic graphs, routers, supervisors, event choreography, handoffs, state, checkpoints, and human approval.

Multi-agent systems: choose the coordination cost before the pattern
Supervisor
  ├── Research agent
  ├── Pricing agent
  ├── Legal agent
  └── Writer agent

This diagram tells you almost nothing about whether the system will work.

It does not show who owns state, what happens when pricing succeeds and legal times out, whether agents can repeat side effects, or how context moves between them. It also hides the easiest architectural question to miss. Why are these separate agents at all?

Centralized orchestrator-worker, event-driven choreography, and human-in-the-loop are useful categories. Before choosing one, account for the coordination cost it creates.

Three production orchestration choices: deterministic workflow, supervisor, and event choreography
Choose the least autonomous pattern that completes the work.
Three-layer topology of an AI system showing input, decision, and control surfaces
Orchestration is the decision surface. Controls keep its state and tools bounded.

Use the least autonomous design that completes the work

If the steps are known, write them down as a workflow.

receive_request
  ↓
load_customer
  ↓
validate_products
  ↓
calculate_price
  ↓
draft_offer
  ↓
approve_discount

The model may extract products from an email and draft the offer. Customer validation, price calculation, and discount rules can remain deterministic.

This design is easier to test because the graph constrains the allowed order. It is cheaper because an LLM does not plan steps the application already knows. It is safer because a prompt cannot invent a new path around the approval node.

LangChain's custom workflow pattern explicitly supports this mix. A LangGraph node can be ordinary code, a model call, or an agent. Sequential steps, branches, loops, and parallel execution stay visible in the graph.

Use a general agent loop when the path genuinely depends on information discovered during execution. Even then, put budgets and permission boundaries around it.

Count probabilistic boundaries

If one agent completes its step correctly 95% of the time, four dependent agents do not give the workflow 95% reliability.

Under a simplified independence assumption, four required steps at 95% each produce about 81% end-to-end success. Real errors are not independent, but the calculation exposes the direction of travel. Every routing decision, context handoff, and synthesis step creates another place to fail.

route correctly
  × specialist succeeds
  × context transfers correctly
  × result is reconciled
  × final action is valid

Do not use that multiplication as a forecast. Use it to challenge an architecture that keeps adding agents without adding evidence.

Track success at every boundary. An end-to-end failure should identify whether routing, specialist execution, transfer, or synthesis failed.

Pattern 1 uses a deterministic graph

Choose a graph when the business process has known stages and controlled branches.

from typing import Literal, TypedDict
from langgraph.graph import END, START, StateGraph

class QuoteState(TypedDict):
    request_id: str
    risk: Literal["low", "high"]
    draft: dict | None
    approved: bool | None

def route_after_draft(state: QuoteState) -> str:
    return "approve" if state["risk"] == "high" else "publish"

graph = StateGraph(QuoteState)
graph.add_node("prepare", prepare_quote)
graph.add_node("approve", request_approval)
graph.add_node("publish", publish_quote)
graph.add_edge(START, "prepare")
graph.add_conditional_edges("prepare", route_after_draft)
graph.add_edge("approve", "publish")
graph.add_edge("publish", END)

The inspectable route matters more than the syntax. Risk classification decides whether approval occurs. The model does not get to skip it.

Graphs work well for regulated operations, long-running jobs, and processes where recovery must resume from a known state.

Pattern 2 uses a supervisor with workers

A supervisor is useful when one component can decompose a request and delegate bounded tasks to specialists.

The supervisor should own the execution budget and final result. Workers should receive only the context and tools needed for their task. Treat them as typed services, not personalities in a group chat.

{
  "task_id": "research-17",
  "task_type": "policy_lookup",
  "inputs": {"jurisdiction": "PL", "topic": "refund"},
  "allowed_tools": ["search_policy_repository"],
  "deadline_ms": 5000,
  "output_schema": "PolicyFindingV2"
}

The worker returns a structured result with citations and status. The supervisor should not infer whether the worker timed out from a missing chat message.

Central control simplifies audit and budget enforcement. It can also become a latency bottleneck and a single failure point. Persist supervisor state and make worker tasks independently retryable.

Pattern 3 uses a router

A router classifies the request and sends it to one or more specialists. It fits workloads where categories are distinct and most requests need one specialist.

LangChain documents routers as one of its main multi-agent patterns. The quality of the router is a separate evaluation problem. Measure confusion by route and the cost of misrouting.

Use deterministic routing for known keys such as tenant, product, locale, or permission domain. Use a model only for semantic ambiguity. Add a fallback route for low confidence.

Parallel fan-out is useful when sources are independent, but it changes the failure contract. Decide whether the synthesis waits for all branches, a quorum, or a deadline. Record missing branches in the output instead of presenting a partial answer as complete.

Pattern 4 uses event choreography

In choreography, components subscribe to events rather than receiving commands from one supervisor.

mortgage.application.received
  ├── identity agent
  ├── affordability agent
  └── fraud agent

assessment.completed events
  ↓
decision aggregator

This allows parallel work and loose deployment coupling. It also inherits the hard parts of event-driven systems. Events can arrive twice, out of order, or after a deadline. Consumers need idempotency. The aggregator needs a correlation ID and a rule for incomplete sets.

Do not put private model context on a shared event bus. Publish a minimal contract and let authorized consumers retrieve what they need.

Choreography works best when participants are genuinely independent. If every agent needs to wait for the previous one, a graph or supervisor is easier to reason about.

Pattern 5 uses state-driven handoffs

A handoff changes who interacts with the user or which capabilities are active.

LangChain handoffs persist a state variable such as current_step or active_agent. Tool calls update that state and route control. This can enforce a sequence, such as collecting a warranty ID before enabling a refund capability.

State-driven handoffs are different from passing an entire conversation to a new independent agent. The safer design often keeps one durable state and changes prompt, tools, or routing based on that state.

Validate handoff preconditions in code. A model request to switch into a refund state should fail if identity verification is incomplete.

Human approval is an interrupt, not a chat message

"Ask the user before sending" inside a system prompt is not an approval mechanism.

An approval needs a persisted pending action, an authenticated reviewer, an immutable view of the proposed arguments, and an explicit decision.

LangGraph interrupts pause execution and save graph state until external input resumes the run. LangChain's human-in-the-loop middleware can intercept selected tool calls and support approve, edit, or reject decisions.

HumanInTheLoopMiddleware(
    interrupt_on={
        "read_customer": False,
        "send_offer": {"allowed_decisions": ["approve", "reject"]},
        "issue_refund": {"allowed_decisions": ["approve", "reject"]},
    }
)

The reviewer needs the evidence behind the action. Show source records, policy checks, amount, recipient, and downstream effect. A button next to the agent's prose is not enough.

Persist state outside conversation history

Multi-agent systems need explicit state for completed work, pending work, budgets, approvals, and side effects.

{
  "workflow_id": "quote-1842",
  "completed": ["extract", "customer_check"],
  "pending": ["price", "legal_review"],
  "failed": [],
  "tool_calls_used": 4,
  "cost_used_usd": 0.11,
  "approval": null
}

Conversation messages are evidence, not the control plane.

LangGraph persistence saves graph state as checkpoints organized into threads. Checkpoints support resume, human review, debugging, and recovery after node failure. In parallel steps, pending writes can preserve work from successful nodes rather than rerunning everything.

Use a durable checkpointer in production. In-memory state disappears at exactly the moment recovery matters.

Define the failure contract for each node

Every node or worker should declare:

timeout
retry policy
idempotency behavior
fallback
compensation action
escalation target
output schema

A read-only search can retry. A payment action needs an idempotency key and reconciliation. A draft can fall back to another model. A legal decision may have no automated fallback and should escalate.

Bound loops with maximum model calls, tool calls, tokens, cost, and elapsed time. LangChain middleware provides call limits, retries, fallbacks, PII controls, and approval hooks. Apply these as workflow policy, not generic defaults across every task.

Choose based on ownership and failure

The right pattern is the one whose state and failures your team can operate.

Use a deterministic graph for known processes. Add a router when requests divide into stable categories. Use a supervisor when bounded specialists need central planning. Use choreography when participants are independent and event semantics are already part of the platform. Use handoffs for stateful user journeys. Add humans where risk requires a decision the system should not own.

Most production systems combine two or three of these patterns. The mistake is combining all of them before one workflow has proved its value.

The next article defines the security boundary around those tools and actions: Agent governance starts with the blast radius

Start with the work

Bring the workflow that needs attention.

Pick a time for a working session or send a short brief. Either way, we will come prepared to understand where the work gets stuck.

Talk through the work

Book a 20-minute consultation.

Bring the workflow that feels slow or fragile. We will determine whether it is a sensible candidate for an AI system.

Bartosz LuderaBartosz LuderaFounder, Harnessloop

Choose a time for a 20-minute consultation.

Send a workflow brief

Prefer to write it down?

Tell us where work waits, repeats, or falls through the cracks.