Multi-Agent Orchestration


Picture a multi-agent demo that runs fine on your laptop. Then it meets a remote tool server and a load balancer. Sessions vanish. Nobody can say which agent botched the call.

Search for multi-agent orchestration frameworks 2026 and half of what you find is hype. Here I stick to what’s dated and documented. Older ideas get called older.

Why Multi-Agent Setups Break at the Seams

Most failures sit between the parts. Agent to tool. Agent to agent. The model itself is rarely the weak spot. Early setups kept state inside long connections and passed messages in whatever format each team picked. One team, one server, no trouble. Then traffic spreads out and it hurts.

Take tool servers. MCP, the Model Context Protocol, is an open standard for hooking agents up to tools. Older versions opened with a handshake and issued a session ID. That tied a client to one server instance. Agent-to-agent calls had the same flavor, just lots of one-off connections that never matched.

What Actually Changed in 2026

Three changes are documented well enough to trust. A2A, a standard for agents calling agents, reached 1.0 in March. MCP dropped protocol sessions in July. OpenTelemetry added agent spans, though they’re still a draft. All three trim glue code. None of them designs your system for you.

Tool servers lost their sessions

The July 28, 2026 MCP spec threw out the handshake and the session header. Each request now carries its own version and capabilities. Any instance can answer it, even behind a plain round-robin load balancer.

Long jobs changed as well. Tasks had been experimental since the 2025-11-25 revision. They’re now an official extension. A server returns a task handle, and the client polls with tasks/get. If you really need state, create an explicit handle and have the model pass it back.

Migration will cost you something. The maintainers say so themselves, mostly for code that leaned on session IDs.

Agent-to-agent calls got a stable spec

A2A isn’t new. Google announced it in April 2025 and handed it to the Linux Foundation that June. What’s new is v1.0, out in March 2026, the first stable release.

It brings signed Agent Cards for checking identity, plus multi-tenancy. Agents on different platforms can pass sub-tasks around without sharing memory. Rough rule: A2A is agent to agent, MCP is agent to tool. A registry story is still on the roadmap, so don’t count on it yet.

Traces learned agent vocabulary

Normal tracing shows HTTP calls. It says nothing about an agent planning or running a tool. OpenTelemetry’s conventions now name those steps: invoke_agent, invoke_workflow, plan, execute_tool. They’re marked Development, so names may still move.

Prompt and response text is opt-in, since it can carry personal data. MCP also describes passing trace context in request metadata. And check your own framework. It has to emit these spans before you’ll see them.

One Workflow, Four Responsibilities

Keep four jobs separate. Orchestration decides the order of steps. Messaging moves events around. The model writes text. Your application remembers facts. A made-up refund flow shows it. It’s hypothetical, not a customer story.

A user asks for a refund in chat. Messaging carries that in. An intent agent runs, then a search over the refund policy. The model provider drafts a decision. Next, a billing tool called over MCP issues the refund, and if it’s slow, it hands back a task handle. The refund ID goes into your database. The model shouldn’t own that fact. All along, progress events stream to the user and each step lands in a trace.

Where DNotifier Fits

DNotifier provides one SDK and one API with multi-model support. In the refund flow it covers orchestration, messaging, and observability. Your billing tool, your database, and your model provider stay yours.

Start with messaging. The SDK does real-time messaging over WebSocket and HTTP, and the docs mention pub/sub-style delivery. That handles the request coming in and the updates going out.

Orchestration next. You define named agents and a workflow, then run it with runWorkflow. Agents share a state map for one execution. A step can call semantic search over a knowledge base, which suits the policy lookup. The docs also list provider and model selection, so the drafting step can use whichever model you pick.

For tracing, turn on observability: true and steps appear in a workflow dashboard under an execution ID. Add logs: true and the SDK records sessions and token usage.

Some things stay on your side. That shared state lasts one execution only. The public docs don’t describe durable checkpoints or A2A support, so keep anything you need to recover in your own store. The site does list MCP servers among connected sources. Check the docs for setup before you plan around it.

Trade-Offs and Mistakes

Standards cut glue code but bring migration work and some governance chores. Two mistakes come up a lot. People treat drafts as final. And they park state in the wrong layer.

Plan the MCP upgrade early. Breaking changes are in there, and Roots, Sampling, and Logging are deprecated, with at least twelve months of support left. Pin your trace attribute names too. They’re still in Development.

A linear job doesn’t need five agents. Every handoff is another place to fail. Keep business facts out of prompts and sessions, and put them in a database. Leave prompt content out of traces by default. Turn it on on purpose.

FAQs

Did MCP make applications stateless?

No. The protocol dropped sessions, but your application can still carry state. The usual way is an explicit handle that the model passes between tools.

Do I need both A2A and MCP?

Often, yes. They answer different questions. A2A lets agents talk to each other, and MCP connects agents to tools and data.

Are OpenTelemetry’s agent conventions stable?

Not yet. They’re marked Development, so pin the attribute names you emit and expect some change.

One Thing to Take Away

Standards are absorbing a lot of the plumbing. That leaves you the design questions. Who orchestrates? Who owns state? Where do traces go? Settle those first, tools second.

If you want to see orchestration, messaging, and observability in one SDK, take a look at www.dnotifier.com.


Leave a comment