Writing
Agents Need Protocols Between Each Other, Not Just Tools
Tool calling solves the inside of the loop. It says nothing about one agent reaching another built by a different team on another platform.
Writing
Tool calling solves the inside of the loop. It says nothing about one agent reaching another built by a different team on another platform.
Notes
For the last year, most agent architecture has been discussed from the inside out: give a model tools, connect it to data, let it call functions, and keep enough state around the loop for it to finish a task.
That solves only one side of the problem.
Google's Agent2Agent protocol, announced on April 9, treats another boundary as a first-class engineering problem: what happens when one agent needs another agent, built by a different team, running on another platform, with its own tools, memory and internal logic? Google launched A2A as an open protocol with contributions and support from more than 50 technology and services partners.
I think this is the more interesting question for large agent systems. If agents become useful in production, we probably will not build one enormous agent that understands every domain. We will build specialised systems and need them to cooperate without exposing all of their internals to each other.
MCP gave the industry a useful model for connecting AI systems to external capabilities and context. An MCP server can expose data sources and operations through a common interface instead of every AI application implementing another proprietary connector. Anthropic introduced it in November 2024 specifically as an open standard for connecting AI systems with external data and tools.
A2A is aimed at a different boundary. Google explicitly describes it as complementary to MCP: MCP provides agents with tools and context, while A2A is intended to let independent agents communicate and coordinate.
That difference matters.
A tool is usually passive. It has a defined operation, receives parameters and returns a result. The caller is expected to understand roughly what the capability does.
An agent can be opaque. It may have its own model, prompts, tools, data sources, policies and long-running state. The calling system may know what outcome it can request without knowing how that outcome is produced.
Treating such a system as one giant function call throws away part of what makes it useful.
One of the more interesting pieces of A2A is the Agent Card.
Google's launch specification describes a JSON document through which an agent can publish its capabilities and connection information. A client agent can discover that description and decide whether the remote agent is suitable for a task.
That sounds simple, but the abstraction is important.
In normal software integration, we tend to expose methods: create_invoice, search_candidates, submit_render_job. With an agent, the useful boundary can be closer to capability: "I can research candidates across these systems," or "I can manage this part of the asset pipeline."
The internal workflow does not need to be shared.
This is closer to how teams work. I can delegate a task because I know someone's responsibility and expected output. I do not need access to every intermediate decision they make.
If multi-agent systems scale, I suspect this distinction between capability and implementation will become increasingly useful.
The second part of A2A that stands out to me is its task model.
Google defines a task as a protocol object with a lifecycle. A remote agent can work on it over time, report status changes and eventually return an artifact. The protocol is designed for both quick exchanges and jobs that may take hours or days with human involvement.
That is a much better fit for real production work than pretending every agent interaction is a synchronous API request.
A render job is not a request-response interaction. Neither is a large repository migration, research task, asset conversion batch or localisation pass. These jobs have states: accepted, running, waiting for input, failed, completed. They may produce files rather than text. They may require another system to continue later.
Once agents operate on this kind of work, the communication layer has to represent those states explicitly.
The model can decide what to do. The protocol still needs to tell the rest of the system what is happening.
There is a natural tendency in AI products to keep adding capabilities to one agent.
Start with search. Add code execution. Add company files. Add browser control. Add CRM access. Add the production database. Soon the same model is carrying a large tool catalogue, several security domains and a growing amount of context.
That architecture becomes difficult to reason about.
Specialisation offers another route. A finance agent can own finance-specific tools and permissions. A development agent can own repository access and test infrastructure. A production agent can understand asset state and render operations. A user-facing coordinator can delegate work without receiving every internal tool definition from every domain.
A2A is interesting because it was designed for agents that do not need to share memory, tools or context. Google describes this explicitly as one of the protocol's design principles.
For production architecture, that can become a security boundary as much as a software boundary.
In a Blender, rendering or game-development pipeline, I would not split every small operation into a separate agent. That would create orchestration overhead without buying much.
But domain boundaries already exist.
An asset-management system has one type of state. Blender has scene state. A render farm has job state. An engine build system has another. A localisation pipeline may be maintained by a different service entirely.
Today I would normally connect those systems with APIs, scripts and queues. Agents can sit above those deterministic interfaces and interpret higher-level intent.
If those agent layers become independently useful, a common protocol between them starts to make sense.
For example, a production coordinator could ask an asset agent to prepare a group of models. That agent may use MCP tools or direct APIs internally to inspect metadata and validate files. It could then return an artifact and task status through A2A. The coordinator does not need to know whether the remote agent used Blender Python, a database query or another model to finish the work.
That separation keeps the integration surface smaller.
A standard message format does not make multi-agent systems reliable.
If one agent sends another a wrong result, the error can move across the system just as easily as a correct result. Capability discovery does not prove competence. A completed task status does not prove that the resulting artifact is valid.
The evaluation problem from single-agent systems therefore becomes larger, not smaller.
Every delegation needs a contract around expected output. Important artifacts still need deterministic checks. Permissions still need to stay close to the system that owns them. Long-running work needs timeouts, cancellation and observable status.
There is another issue: identity. Once agents from different platforms can request work from one another, authentication and authorisation stop being optional implementation details. Google says A2A was designed around enterprise authentication and authorisation patterns rather than inventing a separate security model.
I would treat remote agents more like external services than trusted coworkers.
The architecture I expect is becoming clearer.
MCP can standardise how an agent reaches tools and context. A2A can standardise how one agent delegates to another. The agent runtime manages reasoning, execution and state inside each system. Deterministic software continues to own permissions, source-of-truth data and validation.
Those layers solve different problems.
If this direction holds, agent engineering will start to inherit more ideas from distributed systems: discovery, routing, capability negotiation, task ownership, retries, status propagation, authentication and failure isolation.
That seems more realistic to me than expecting one model process to own an entire enterprise workflow.
OpenAI's Agents SDK points in the same architectural direction from inside a single application: it already includes multi-agent handoffs, orchestration and tracing rather than treating an agent as one isolated completion.
A2A is only two days old, and Google still describes the current specification as a draft, with a production-ready version planned later in the year. More than 50 technology and services partners supporting the launch is a useful signal, but it does not prove that the protocol will become the standard.
The larger idea matters even if A2A itself changes.
Agents need a common way to use tools.
Once there is more than one agent, they need a common way to work with each other too.
More