Skip to content

What is the A2A Protocol? The Agent-to-Agent Standard Explained

A deep dive into the A2A Protocol: how it solves multi-agent silos with Agent Cards, task lifecycles, and horizontal coordination alongside MCP.

Tuan Tran Van
11 min read
Contents (9 sections)
  1. Why Multi-Agent Systems Need an Open Protocol Over Custom Glue Code
  2. Opaque Execution and Autonomy Boundaries in A2A
  3. Agent Discovery via Standardized Agent Cards
  4. The Task-Oriented Model: Task Lifecycles, Messages, and Artifacts
  5. Communication Patterns and Wire Bindings: JSON-RPC, gRPC, and REST
  6. A2A and MCP: Horizontal Coordination Meets Vertical Tool Depth
  7. Production Challenges and Engineering Tradeoffs
  8. When Does Your Architecture Actually Need A2A?
  9. References

The Agent-to-Agent (A2A) Protocol is an open specification designed to standardize how independent AI agents discover capabilities, delegate tasks, and collaborate securely across organizational boundaries without proprietary vendor lock-in.

In modern enterprise architectures, AI systems are expanding rapidly but remain largely fragmented. A2A functions as a universal communication standard—much like SMTP did for email routing—allowing specialized agents across heterogeneous runtimes to exchange context, track asynchronous jobs, and negotiate tasks safely.

Governed by the Linux Foundation via the Agentic AI Foundation (AAIF), the v1.0 standard marks a fundamental transition away from closed "walled gardens" toward an open, federated network of specialized AI systems.

Backing from Google, Microsoft, Salesforce, ServiceNow, and Tyk gives the standard real industry weight. Under the hood, the protocol buffers definition in a2a.proto is the sole normative source for the specification, while JSON Schemas are generated as derivative artifacts for web clients. That separation guarantees deterministic data structures and strict typing across polyglot systems in Go, Python, and Java.

Federated network of autonomous AI agents communicating via open A2A Protocol

Why Multi-Agent Systems Need an Open Protocol Over Custom Glue Code

Building production multi-agent systems without an open standard results in severe technical debt. Engineering teams typically author bespoke glue code to connect agents running on diverse frameworks such as LangChain, CrewAI, or PydanticAI. When an enterprise attempts to coordinate automated code review, financial audits, and HR dispatchers, maintaining individual point-to-point connections quickly becomes unsustainable.

Comparison between point-to-point custom glue code and open A2A Protocol standard

Custom integrations also create tight vendor lock-in. Organizations risk coupling their core operational workflows to a single SDK vendor or proprietary orchestration cloud. If a model provider changes APIs or a new specialized third-party agent becomes available, refactoring brittle glue code requires substantial engineering investment. A2A addresses this by defining a vendor-neutral interface where agents communicate via standardized message semantics rather than private internal contracts.

A federated multi-agent network decouples agents along architectural boundaries. Systems can dynamically locate and route requests to specialized peers based on declarative capability manifests. This modularity reduces coordination friction, preserves loose coupling, and enables compound AI architectures to scale predictably across corporate divisions.

Opaque Execution and Autonomy Boundaries in A2A

The guiding design tenet of the A2A Protocol is Opaque Execution (Section 1.2). Collaborating agents interact purely on the basis of published capabilities and skills, with zero access to each other's internal state machines, system prompts, long-term memory stores, or local tool configurations. This boundary protects sensitive organizational IP and isolates operational risk across the multi-agent dependency chain.

Agent autonomy is preserved through decoupled reasoning. When an orchestrator delegates a subtask to an A2A agent, that agent retains full discretion over its internal planning, tool invocations, and reflection loops. It only returns agreed-upon deliverables (Artifacts) and contextual status updates. Concealing intermediate "thought streams" (chain-of-thought) is critical in enterprise settings where cross-departmental boundaries, client confidentiality, and security boundaries must be strictly enforced.

This separation of concerns guarantees data integrity. A third-party security auditing agent can evaluate a configuration and return a compliance report without requiring access to production databases or proprietary system prompts. A2A establishes a contract where opaque actors trust one another strictly through verified schemas and explicit data payloads.

Agent Discovery via Standardized Agent Cards

Before agents can collaborate, they must discover peer identities and capabilities. A2A standardizes this through an Agent Card, a JSON manifest hosted at the well-known URI /.well-known/agent-card.json (per RFC 8615). The manifest outlines metadata such as name, description, supported transport bindings (supportedInterfaces), operational feature flags (capabilities), and a catalog of exposed skills[].

Agent Card discovery flow and JWS cryptographic verification in A2A Protocol

Each entry in skills[] documents accepted input and output media types (inputModes, outputModes) alongside concrete execution examples. From a governance standpoint, a root-level security array maps specific scopes defined in securitySchemes directly to individual skills, enabling fine-grained role-based access control. To prevent spoofing and man-in-the-middle tampering, Agent Cards support AgentCardSignature using JSON Web Signatures (RFC 7515) calculated over payloads formatted according to the JSON Canonicalization Scheme (RFC 8785).

Here is an example of an A2A v1.0-compliant Agent Card for an enterprise flight booking agent:

json
{
  "name": "Corporate Flight Booking Agent",
  "description": "Handles domestic and international flight reservations.",
  "version": "1.0.0",
  "supportedInterfaces": [
    {
      "url": "https://flights.internal.corp/a2a",
      "protocolBinding": "JSONRPC",
      "protocolVersion": "1.0"
    }
  ],
  "capabilities": {
    "streaming": true,
    "pushNotifications": true,
    "extendedAgentCard": false
  },
  "skills": [
    {
      "id": "book_flight",
      "name": "Book a flight",
      "description": "Books a flight given origin, destination, and dates.",
      "tags": ["travel", "booking"],
      "inputModes": ["application/json"],
      "outputModes": ["application/json"]
    }
  ],
  "securitySchemes": {
    "corp_oauth": {
      "type": "oauth2",
      "flows": {
        "clientCredentials": {
          "tokenUrl": "https://auth.corp/token",
          "scopes": {
            "flights:write": "Book flights"
          }
        }
      }
    }
  },
  "security": [{ "corp_oauth": ["flights:write"] }]
}

The Task-Oriented Model: Task Lifecycles, Messages, and Artifacts

Every collaborative workload in A2A is structured as an explicit Task governed by an unambiguous finite state machine. A major evolution in the v1.0 specification is the removal of the legacy kind discriminator and the retirement of the final: true boolean in favor of closing the network stream once a task enters a terminal state.

Task lifecycle state machine and streaming artifacts in A2A Protocol

The lifecycle of an A2A task progresses through several deterministic states:

TaskState (v1.0)ClassificationTechnical Meaning
TASK_STATE_SUBMITTEDIn-flightThe task has been accepted by the server and enqueued for execution.
TASK_STATE_WORKINGIn-flightThe agent is actively processing and generating progression events.
TASK_STATE_INPUT_REQUIREDInterruptedExecution is suspended awaiting client parameters or clarifications.
TASK_STATE_AUTH_REQUIREDInterruptedExecution is halted pending credentials or token refresh.
TASK_STATE_COMPLETEDTerminalThe task finished successfully; final Artifacts are ready.
TASK_STATE_FAILEDTerminalAn unrecoverable execution error occurred (JSON-RPC -32xxx codes).
TASK_STATE_CANCELEDTerminalThe task was terminated upon request via CancelTask.
TASK_STATE_REJECTEDTerminalThe agent refused the task due to rate limits or security policy.

A2A draws a clean distinction between Messages and Artifacts. A Message is a conversational entity composed of multimodal Part blocks (text, media URIs, or JSON structures) exchanged dynamically between the user and agent. An Artifact represents a tangible, final outcome produced by the task (such as a generated spreadsheet or PDF contract). For streaming interactions, the protocol emits ArtifactUpdateEvent with append: true and lastChunk flags, allowing consumers to process large binary payloads incrementally without buffering entire files in memory.

Communication Patterns and Wire Bindings: JSON-RPC, gRPC, and REST

The A2A specification is organized into three distinct tiers:

  1. Layer 1 (Canonical Data Model): Core domain entities (Task, Message, Part, Artifact) modeled via Protocol Buffers (a2a.proto).
  2. Layer 2 (Abstract Operations): High-level operational semantics (SendMessage, GetTask, CancelTask, SubscribeToTask).
  3. Layer 3 (Protocol Bindings): Transport-level implementations over real network stacks.

Because of this architectural separation, clients receive semantic equivalence regardless of the selected wire binding. The standard binding for web and public internet traffic is JSON-RPC 2.0 over HTTPS. In v1.0, the core interaction method is named SendMessage (replacing the legacy message/send convention). Protocol versioning and optional capabilities are communicated strictly via HTTP headers like A2A-Version and A2A-Extensions:

http
POST /a2a HTTP/1.1
Host: agent.internal.corp
Content-Type: application/a2a+json
A2A-Version: 1.0
Authorization: Bearer <token>
 
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "SendMessage",
  "params": {
    "message": {
      "role": "ROLE_USER",
      "parts": [{ "text": "Analyze Q1 revenue data." }]
    },
    "configuration": { "returnImmediately": true }
  }
}

For real-time progress and intermediate event delivery, A2A relies on Server-Sent Events (SSE) initiated via SubscribeToTask. For long-running asynchronous workflows that outlive standard HTTP socket timeouts, the protocol supports Push Notifications (Webhooks), allowing the server agent to POST StreamResponse payloads back to a registered callback URI. In low-latency internal microservice fabrics, gRPC bindings provide serialization speed and type-safe stubs compiled directly from a2a.proto.

A2A and MCP: Horizontal Coordination Meets Vertical Tool Depth

A frequent misconception in the developer ecosystem is that A2A Protocol and the Model Context Protocol (MCP) are competing alternatives. In practice, they operate on orthogonal architectural axes and are designed to be deployed together:

Dual-axis architecture: Horizontal A2A coordination and vertical MCP tool depth

  • MCP operates vertically (Vertical Axis): Connects an individual agent host to underlying databases, local filesystem primitives, and developer tools (Client-to-Tool).
  • A2A operates horizontally (Horizontal Axis): Connects autonomous agents across distributed systems as sovereign peers (Agent-to-Agent).

Consider the workflow in an Auto Repair Facility:

  • The Shop Manager Agent coordinates over A2A to take customer requests and negotiate work orders with peer agents.
  • The Mechanic Agent receives the assigned task and uses MCP to query specialized local interfaces, such as an OBD-II diagnostic tool or hydraulic lift controller.
  • After discovering a faulty catalytic converter, the Mechanic Agent uses A2A to reach out horizontally to a remote Parts Supplier Agent operated by an external distributor.
DimensionA2A ProtocolModel Context Protocol (MCP)
Primary ScopePeer collaboration across independent agentsContext and tool provisioning for a single agent
Architecture AxisHorizontal (Agent-to-Agent)Vertical (Agent-to-Tool / Host-to-Server)
Execution PatternAsynchronous, long-running task lifecyclesSynchronous or streaming request/response
Transport StackHTTPS (JSON-RPC), gRPC, Webhooksstdio (local pipes), HTTP SSE
State BoundaryOpaque execution; shares only context and artifactsTransparent tool schemas and structured data

Production Challenges and Engineering Tradeoffs

When deploying A2A in production, engineering teams run into a few practical gaps in the v1.0 spec:

  • Custom skill payloads lack native schemas: While A2A defines task containers, it leaves skill argument validation open. To prevent parsing failures across independent vendor stacks, authors document input schemas inside the skill's description or adopt custom schema extensions (Section 4.6 of the specification).
  • Delegation creep in multi-hop chains: When Agent A delegates to Agent B, which then calls Agent C, passing ambient user credentials introduces privilege escalation risks. Production setups use OAuth 2.0 Token Exchange (RFC 8693) alongside Rich Authorization Requests (RFC 9396) to downscope security tokens at each hop, limiting downstream agents to least-privilege operations.
  • Distributed tracing across boundaries: Tracking an asynchronous task executed across corporate boundaries requires disciplined telemetry. Propagating the standard W3C Trace Context header (traceparent) on every outbound A2A request is mandatory to reconstruct OpenTelemetry traces when an agent down the line fails.
  • Standardized fault handling: Clients must parse standard JSON-RPC error codes (such as -32001 for TaskNotFound) rather than generic HTTP 500s, letting callers distinguish temporary timeouts from permanent policy rejections.

When Does Your Architecture Actually Need A2A?

Not every AI use case warrants protocol-level overhead. If you are developing a self-contained application with a single agent using standard prompt chaining, in-process frameworks are sufficient. A2A becomes an architectural requirement under the following conditions:

  • Cross-Organization or Cross-Team Handoffs: Autonomous agents owned by separate departments or external vendors must collaborate without sharing internal codebases.
  • Polyglot Agent Fleets: You need to replace or maintain micro-agents built in different programming languages without rewriting coordination logic.
  • Durable Asynchronous Operations: Workflows involve long delays, human-in-the-loop validation, or complex multi-step tasks requiring verifiable state recovery.
  • Enterprise-Grade Capability Discovery: Your ecosystem comprises dozens of specialized agents whose skills must be located and queried dynamically via Agent Cards.

In enterprise architectures, deploying A2A agents behind a managed API Gateway (such as Tyk or Envoy) is an operational best practice. The gateway manages transport-level concerns—including mTLS encryption, token verification, signature auditing, and rate limiting—freeing agent implementations to focus entirely on LLM reasoning and domain logic.

References

Share this article