ACP the shared surface for coding agents
by John Robinson @johnrobinsn
If you have tried to programmatically drive an AI coding agent from your own program — a CI pipeline, an orchestrator, a client that is not the vendor's blessed UI — you have felt a specific pain. You spawn Claude Code as a subprocess, parse its streaming output, handle its permission prompts, capture its file operations. That works. Then you want to also drive Codex CLI. Different stream format. Second adapter. Then Grok Build, or opencode. Each CLI streams output in its own JSON shape, exposes permissions through its own schema, and manages sessions with its own semantics. Your wrappers have to implement each vendor's take on all of them.
For models, this problem was settled a while back. Whatever LLM you are using, openai.chat.completions.create(...) covers most hosted APIs; HuggingFace Transformers covers most open-weights ones. Providers ship their own SDKs, but they have converged on similar-enough patterns that swapping models is mostly a matter of API keys and parameter names.
Agents did not converge until August 2025, when Zed Industries shipped the Agent Client Protocol (ACP) — a JSON-RPC 2.0 wire, spoken over stdin/stdout, between a client and an agent. Persistent subprocess. Bidirectional. Session-based. Zed's stated goal was to let any editor drive any coding agent without writing a per-vendor wrapper. Enough of the ecosystem showed up on both sides — Zed, JetBrains, and Cursor on the client side; Claude Code, Codex CLI, Gemini CLI, opencode, Kimi CLI, Goose, and Cline on the agent side — that the standardization stuck. (A similar move happened for language tooling in 2016 with LSP; ACP borrows its shape directly.)
ACP was born for IDE integration. But the standardization applies wherever you want to programmatically drive an agent — a meta-harness, a batch orchestrator, a novel client shape. And the wire is not just text: it is multimodal by default, which matters more than it first sounds.
This piece is about what ACP is, why it matters, and where the interesting layer above it still is not built.
Why it stuck #
Editors had the obvious version of this pain. Zed did not want to write a bespoke wrapper for every coding agent that shipped, and JetBrains did not want to either, and neither did Cursor. But agent vendors had the mirror-image problem. If Anthropic wanted Claude Code to be usable from every editor, they had to lobby each editor's maintainers individually to build integration. Same for OpenAI with Codex, Google with Gemini. The M×N pain lived on both sides.
That is why ACP took hold without a formal joint-standards process. Zed published the wire; Google made Gemini CLI speak it natively; Anthropic bridged Claude Code via an official adapter (claude-agent-acp); Codex got codex-acp; the community wrote adapters for opencode, Kimi CLI, Goose, Cline, and others. LSP grew the same way over 2016–2020 — no consortium, just each side adopting once the payoff became obvious. What matters is that the wire is stable and public, and enough of the ecosystem shows up on both sides that new participants have an obvious thing to conform to.
The multimodal surface — the part not enough people are talking about #
The other thing to know about ACP — the part that has not been talked about enough — is that it is multimodal by default. Messages are not strings. They are ordered arrays of ContentBlocks, and a ContentBlock is one of five types:
text— plain text, the baseline every agent must supportimage— an image, with base64-encoded bytes and a MIME typeaudio— audio, same shaperesource— a full embedded resource (a file's contents, inline, with URI + MIME type)resource_link— a pointer to a resource by URI
This is not text-with-attachments. It is not a bag of blobs stapled to a chat message. The protocol is multimodal at the primitive level. A user prompt can be [text, image, resource_link] interleaved in whatever order the semantics demand. An agent's streaming response can carry [text-chunk, image, resource-with-diff-content, text-chunk] — the client sees each block arrive in order, and can decide what to do with each one.
A concrete example — a session/prompt from a client that wants the agent to look at both a screenshot and a log file at once:
{
"method": "session/prompt",
"params": {
"sessionId": "demo-session-01",
"prompt": [
{"type": "text", "text": "Look at this and the attached log..."},
{"type": "image", "mimeType": "image/png", "data": "<base64>"},
{"type": "resource_link", "uri": "file:///var/log/build.log",
"name": "build.log", "mimeType": "text/plain"}
]
}
}
An agent response could stream back a similar mix: text explanation, an SVG diagram of the call graph, an embedded patch file. All standard content blocks. All in one session. All part of the wire.
If you have not thought about coding agents this way — if the mental model has been chat interface that produces code — this is the moment to update it. The wire has room for the model to hand you back visuals, audio, generated files. The wire has room for you to hand the model images and audio and files as input. Everything below is present-tense capability, not roadmap.
Why the multimodal surface matters in the medium term: the frontier work in agent capability — vision models that read screenshots, audio input that carries prosody the LLM can reason about, generated diagrams as an output modality — has a wire protocol ready for it. Nobody has to redesign the transport when GPT-6 (or whatever) starts reasoning about screen state as a first-class input. ACP already carries it.
What this unlocks for meta-harnesses #
The ACP payoff for editors is obvious — write your ACP client once, get every ACP-compatible agent. But the interesting layer sits one level up.
If you are building anything that drives coding agents from above — a CI system that dispatches agent tasks, a batch orchestrator that fans work out across agents, a novel client shape (voice-driven, multi-device, ambient) — you now have a substrate. You do not have to write vendor-specific integrations. You do not have to reverse-engineer subscription auth for each provider. You do not have to keep up with each vendor's session-state semantics. You spawn the adapter as a subprocess and speak ACP to it. The adapter (claude-agent-acp, codex-acp, native gemini --acp) fights the vendor-compatibility battle for you.
That is a materially different position than eighteen months ago. Building a meta-harness in early 2025 meant writing a stack of vendor-specific glue code, each piece of which broke every time the vendor changed anything. Building one now means writing an ACP client, and the vendor-specific glue is somebody else's problem.
Same protocol, different client shape. An IDE renders the same content blocks into editor panels, inline previews, multi-buffer diffs. A voice-first client could route text to TTS, images to a paired screen, diffs to a phone card. The wire does not care which one you build.
The uncontested frontier #
ACP defines a lot. It does not define everything. Some things worth naming:
ACP does not define cross-agent or cross-device continuity. Within a single agent, ACP handles resumption cleanly — session/load lets you pick up where you left off. What ACP does not define is a "session" that spans multiple agents (Claude Code for the edit, Codex CLI for the review), or that follows you between devices, or that lives in the client rather than the agent.
ACP does not define multi-agent coordination. One ACP connection is one client talking to one agent. If you want to dispatch a task across three agents in parallel and reconcile their outputs, you build that. ACP is the wire for each individual connection.
ACP does not define cross-modal state or user context. Which projects you have going. What your focus rules are. Which trusted actions the agent can take without asking. What device is closest to you right now. What was decided in a previous session. All of that lives above ACP, in whatever the client chooses to build.
That is the interesting layer, and it is uncontested. Editors are competing in the visual-IDE-client slot. Nobody is competing for the other client shapes yet. Nobody is competing for the state-and-context layer above the wire. Both of those are wide open.
The takeaway #
If you are building agent orchestration in 2026, the honest advice is: invest in ACP. Do not write vendor-specific wrappers. Do not roll your own subprocess protocol for driving Claude Code or Codex CLI. Speak ACP, spawn the adapter, let the ecosystem do the compatibility work.
What you build on top of ACP — session persistence, multi-agent coordination, non-IDE client shapes, cross-modal routing, user context — is where the differentiated value lives. That layer is not going to be defined by a wire protocol committee. It is the product design. It is the thing that makes a coding-agent experience feel like yours.
That is what I am building. Vibr8 is a voice-first, multimodal meta-harness that uses ACP to drive and orchestrate multiple coding agents from a single surface. The bigger frame it lives inside — a modality-adaptive interface that routes by circumstance rather than by app — is the argument in AI will be your only UI. ACP is the substrate that made the voice + multi-agent layer tractable as a solo project rather than a team-scale integration effort.
The wire is standard. What you do above it is the frontier.
Interested in being a beta tester when Ring Zero (vibr8) is released? Sign up for the mailing list now.
Follow along at storminthecastle.com — or reply if you're building in this space.
Share on Twitter | Discuss on Twitter
John Robinson © 2022-2026
- Previous: Towards an always-listening voice agent