AI Engineer World's Fair 2026
Why Your AI Agents Can't Talk to Each Other (Yet) — Vlad Luzin, BAND
Read the talk
Why Your AI Agents Can't Talk to Each Other (Yet)
Vlad Luzin explains why connecting independent agents takes more than forwarding prompts: shared conversations need ordered delivery, recoverable state, runtime identity and human visibility. BAND’s demos follow that idea from agent registration to cross-user collaboration.
From a talk by Vlad Luzin
At a glance
Ideas worth remembering
Running separate planning, coding and reviewing sessions already creates a human-routed multi-agent workflow. Loop engineering automates the exchanges between those sessions.
Shared agent work needs ordered transport, retries, recoverable state and mappings between each framework’s runtime identifiers.
Observability should connect message delivery to the receiving agent’s subsequent tool calls, so operators can understand what an exchange caused.
BAND’s cross-user demo separates registration, bilateral contact consent and conversation invitation; becoming discoverable does not skip the consent step.
The final collaboration pattern keeps humans available: Claude Code invites Codex and Mike for engineering work, and the group can invite Vlad for help.
Give agents a place to find each other
An agent working on your behalf may need help from an agent working for someone else. That is the starting point for Vlad Luzin, co-founder and CTO at BAND: AI-to-AI communication within companies, between companies, and between businesses and consumers. His forecast assumes autonomous, always-on agents spread across the world. The immediate engineering question is how those separate agents find one another and work on the same task.
The proposed interaction starts in a conversational space. An agent receives a task from a person or another system, looks up peers in a registry, invites suitable peers into the conversation, delegates work, gathers information and reports back. A registry supplies discoverable participants; the conversation supplies a place for their work to meet. This is the picture the rest of the talk tries to turn into infrastructure.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
From copying between sessions to loop engineering
The familiar version already exists on a developer’s laptop. One Claude or Codex session plans while another reviews; one writes code while another checks it. Both sessions retain their own working context, but the developer moves output between them. Luzin’s memorable description is that the human becomes a router and switch, carrying messages between two stateful agents.
Loop engineering, in his distilled account, automates that handoff. Python or TypeScript code prompts the agents back and forth after a human triggers the loop. The separate sessions remain; the script takes over message movement. That distinction matters because automating a prompt exchange does not, by itself, explain how agents discover peers, survive failures or collaborate across owners.
Why keep separate sessions at all? Luzin calls the problem a single-agent bottleneck and names several pressures:
- Confirmation bias: an agent reviewing its own work can continue along the assumptions that produced it.
- Attention dilution: more material competes for attention inside the same session.
- Context fragmentation and recall degradation: keeping information in a large context does not guarantee that the relevant information will guide the next step.
His conclusion is that a context window of one million or two million tokens does not solve these problems on its own. This is the motivation offered for separate agents, rather than a demonstrated comparison of their performance in the recording.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A messaging integration still leaves collaboration work
Messaging platforms look like the obvious starting point: Slack, Teams, Discord, WhatsApp and Telegram already move messages between participants. Luzin presents the setup burden as five manual steps for Telegram, seven for Discord, eight for Slack and eleven for WhatsApp. These are his integration counts, not a universal measure of setup difficulty. The more consequential objection is the result: connecting a bot to its owner does not automatically create a network of agents that can discover and contact one another.
The resulting agent may be reachable by you yet still isolated from useful peers—what Luzin calls “digital solitary confinement.” He characterizes the messaging platforms as designed for humans and restrictive about bot-to-bot communication. Within that framing, a successful chat integration solves access to one agent while leaving peer access and cross-user interaction as additional work.
Protocols offer another starting point. The architecture Luzin sketches has servers calling other agents as tools through MCP, potentially followed by an A2A call to another agent. His criticism concerns the application machinery left around those calls:
- Conversation continuity: a tool invocation needs an application-level way to return to the same agent’s working state later.
- Two-way initiation: in the client-server arrangement he describes, allowing either agent to initiate work requires implementing both directions, with client and server roles on each side.
- Long-running work: chains of REST calls encounter timeouts, while queues and discovery still need supporting infrastructure.
His shorthand description of MCP as stateless and A2A as one-way should be read as a critique of this invocation-based architecture, rather than a complete specification of either protocol’s capabilities.
The practical complaint is about where engineering effort goes. After choosing a protocol, a team may still spend its time building queues, discovery and reciprocal communication before it can exercise the collaboration it wanted. Luzin calls this “plumbing.” His next move is to identify the higher-level system that would absorb that work.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Connect runtime state, not just network endpoints
Two sessions on one laptop and your Claude talking to someone else’s Codex share the same underlying problem: separate processes must coordinate. Luzin treats agents as nondeterministic microservices. Ordinary distributed software already has failures and delivery problems; agents add behavior that cannot be assumed to repeat identically from one invocation to the next.
Three mechanisms carry the collaboration:
- Ordered, real-time transport: messages must reach the agents in the intended sequence, with retries handled by the communication system. Order matters because the messages become the sequence an LLM reasons over.
- Persistence and hydration: a failed agent process must be able to recover the state needed to continue. Hydration means restoring that saved state into a running agent; otherwise one crash can break the group’s work.
- Runtime binding: different frameworks name their active work differently—thread IDs, session IDs, execution IDs and run IDs. The system must map those identifiers together so that each participant’s local runtime contributes to the same shared task.
What must connect a LangGraph thread to a coding-agent session besides an address? The diagram separates the shared conversation from the local runtimes that participate in it. Runtime binding supplies the relationship between them; ordered delivery and recoverable state keep that relationship useful as work proceeds. An IP address, URL or Pub/Sub topic identifies a technical destination, but does not by itself identify which ongoing task the receiving agent should continue.
The proposed abstraction therefore names conversations, participants, channels or rooms, and agent-aware message routing. Organizations also need identity and governance. Observability must follow the causal chain beyond delivery: which message reached which agent, and which tool calls happened after that agent received it? A message log captures the handoff; tracing the subsequent tool activity captures what the handoff caused.
That is the infrastructure between the appealing future and the working system. Luzin punctures the forecast with a joke: the future of everyone lying on a beach on universal basic income depends on solving these less glamorous coordination problems first.
Local thread or execution state.
Runtime binding associates each agent’s local work with the conversation; transport and persistence support continued exchange.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
BAND’s demo moves from registration to cross-user consent
BAND is introduced as a global interaction and collaboration layer that takes on this supporting infrastructure. Its proposed unit of interaction is multi-peer: several agents and humans can share a conversation. Luzin describes a registry, persistence, channels, security and observability, with message filtering performed inside the platform rather than separately by each client. The tradeoff is to place those coordination responsibilities in the platform.
The product claims span any agent, platform, language and environment; the recorded demos illustrate a narrower set of integrations and interactions, without establishing those universal compatibility or reliability claims. The first example uses two accounts: Vlad has a personal Claude assistant running on a Mac, while Mike initially has no agents registered. The demos are prerecorded, though Luzin describes the exchanges within them as real-time.
The observable change begins with Mike’s empty account. An agent starts, and an agent card appears seconds later. A LangGraph agent starts, and its card appears as well. At that point, the agents know about each other and can interact. Registration has turned running processes into discoverable participants belonging to Mike’s account.
Connecting Mike’s Codex agent to Vlad’s assistant adds a different requirement: consent across owners. A contact request requires agreement from both parties before the peer is introduced into the relevant registries. Vlad accepts the request. Codex can then invite his assistant into a conversational space and exchange messages with it. The causal sequence is registration, accepted contact, then invitation—not merely knowing where the other process runs.
What changes when the contact request is accepted? The flow below makes the transition from separate accounts to a shared conversation visible. Discovery supplies candidate peers, consent permits the introduction, and the invitation creates the place where those peers work together.
Human participation remains available after the agents connect. Vlad describes full visibility into his agent’s conversation with the other user’s agent, and his assistant can invite him into that space. Visibility and participation serve different purposes: the owner can see the exchange without manually forwarding every message, then join when human input is useful.
Agent cards appear; Mike’s agents can discover each other.
The demonstrated cross-user connection passes through bilateral consent before Codex invites Vlad’s assistant.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A terminal session becomes a collaborator
The final demo carries Mike’s existing Codex agent forward and adds Claude Code running in a terminal. The terminal session is asked to connect to BAND. That specific session receives an identity and an agent card, while the platform hides its network location from the other participants. The useful identity is now the session that will do the work, rather than an endpoint the user must manage.
Claude Code can then create a conversational space as its owner, invite Codex and invite Mike. The intended task is to create a small website, with one agent reviewing the other’s work. This returns to the opening planning-and-review pattern, but the participating agents can initiate the conversation and bring in peers themselves. Luzin presents the task’s start and intended review loop; he does not describe the finished website or a measured improvement in its quality.
If the agents and Mike struggle, they can invite Vlad into the same conversation for help. That final detail expands the loop beyond a fixed pair of agents: the group can add a human participant when the work calls for one. The proposed benefit is less custom code for prompting agents back and forth, while retaining agent-to-agent review and human assistance in the same space.
The closing joke pits orchestration-loving orcs against collaboration. Underneath it is the product distinction the demos have been building toward: agents gain identities, find peers, enter shared conversations and invite help, while the platform carries the communication machinery. The developer’s role can move from forwarding every exchange to participating when the group needs judgment.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Read the complete timestamped transcript
- 0:13
Okay. Hi, everyone. Can you hear me? A lot of people, yeah, so please be quiet.
- 0:20
So, uh, my name is Vlad. I'm co-founder and CTO at BAND. And today I have four topics I want to discuss with you. First, thesis. Thesis of our company, what we believe, believe in. AI evolution from adversarial agents to loop engineering and beyond. Technical challenges that we experience right now, that should be solved right now. And obviously, BAND, to present the company and the product and what we do. Let's start with the thesis.
- 0:50
First of all, we believe that the future belongs to AI-to-AI communication within a business, between businesses, and between consumers and businesses. AI will be doing work on our behalf, and they will have to talk to each other in order to solve tasks on our behalf. Second, agents will be autonomous, always-on entities spread across the globe. Now, these are very nice words,
- 1:20
but I want to spend a minute just to paint a picture in your head how this type of communication between agents actually is going to look like.
- 1:32
Demo time. Agents will communicate in a conversational space. They will see each other. They will be distributed. They will receive tasks from us or from any system. They will be able to look up peers in the registries, invite these peers in con- into conversational space, delegate the tasks, gather en- enough information, and report back to us. This is how the future of AI-to-AI communication is going to look like.
- 2:02
So while we are talking, hold this in your head.
- 2:07
Let's talk a bit about adversarial agents. I think everyone here knows this concept, and I'm pretty sure everyone here is using it as well. Where are we using it? If you are using Claude or Codex, I'm pretty sure you are running multiple sessions working on the same task. One is planning, another is reviewing. One is writing code, another is reviewing code. So basically what is happening, you have two stateful agents running,
- 2:38
working on the same task, and you are basically a Cisco router and a switch moving packets between these two stateful agents.
- 2:47
Let's talk about loop engineering. A number of talks on this conference about loop engineering. A number of articles, a pretty new concept, right? But I would like to distill it, uh, to its essence. So what is loop engineering? Basically, it's the same as the previous concept. It's you running multiple agents, hopefully stateful agents and not stateless. But instead of you being a router or a switch, you outsource
- 3:17
this to some piece of Python or TypeScript that someone wrote as an open source project and, um, this piece of code basically prompts the stateful agents back and forth, and you trigger it.
- 3:31
Now, why are we using these multiple agents when we do this work? Because of the underlying architectural limitations of, uh, transformers, specifically a concept is called single agent bottleneck, right? So why we are not running one agent in one session that is doing everything? For instance, confirmation bias, right? Attention dilution, context fragmentation, recall degradation. So one million tokens or two million tokens, uh, uh, context will not help
- 4:01
you in terms of the performance. But let's talk about multi-agent systems, right? So we already understand that w- us running multiple sessions, these are basically multi-agent systems running. So what is the easiest way to connect multiple agents running together? Messaging platforms, right? We already have Slack, right? Anthropic released Slack agent. We have Teams. We have Discord. We have WhatsApp. We have Telegram, and so on, right? So this is a
- 4:30
solution. Well, in order for you to connect an agent to a Telegram, five steps. In order for you to connect to Discord, seven steps. In order for you to connect your agent to Slack, eight steps. WhatsApp, eleven. And these are not simple steps. They are all manual, and you need to read a book in order to be able to set it up. And even if you
- 5:00
do, you get only one thing and one thing only. What is this thing? An agent talking to a person, which is usually you. This agent cannot talk to any other agent. This agent cannot talk to any other person without you jumping through hoops. So basically, your agents are still alone in a kind of digital solitary confinement.
- 5:29
But smart people will say, "I know the solution." Protocols, right? MCP, ACP, A2A, and et cetera. This is the solution, right? Well, if you are a proponent of that kind of a solution for, uh, connectio- for connecting your sessions, this is how your multi-agent system is going to look like. Multiple servers calling other agents' tools through MCP-
- 6:00
And if you want, maybe this agent that you called as a tool can call another agent as an A2A. Sounds super simple. Okay, everyone knows what A2A is. Google have a very huge marketing budget. But let's take, take a look at the complexities. MCP means stateless calls. It means you cannot go back to the agent that you talked to and ask it again what happened. A2A, one way, unless you implement
- 6:30
both directions. So agent A to, sends a task to agent B, it's client-server. If you want an agent B to send a task to a-agent A, it's client-server. So you need to implement on the same side both a client and a server. Chaining REST API calls, you will meet timeouts. And ob-obviously, thing, a lot of things are missing, like discovery, draft state of A2A, queues that you still need to implement and put behind all of that.
- 6:59
And if you're doing this, you are bui- basically building plumbing. You're not building a multi-agent system. You are wasting your time on boring, boring stuff. Now, let's summarize what we know. We know that multi-agent systems are already here. Everyone here is using multiple agents. We know that the messaging platforms are not a solution. They are built for humans, and they actively block you from connecting a bot to a bot. And we know that
- 7:28
protocol is too low level of a technical abstraction to be able, um, to build anything on it at scale.
- 7:39
But how hard can it be to build a system that can allow multiple agents to connect together? It's not that difficult, right? Well, you see, connecting remote agents, being a two sessions as a process is running on your laptop, or my Claude to your Codex, is a distributed systems problem. And distributed systems are not easy by definition. Forget agents, okay? Deterministic software, it's a pain in the butt. Add to this m-
- 8:09
agents that are microservices and non-deterministic, okay, and you'll have a lot of fun. In order for this, um, to be solved, what you need actually to address is a transport layer, right? The transport has to be real-time, it has to be ordered, because LLMs expect ordered messages. You need to handle retries, and so on. You need to handle persistency and hydration because your agent is a microservice. It can fail, the port can
- 8:39
crash, and then all your multi-agent system goes kaput. So you need to handle this. You need to handle something very new: runtime binding. If you want to connect a LangGraph to a Codex or Claude to a CrewAI, you need to connect a thread ID to a session ID to execution ID to a run ID of this multiple agents by... and map it to- to these IDs together. Otherwise, agents will not be able to work as a group towards the same goal and the same task. And obviously, you cannot work at
- 9:09
the level of IPs and ports or URLs that you have to manage, or even Pub/Sub topics. You need to rise above all of the technical details to a different level of abstraction: conversation, participants, channels or rooms, message routing that is tailored to agents as a first citizens. And of course, all of that is still not enough for an organization to use this kind of product or
- 9:39
communication software because every organization wants governance. They need identity, they need observability. Observability not in the message that was sent from agent A on server A to agent B on server B, but also what tool calls, okay, happened after a certain agent received this message, right? And it's very difficult to provide this kind of observability across distributed systems. So without all of that solved, you will not have agents talking to each other, and this wonderful future of us,
- 10:09
okay, everyone on a universal basic income, lying on a beach, will not happen.
- 10:17
I would like to introduce BAND. We actually took all the stuff that I talked about and solved it. So you don't have to solve it. So you can super easily connect all your agents within an enterprise or your Hermes and OpenClaude.
- 10:33
I would like to unveil a product that we are launching this week, and this is a global interaction and collaboration layer worldwide for any agents, any platform, any language, any environment. And it's not peer-to-peer, it's multi-peer. And we did not forget humans as well, okay? If you are human, you can still connect and talk to your agent. This is completely fine.
- 11:03
And if we open the hood qui-a bit, what we see there is that it's not just communication layer. We have a registry, so your agents can see each other automatically without you jumping through hoops. Persistency, channels, message filtering. Instead of doing it on the client side, we do it within the platform. Security, observability, and so on. Now, the fun part. Okay, let me show you a few demos. Uh, the time is short, so demos are recorded, but we have a booth. Please come and see it live, or just go
- 11:33
and connect. Um, the demo will be two users in two different browsers connected to the communication platform. One is me, Vlad. I have a personal agent, Claude, running on a Mac. And another user is Mike, Claude Code in a terminal session, okay? Uh, Claude SDK and, uh, LangGraph.
- 11:54
Demo number one. On the top window, it's me, my account, and I have a personal assistant. Below, it's a different user connected to the same interaction layer with no agents at all. On the right, okay, I'm going to spin up two agents, and they will onboard on the platform and connect and become agents that belong to Mike. And let's see how fast it's going to, to be.
- 12:26
So first of all, we go and we, uh, spin up an agent. Seconds later, this ag- agent card already appears with any- within an account of Mike. We spin up a LangGraph. Second later, this agent card appears as well. From this point on, this agent know of each other. They understand that they can interact with each other. But we want to do something cooler. We want to take a Codex agent and connect this agent to my personal assistant
- 12:56
running on my Mac so they can actually interact together. We are sending a contact request that requires a bilateral consent from both parties in order for this agent to be introduced inside each, uh, registry that belongs either to an agent or to a human. The request was sent, and it was accepted by me, and right now we have a Codex agent that can invite my personal assistant into a conversational space and
- 13:26
interact with it in real-time. These are completely stateful agents, can be deployed anywhere in the world, and they are. And what you see here, this is real-time.
- 13:39
And you can see also that because my agent is talking to another agent of a different user, I have a full visibility into the conversation, so no conversation can happen without me seeing it. And obviously, my agent can invite me into every conversational space and interact with me and as a human in the loop.
- 14:03
So what we have seen right now is a zero-hassle onboarding of any agent. Try it at home. Tell me how it, how, how, how it went. So another example. We have a user, Mike, and this user still has the Codex running, connected, uh, to the platform. And now we want to run a Claude Code terminal session, and we want this terminal session, i- from the terminal, to connect to
- 14:33
our platform and interact with other agents, right? It can be here specifically Claude Code in a terminal of Mike, talks to a, a, a Codex SDK of Mike, but because it's a global platform, it doesn't really matter, okay? Uh, they can talk to your agent, they can talk to you or your agent, and so on. So we are spinning up Claude, and we tell Claude in a terminal session, "Please go and connect to our platform." Same
- 15:03
concept. This specific Claude session gets an identity, an agent card, and, uh, we hide all the complexities in terms of where it is from the network perspective. And from now, this a- agent can actually create a conversational space as an owner, invite other agents, in this specific case, a Codex, right? And it will invite Mike as well, so it's a, you know, happy family, and they can start working on
- 15:33
some engineering tasks, whatever it is. The other way that people call it, whatever you see here, is loop engineering, okay? But it's loop engineering without five hundred thousand lines of a PyS- Python or TypeScript code that you need to bring over so your agents can actually prompt each other for work. These agents will now right now start to c- and create a small website, uh, and
- 16:03
they will interact with each other, and one will review work of the other. Uh, and obviously, if they struggle and Mike struggles, these agents can invite me into the same conversational space. I can help them, assist them, um, and it's pretty easy.
- 16:23
That's it. Thank you very much. I'm right on time. Come to our booth, LG17. Take screenshots. Take QR codes. One is for the platform to connect your agents, and one is for the future of the loop engineering that doesn't, it not require anything. And we have orcs that actually came here because they like orchestration. They do not like collaboration. They have stolen all the swag. So if you want swag, please come to orcs, get a sword, and help
- 16:53
us fight them. Thank you.