Get Out of the Model's Way — Kevin Hou, Google DeepMind
Read the talk
Get Out of the Model’s Way
Kevin Hou explains how Antigravity moves from human-managed agents to model-led teams, using dynamic subagents, event listeners and generated interfaces to turn better models into more capable products.
From a talk by Kevin Hou
At a glance
Ideas worth remembering
Scaling with intelligence means that improvements to the model should produce visible improvements in the user’s workflow.
Agent teams move task decomposition and worker selection into the lead agent, which dynamically configures specialists and can choose their models.
The evaluation example combines a computed delta, 100 parallel hypothesis investigations and a generated interface for inspecting the findings.
Sidecars provide persistent event listening; generative UI provides task-specific interaction. Together they extend an agent beyond responding to a prompt with text.
Give Messi the ball
Imagine coaching Argentina in the eighty-ninth minute with Messi on your team. Kevin Hou’s play is simple: “give Messi the ball and get the heck out of the way.” The analogy introduces a product design problem: a capable model can take on substantial work, but the surrounding application has to give it room to do so. Hou, who leads part of Antigravity’s engineering team at Google DeepMind, calls this “scaling with intelligence.” As the model improves, users should experience a more capable product.
Antigravity began as an agent-first coding product for technical and nontechnical users. Its initial IDE included an agent manager for orchestrating multiple agents. The team subsequently extracted the agent into a CLI and, with Antigravity 2.0, separated the IDE and agent manager into two applications. The standalone manager brings projects, subagents, worktrees, scheduled tasks and voice interaction into a single place. That separation is the first concrete expression of the principle: managing work should not require staying inside an editor.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
From autocomplete to orchestration
The progression starts with the tools Hou worked on in 2022: autocomplete and chat sidebars supported by embeddings, rules files and abstract syntax tree parsing. Much of the surrounding application followed deterministic logic because that was what the models could handle. As capabilities changed, each generation of tools needed different building blocks:
- Agents in 2024: MCP, custom tools and permission systems supported models that could take actions across multiple steps.
- Agent managers in 2025: Humans began supervising several agents in parallel, supported by skills, hooks and artifacts.
- Agent teams in 2026: Antigravity’s proposed next layer combines dynamic subagents, generative UI and sidecars, allowing the model to organize more of the work itself.
Moving between those layers can mean taking away something users like. Giving an AI access to a terminal raised obvious fears about destructive commands. Hou describes better model judgment and investment in permission systems as the combination that made terminal access useful. The lesson depends on both pieces: an agent gains the ability to execute commands, while permission controls help govern what it may run.
Removing the chat sidebar from Windsurf produced a different kind of resistance. Users lost a familiar interaction and received an agent in its place. The justification was a change in what the model could do: multi-step research and execution made it possible to carry a task forward rather than stop at an answer. Hou treats these decisions as bets on a potentially better workflow, acknowledging that the team does not always get them right.
Separating the agent manager from the IDE extends that same bet. Hou compares the IDE’s relationship to the manager with the debugger’s relationship to the IDE: useful when someone needs to inspect a lower layer, but unnecessary for every task. In this proposed hierarchy, orchestration becomes the everyday workspace, and the editor remains available for deeper intervention. Agent teams, swarms and software factories are different names for the future the team is betting on.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Let the lead agent assemble the team
The move to model-led teams comes from a relationship between model development and product development. Antigravity 1.0 put the human in charge of parallel agents; experience with that product then helped inform Gemini’s ability to manage agents. Hou describes Gemini Flash as becoming capable of leading teams as well as executing individual tasks, with improvements in speed and cost making that orchestration more practical. Task decomposition and collaboration still have substantial room to improve.
In the public-preview agent teams mode, the user enters /teamwork and describes a task. Specificity helps, and the lead agent can ask for missing information. The user then works with that lead agent, which assembles and manages the team. The division of labor is generated for the task: possible roles include frontend engineering, backend engineering, infrastructure, QA and design.
Each subagent can operate independently, and the lead agent can assign it a different model. This gives model judgment two jobs: deciding how to divide the task and deciding what kind of worker should handle each part. Hou’s Avengers analogy fits the intended behavior—a collection of specialists assembled around a mission, with roles chosen as the work demands.
The interface can adapt to the same task. Asking for progress might produce a Kanban board; asking for a view of activity over time might produce a debugger-like timeline. Both are generated on demand. The agent manager therefore does more than launch workers: it can produce a way to inspect their work that matches the user’s immediate question.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A kernel that runs Doom—and its bill
The team exercised this approach on a browser-based raw photo editor and a messaging application. Each used hundreds of subagents and took almost half a day. The more ambitious run built an operating system kernel from scratch that could run Doom, which a colleague demonstrated at Google I/O. Running the game supplies a concrete observable result: the generated kernel supported enough functionality to execute it.
Hou reports 93 subagents over 12 hours, 15,000 requests, two billion tokens and a cost under $1,000 for the kernel run. These are figures for one demonstrated project, rather than a general estimate for building an operating system. The result establishes an ambitious working example; it does not establish production readiness or show that increasing the number of agents always improves results. Even Hou’s description of the cost is measured: “mildly affordable.” A successful long-running team can still consume substantial time and inference.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Turn an evaluation delta into parallel investigations
The next example applies the same building blocks to internal research. A side-by-side evaluation begins with tasks, model rollouts and two tables of results: a control and an experiment. Comparing the tables reveals differences, but the researcher still needs to understand why they occurred and what to change in the next experiment. Traditionally, that investigation involved considerable work in Jupyter notebooks.
Follow those two tables through the new workflow. The researcher loads a skills file and asks about the evaluations in natural language. With those skills and knowledge of Google’s monorepo, the agent computes the delta. It then creates a research specialist to propose 100 hypotheses explaining the differences. Each hypothesis receives its own subagent, so the investigations proceed in parallel before their findings return as one report. A comparison has become a coordinated search for explanations.
How does one comparison expand into many investigations and become one usable result? The diagram follows the fan-out and return path. Parallel workers explore separate hypotheses; their findings converge before the researcher inspects the report. The hypotheses remain candidate explanations for review, rather than becoming established causes merely because agents investigated them.
The result also includes a generated interface with dropdowns, filters and ways to segment and slice the data. This changes what the researcher can do with the answer: inspect subsets and explore findings rather than read a single fixed response. The report and interface can be regenerated as needed. Hou reports that researchers automated 90% of this workflow and that a previously manual process now takes minutes; the talk does not define the denominator for that percentage or provide a measured baseline duration.
Previously, this kind of automation required someone to construct an asynchronous pool of agents, configure judges and connect data pipelines. Antigravity makes a dynamic subagent graph and generated presentation part of the product’s available building blocks. The skills file supplies workflow knowledge; the model organizes the investigation; the interface makes the resulting data easier to explore.
Two tables contain results from task rollouts.
The agent computes a delta, fans out across 100 hypotheses and brings the findings back into a report with an interactive view.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Dynamic workers and long-lived listeners
The examples lead into a more precise description of two building blocks:
- Dynamic subagents: The main agent configures, prompts and seeds workers as needed. They can run in parallel, take specialized roles and operate in environments such as sandboxes or remote execution systems. The intended scaling path is better task decomposition and collaboration as the coordinating model improves.
- Sidecars: A long-lived utility process listens for outside events. It gives the model a way to arrange triggers from SMS messages, webhooks, cron jobs or GitHub pull requests. Antigravity already uses this mechanism for scheduled tasks.
The distinction is useful: a subagent performs delegated work, while a sidecar keeps listening for something that should initiate work. Scheduled tasks are one instance of the broader listening mechanism. Hou presents sidecars as a new plugin protocol and announces a forthcoming specification for others to build on; the talk describes its purpose and existing internal use, rather than its implementation API.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
An interface that is not fixed in plastic
The final building block is generative UI. Hou reports that Gemini Flash in Antigravity produces almost 900 tokens per second, which he describes as 10 times faster than many other frontier model experiences. The comparison’s workloads and measurement conditions are unspecified, but the product implication is clear: fast generation makes it practical to create a task-specific interface within seconds.
Antigravity renders generated UI inline in the conversation. The examples range from playing Doom to inspecting bar charts, graphs and tables. Earlier, the same mechanism produced a Kanban board for task status and an interactive evaluation report. These views serve different questions, so their controls and presentation can change with the work.
Hou’s claim that human-written specialized interfaces may be obsolete is a hypothesis about where this capability leads. His analogy is Steve Jobs’s criticism of physical keyboards and control buttons that were “fixed in plastic” and identical across applications. A generated interface can instead supply the controls needed for the current task. The evaluation example makes that distinction concrete: a researcher gets filters and data slices because the work calls for inspecting results, while a project manager can request a board or timeline.
The closing design question is what to add when implementing features has become easier. Dynamic subagents let the model organize workers; sidecars let it respond to events; generative UI lets it present and expose the results in a useful form. Hou recommends choosing building blocks that can benefit from the next model’s capabilities. Getting out of the model’s way therefore requires product decisions: provide mechanisms through which better judgment, faster generation and more capable coordination can change what users accomplish.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Read the complete timestamped transcript
- 0:33
All right. Hello, everyone. Um, my name is Kevin. I'm gonna be talking about Antigravity. So are there any World Cup fans out there?
- 0:42
Woo!
- 0:43
Woo! Imagine you are coaching Argentina, and you're in the eighty-ninth minute, and you have Messi on your team. What play are you running? It's called give Messi the ball and get the heck out of the way. LLMs aren't just role players anymore. They can be your star player if you build the right product around them. And to let your star player cook, you have to get out of the model's way.
- 1:12
We might wanna get the slide. Are the slides up?
- 1:15
Yeah.
- 1:15
Oh, they are. Great. Um, so Antigravity is Google's agentic coding product for technical and non-technical users. Uh, we launched back in November of twenty twenty-five and have been accelerating devs both within Google and externally ever since. My name is Kevin Hou, and I lead part of the engineering team on Antigravity. So let's talk a little bit more about what Antigravity is. We have and always will be unapologetically agent-first. So we debuted the Antigravity
- 1:45
IDE last year with a brand-new agent manager concept, and it was a platform to manage and orchestrate many agents. Since then, we've actually extracted our agent and launched our own Antigravity CLI. And last month at Google I/O, we had the pleasure of launching Antigravity two point zero. In the theme of getting the model out of the way, we actually decoupled the IDE from the agent manager, so now you have two separate applications, um, and now you can use the agent manager in a
- 2:15
standalone app.
- 2:18
And since pictures are worth a thousand words, here's a screenshot of Antigravity two point zero in action. As you can see, not only is it your own dedicated mission control for your agents and projects, you have sub-agents. You have all the new models. You have work trees, scheduled tasks, voice mode. There are so many things to unpack with the product. But I don't wanna spend today telling you about the product. I wanna tell you a little bit more about the behind-the-scenes, some of the principles that went into it, and notably, some of the things that led to its roadmap.
- 2:48
So as some of you, mm, for the long time AI Eng fans, uh, this is actually my fifth time speaking at AI Eng, um, and I've been building developers tools since twenty twenty-two. And the one thing that has stood above all other lessons that I've talked, talked about is the idea of scaling with intelligence. This means that as the model gets better, so should your product. And the frontier edge of whatever model you are serving should be apparent inside of your user's product experience. So let's get into more concrete examples of what
- 3:18
this means. So for those of you that follow me on X or hear me just yap generally for the last four years, you'll know that I've been working on a number of these sort of transformations year over year over year. In twenty twenty-two, I was working on autocomplete and chat sidebars. This was based on embeddings, rules files, AST syntax tree parsing. Basically, everything inside of that app is deterministic because that's all that the model could really handle. And in twenty twenty-four, when
- 3:48
agents came onto the scene, it completely changed how developers were going to do work. With it came new primitives like MCPs, custom tools, and permission systems. And with twenty twenty-five, we introduced Antigravity's agent manager with many other products following suit in that similar form factor, with users managing many agents at once in parallel. And this led to things like skills, hooks, artifacts, and a couple other primitives, um, and that sort of defined the twenty twenty-five era.
- 4:18
So let's talk a little bit about twenty twenty-six and what those primitives might be. Before we answer this question, I want to take you back to some of these battle scars that are a little bit closer to home. Scaling with intelligence really is not easy. It's really hard to take away something that users love and are familiar with to lead them down potentially, and that's a big keyword, a better path. We aren't right a hundred percent of the time, but there are two that jumped to mind when I was putting together the slides for this talk. The first one is
- 4:48
giving AI a terminal. We all remember fears about Son of Anton deleting your entire code base and doing catastrophic things to both, you know, your startup, your company, et cetera, et cetera. But as models got better and people invested in primitives, such as permission systems, users ended up building faster, they ended up shipping more, and they did so safely. So we were able to overcome this, and as models got smarter, they were able to make better decisions about what they should and should not run in your terminal.
- 5:17
The second instance is, um, this tweet, which is very representative of sort of the yelling that I got, uh, when we removed chat from Windsurf. So a lot of users were yelling at our team because we took away something that was very dear to them, the chat sidebar, and replaced it with only an agent. Now, at the time, this is something that was familiar and rather difficult to swallow. But when we look back, models have advanced, multi-step research, agentic research and execution became the new paradigm, and here we are today using and loving
- 5:47
all these agentic products. And so now I bring you to today's battle. What is going on today? So we decoupled the agent manager from the IDE, and with Antigravity two point zero, we split them into separate applications. We believe that the IDE is to the agent manager what the debugger was to the IDE. You don't always need a debugger, but it definitely is helpful to have it if you need to go a layer beneath and go one step deeper into that abstraction stack. And our prediction is that this idea of agent
- 6:16
orchestration, you can call it agent teams, you can call it swarms, you could call it software factories, is the future, and we're willing to bet on that future. So here are the primitives for what we're calling the agent teams twenty twenty-six era. These are things like sub-agents, generative UI, and sidecars. And we'll talk more concretely about what those things are and some examples of how they manifest inside of the product. But it's really important to first understand the why. What brought about
- 6:46
these changes and what model changes, what model properties actually led to the development of these new things? And as a product team, do you force the new era of primitives or is it something that comes to you by using the model and experiencing the model? The answer is kind of both, right? And the privilege of being inside of Google DeepMind is that we do have that relationship between the product and the model. So you remember the crux of Antigravity 1.0 is to manage agents in parallel, to put the human in the
- 7:16
driver's seat. And if you remember my last talk, I talked a lot more about this research product flywheel. And now, as promised, because of the Antigravity product, Gemini has now learned a thing or two about how to manage a team of agents. There's still a lot of headroom to make multi-agent systems better, more collaborative, better at deconstructing tasks into smaller tasks, but we've got a really good head start with Gemini. And all the basics have been imbued to the model so that we can
- 7:46
build a product like Antigravity 2.0.
- 7:50
Gemini 3.5 Fla-- Yeah. Gemini 3.5 Flash was launched back in April, and this brought to market a lot of those capabilities that we had been working on in the background with Antigravity. And Flash now isn't just good at executing tasks, it's actually really good at leading teams. It's faster and cheaper, pushing the Pareto curve of what is intelligent versus the speed and the cost at which you run those things. And putting this all together, we were really excited to announce agent teams in public preview inside of Antigravity.
- 8:21
All you have to do is simply type the slash command, slash teamwork, and you'll see a new mode where you can enter and unleash a swarm of agents onto the task at hand. So we'll talk a little bit about how this works. You as a user will specify your task. The more specific you are, the better, though the nature of these agentic communication styles is that if it needs something more, it can actually ask you for more until everything is basically clear. You'll work with that lead agent, and it will manage a team of arbitrary size to get that work
- 8:50
done. And what I like to say, it's kinda like the Avengers, right? It'll take a bunch of specialized roles. It may-- Front-end engineers, back-end engineers, infrastructure specialists, QA design. The list goes on and on and on, and there are infinite possibilities for what each of those sub-agents could take on. Each sub-agent is dynamically generated and can operate independently. Um, and it can even actually select a different model from what the main agent is using, and this is done so by that main agent. Again, we are scaling with intelligence.
- 9:21
And one of the coolest aspects of this is that it can use generative UI. With a model that is as fast as Flash, things can happen nearly instantaneously if you ask, "Hey, what is the status of my task? Show me a Kanban of what's going on." Or maybe, you know, you prefer something a little bit more like, uh, the, the Chrome debugger tool. It can show you a timeline like that. And all these things are generated on the fly because it's able to generate UI on demand. So some of the projects that the system has implemented, um, we've built a photo editor.
- 9:52
You can actually edit raw photos directly inside of your browser. Um, we've also built a messaging app that might look a little bit familiar to those in the room. Um, and each of these took hundreds of sub-agents, uh, and took almost half a day to run. But to really put it through its paces, one of the hero runs that we did was actually building an entire OS kernel. This is something that, uh, we got to show off at Google I/O, but we built a complete OS kernel from scratch and actually played Doom on it. And my colleague Varun was able to demo this at
- 10:22
Google I/O. We were super proud of this particular milestone because it really demonstrated that if you throw more intelligence, you throw more sub-agents, um, at this sort of problem, a model like Gemini 3.5 Flash could do this in a way that was not only very, very powerful, but also scalable and, you know, mildly affordable. Obviously, we're not gonna spend thousands and thousands of dollars to build an OS kernel every day, though it is possible. And some of the stats out of this, it took ninety-three sub-agents over the course of twelve hours, made fifteen thousand
- 10:52
requests, two billion tokens, and it was under a thousand dollars, which was one of the really cool aspects of this project. And so as you can see with this particular example, sub-agent primitives are one of the defining parts about building a twenty twenty-six era of agent teams. So agent teams are just that first example, and I wanna show you another example that our team uses internally that sort of demonstrates some of these new primitives. Um, the second one is about automating research tasks. So we work inside of Gemini. We help
- 11:22
sort of make Gemini better at coding-related tasks, agentic-related tasks, and this is where the real magic starts happening with the product. We have an internal version of Antigravity that researchers, engineers, non-technical folks can use, and when they understand the primitives that Antigravity offers, it becomes a very, very powerful way to automate your own workflows. So we'll take the example of side-by-side eval analysis. So this is a very common workflow, not only at, at DeepMind, but just generally in the industry. You essentially will take multiple rollouts,
- 11:52
one, two, three, four, et cetera, um, and you wanna compare them. So you'll take a set of tasks, you'll do some rollouts, you'll get some results, and they'll essentially be in two different tables. Now, you'll look at the control, you'll look at the experiment, and then you'll have to figure out not only what the difference was, but perhaps what are the reasons for those differences and how can we actually iterate from there and make a better version of-- for the next experiment. Now, traditionally, this was a lot of Jupyter Notebook elbow grease, essentially. But when you start working with the new primitives in twenty twenty-six,
- 12:22
you end up with a lot cleaner of a workflow. So researchers were able to automate 90% of this workflow by simply asking the agent about the evals in question using natural language. Then the agent that is now primed with skills and an understanding of Google's massive monorepo code base is able to crunch the numbers and get back to you with a delta. Now, what's really cool here is instead of just taking that delta then handing it back to the user, it went the extra step. It spun up for a research agent specialist that proposes 100 different hypotheses over
- 12:52
why those deltas might occur. And then it uses subagents to then spin up one subagent for each hypothesis, and basically drills into that particular case in parallel, mapping back to a single response, and then telling the researcher, "Hey, here are some areas that I found. Now, uh, here's a report that you can review." And what's really cool is that it doesn't stop at just the report. It actually puts together a generative UI for you to look through, interact, select drop-downs, filter, segment, slice, and actually interact
- 13:22
richly with that data.
- 13:26
And internally, we care a lot about this sort of workflow, improving the model, improving the product, and understanding the ways that users find success and failure internally at Google. So what used to be a very manual process now takes minutes. So what used to be hand engineering, you'd have to build your own async pool of agents, you'd have to set up your judges, uh, you'd have to tape together data pipelines. All of this now starts becoming grounded in these new primitives that we've established earlier in the slideshow. You have a subagent graph that is completely dynamic.
- 13:56
The generative UI comes in at the end to richly convey the findings in a way that the user best understands or maybe caters to their learning style. And all of these things can be regenerated and redone on the fly. All the user had to do was load up a skills file and ask away. So with teamwork and this eval example, we start arriving at these 2026 primitives that I keep talking about, and these model characteristics really change the way that we have to think about the product and how we have to develop the product. So the three examples that we've talked about,
- 14:26
first, we have the dynamic subagent. And to provide a little bit more color here, basically, no two subagents are the same. The main agent is the one that is orchestrating this entirely on its own. It's configuring and prompting and seeding these subagents on the fly. They can operate in parallel. They can operate in different types of secure environments, be it a sandbox, be it a remote execution system, and they can all take on infinitely nu- an infinite number of specialized roles. So the scaling story here is quite obvious, and from the last two
- 14:56
examples, you can probably tell. As the model gets smarter, your team will become more specialized, it will become more collaborative, and ultimately that means it'll be capable of getting more complex work done for you. And now the second is this new concept. We've alluded to it slightly in the past, but it's called sidecars. This is a new plug-in protocol that we're bringing to Antigravity. A sidecar process is essentially it is a sidecar process. The naming sort of reflects what, what's going on under the hood, but it's a long-lived utility, and it's responsible for
- 15:26
listening. It allows the model to listen to the outside world and set up its own triggers for things that might happen. For example, this could be SMS messages. This could be webhooks, cron jobs, hooking it up to GitHub PRs. The, the list goes on and on, but this is a generic plug-in primitive. Antigravity already uses sidecars for things that are time-based. This is where the c- the scheduled task cron concept comes from. Um, but under the hood, this is all this new sidecar primitive. So we'll be releasing the spec for this so that you all can build on top of this new primitive, um, later this summer.
- 15:57
But there are some really, really cool ways that people internally have been using this sort of concept. And the third and final primitive is generative UI. So we hypothesize that human-written specialized UIs are kind of dead. Gemini Flash on Antigravity clocks in at almost 900 tokens a second. This is 10X faster than a lot of the other frontier model experiences. And in a matter of seconds, you're able to go from whatever you were thinking inside of your head into a prompt, into a use case that is designed and
- 16:26
embedded inside of your conversation view perfectly. And rather than rely on templates or even HTML files, Antigravity can render your generated UI in line. So you can do things like this and play Doom, but this also extends to things like bar charts, graphs, um, tables, anything that you would want in to interact with and maybe inspect a little bit further than just a markdown file or just a conversation.
- 16:53
And generative UI in many ways reminds me of the quote that the late Steve Jobs said when unveiling the iPhone. He justifies the removal of the keyboard and says, "They all have these keyboards. They are there whether you need them or not. And they all have these control buttons that are fixed in plastic and are the same for every application." In an analogous way, we built our product to dynamically scale with the needs of the agent. We skipped the heavy infrastructure and mechanical UIs in favor of sidecars and generative UI, and that creates a
- 17:23
product experience that is not fixed in plastic. So subagents, sidecar triggers, and generative UI are the latest primitives that are powering Antigravity. We've tried our best to stay out of the way and let the model cook, and if you're building a product around an agent, you should consider what are the primitives that are in my product, and how might they scale with the model's intelligence? We all are familiar with shipping features is now quite easy with all of these new tools, and it's about deciding what features to actually
- 17:53
add, um, so that the model-- so that your product can scale with the next release of the next model, which will inevitably be faster, better, and cheaper. And so with the right primitives, you as a builder or you as a product owner, you might be surprised at what the models can do. And in classic fashion, I'm gonna keep using this slide until we've actually conquered the TPU crunch. So you can find me on Twitter. Uh, you can DM me for feedback. We're always looking for new ideas on how to build the latest and greatest. Thank you for watching. Thank you, Swix and Ben, for having me.
- 18:23
It's always a joy to be here, and I'll be at the Antigravity booth if you wanna talk further. If you wanna get to know the product a bit more, some team members will be there. So thank you so much for your time. Excited to meet you all.