AI Engineer World's Fair 2026
Research to Reality with Google DeepMind — Paige Bailey, Google DeepMind
Read the talk
Research to Reality with Google DeepMind
Paige Bailey moves from DeepMind’s science mission to a Gemma model running in a browser, then follows local inference onto phones. The practical question is how model size, runtime and quantized checkpoints make those deployments possible.
From a talk by Paige Bailey
At a glance
Ideas worth remembering
The browser demonstration runs Gemma through Transformers.js locally, producing a comparison table without sending the prompt to a model API.
AI Edge Gallery exposes image descriptions, multilingual audio transcription, function calling and skills on Android and iOS, with device acceleration available on supported hardware.
Deployment depends on model size and checkpoint representation together: Bailey reports a quantized 2B checkpoint smaller than one gigabyte and connects QAT checkpoints to the fast browser demo.
From scientific research to agents that can execute tasks
New models become useful when people can put them into their own projects. Paige Bailey, presenting for Google DeepMind, opens with that practical aim: live demonstrations, questions throughout, and room to explore how the models work. The research context comes first. DeepMind’s mission connects responsible AI development with work on protein structure, medical models, robotics, mathematics and frontier research.
The first product sketch concerns execution. Managed Agents takes a higher-level task described in natural language and gives a fleet of agents a Linux workstation in a sandbox environment. Skills and dependencies can be added to that environment as work proceeds. The workstation matters because it gives the agents somewhere to carry out the task, with software available to support it. Bailey also introduces a computer use API and a speech-to-speech translation API, although this recording moves into local models before developing those demonstrations.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Gemma gives developers a model they can download and adapt
Gemma 4 supplies the open-model side of this story. Bailey describes a family with 2 billion, 4 billion, 12 billion, 26 billion and 31 billion parameters. The two largest versions use different architectures: the 26B version is a mixture-of-experts model, while the 31B version is dense. These are distinct options within the family, rather than one model that must serve every device.
The Apache 2 license is the practical opening for customization. In Bailey’s account, developers can download the models, use them within a company, fine-tune them and build on them for their own projects. That makes deployment location and adaptation part of the developer’s decision: the model can become a component in a locally running application.
“I hate slides,” Bailey says before switching to a community-built browser application. That move makes the deployment question concrete. A downloadable model is useful here because the application can load it directly into the browser and run it on the user’s machine.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A Harry Potter comparison runs entirely in the browser
The browser demo turns a fairly elaborate request into a table. Bailey asks for an emoji comparison of the Harry Potter books, ranked by how funny and exciting they are, with recommendations on what to read. The prompt also asks to incorporate a rationality-themed Harry Potter work. The important change is visible: a prose request becomes a structured comparison containing entries such as The Deathly Hallows, Half-Blood Prince, Order of the Phoenix and Goblet of Fire.
The response feels almost instantaneous to Bailey. She immediately questions its taste—“I’m not sure if I agree with the ranking”—which separates the two things this example demonstrates. The model can follow the requested format and produce recommendations quickly; the literary ranking remains a judgment the reader can disagree with. Perceived responsiveness is an observation from this demonstration, rather than a measured latency result.
The execution path explains why no inference API call is needed. The application runs Gemma locally through Transformers.js, with WebAssembly providing the browser execution mechanism Bailey describes. The loaded model handles the prompt on the machine and returns the table to the browser interface. For this demo, Bailey says the user’s data is not sent elsewhere.
Where does the request travel? The diagram keeps the prompt, runtime, loaded model and resulting table inside the same device. Its useful relationship is the return path: the browser receives an answer from local inference, without handing the prompt to a remote model API.
Ask for emoji, humor and excitement rankings, and reading recommendations.
The Harry Potter request and generated comparison stay on the machine in the demonstrated application.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
AI Edge Gallery brings local models to phones
Google AI Edge Gallery extends the same approach to Android and iOS. Users can download the app and use locally installed models to explore several kinds of input and behavior. The phone becomes both the interface and the machine doing inference.
The capabilities serve different purposes:
- Image descriptions: Give the model an image and ask it to describe the contents.
- Audio transcription: Turn audio into text in multiple languages. Bailey tentatively cites support for over 140 languages; that figure is a qualified claim about language support, not a transcription-quality result for each language.
- Function calling: Try model-driven function calls directly on the device.
- Skills: Explore examples including games, haikus, weather queries and scheduling calendar events.
These examples span content generation and actions. A haiku produces text; scheduling a calendar event involves using a function or skill to do something beyond returning prose. Local inference identifies where the model runs. The recording does not establish that every service reached by a weather or calendar skill also operates offline, so the browser demo’s no-data-sent claim should remain specific to that demonstrated application.
Hardware acceleration is the next deployment choice. On supported phones, AI Edge Gallery can use a device accelerator to perform inference instead of falling back to the CPU. Bailey points to higher-end mobile devices as examples. The mechanism moves the model’s inference work onto available acceleration hardware while keeping execution on the device; support therefore depends on the phone’s hardware.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Choose the model and checkpoint that fit the device
The ending returns to model size as an operating decision. Bailey describes the 26B and 31B models as performing above expectations for their size, including against models an order of magnitude larger. That comparison comes without a named task or benchmark here, so it supports her deployment motivation rather than a general performance ranking.
The proposed cost advantage comes from fitting useful inference into less hardware. A model that runs on one commodity GPU needs a smaller GPU footprint and avoids the distributed inference setup that would otherwise spread its execution across machines. Bailey presents a practical range of deployment targets:
- Single-GPU deployment: The larger variants are intended to run well on a single commodity GPU.
- Laptop deployment: Versions at 12B parameters and below can run locally on a laptop.
- Mobile deployment: The 2B version can fit on a phone.
These are deployment categories, not guarantees for every device in them; the recording does not give a minimum memory configuration.
Parameter count is only part of the fit. The released quantized checkpoints provide smaller model representations, and Bailey identifies QAT checkpoints as what the fast browser demonstration used. For the 2 billion parameter version, she reports a checkpoint size of less than one gigabyte. That quantity describes the checkpoint, rather than the total memory consumed by the browser and inference runtime.
This closes the loop on the Harry Potter table. The browser loaded a model, executed the request locally and returned a structured answer; the closing explanation identifies a compact quantized checkpoint as part of what made that deployment practical. The useful choice is therefore more specific than selecting a model family: choose a model size, a checkpoint representation and an execution path that fit the machine where the work should happen.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
Related talks
- Beating RL With Reflection: GEPA and Optimize Anything
A complementary look at making a model useful for a specific task: reflective optimization changes prompts, agent programs and repository skills, with examples of improving smaller models without increasing their size.
Read the complete timestamped transcript
- 0:12
Trickle in from other locations, and then we'll get started pretty shortly. Um, ideally, there will be some time for lots and lots of live demos. Um, so you'll ideally learn some of the things about our new models, how to use them as part of your projects. Um, and then there also should be time for questions. But it also seems like we'll have a small enough audience where if you ask questions throughout the duration of the presentation, we can also do that. Um, so end goal is to make this as useful as possible for all of y'all. Um, my name is
- 0:42
Paige. I work at DeepMind, and, uh, excited to show y'all what we've been working on too. Cool.
- 0:54
And as mentioned, we'll get started in just another couple of minutes. It looks like people are still getting scanned at the door.
- 2:38
All right. Let's get started. Cool. So, greetings everyone. As mentioned, my name is Paige. If you have questions throughout the duration of the presentation, please feel free to raise your hand and shout them out. Um, ideally, we'll be able to take some, uh, take some as I kind of careen along doing live demos. Um, I also want to give a kind of the requisite caveat at the very beginning that, um, DeepMind's entire mission is to build AI responsibly and for the benefit of humanity,
- 3:08
so trying to cast as much light as we can, um, in this, uh, in this world that increasingly has, uh, increasingly has shade. Um, this translates to things like AlphaFolds, um, much of the work that we do with our open models, including Med-Gemini, um, and a lot of our robotics and science use cases, as well as things like AI for math, AI for frontier research. Um, so in general, everything that I show you today, um, is being used for those
- 3:38
projects that are deeply, deeply embedded in AI for science. Um, and if any of y'all work in the AI for science space, um, please feel free to send questions and ask afterwards about how you can use AI to, to kind of transform and accelerate that, uh, that good work. Um, I don't think it's a secret that Google has been a little bit busy over the course of the last few months. Um, it feels like we've been releasing new models, new features, new products every single week. Um, most notably, uh, just
- 4:08
recently we released a computer use API, which we'll be talking about a little bit, um, something called Managed Agents, which gives you the ability to take, um, kind of a higher order task, uh, describe it in natural language, and have a fleet of agents execute on it in a Linux workstation, like a sandbox environment where you can add skills, you can add, um, kind of dependencies that get pulled in along the way, and a whole bunch more. Um, and then also our speech-to-speech translation API, which we'll take a look at in a second.
- 4:38
And one of the things that I really, really love about Google is that not only are we shipping frontier models, so you might have heard about Gemini 3.5 Flash, which got released at IO, um, but we've also been focused pretty significantly on our open model releases. Um, so how many folks in the room have heard of Gemma 4? Um, quite a few hands. That's excellent. Uh, Gemma 4- Oh, yes, absolutely. Go for it. Um, Gemma 4 is our
- 5:08
open model family. It's the latest iteration. Um, we have many different sizes available, so there's a two billion parameter, a four billion. We just released a twelve billion parameter, um, and then we also have a twenty-six billion and a thirty-one billion, um, mixture of experts and dense model, um, respectively. Um, they're useful for a lot of things, and they're also Apache 2 licensed, which means you can download them, use them as part of your company, fine-tune them, uh, and kind of, uh, expand on
- 5:38
them however you feel like would be most useful. Um, and this Gemma universe is actually pretty cool to see, um, in terms of people building. I, I hate slides, so we're going to see how few slides I can get through today. Um, but, uh, this is an example of something that someone from the community has built using Gemma 4, um, and Fable 5 before it was taken off the market to rewrite some of the kernels. Um, you can load the model directly within the browser. So this is
- 6:08
loading Gemma directly in the browser, using it via WebAssembly. It's sandboxed. Um, and then it can do things that, that feel pretty magical, right? So this model is, uh, this model is currently running completely locally. Um, and if I say something to the effect of, "Create a table with emoji comparing and contrasting all of the, uh, Harry Potter books, um, uh,
- 6:38
based on which are the funniest and most exciting. Um, make sure to give me recommendations on which to read and also, um, um, um, maybe incorporate Harry Potter, Harry Potter in the methods of rationality." Um, or...
- 7:05
And then you get a kind of a, a response that feels almost instantaneous. Um, you have, uh, kind of The Deathly Hallows, Half-Blood Prince, Order of the Phoenix, Goblet of Fire, et cetera. Um, I'm not sure if I agree with the ranking, but that's, uh, but that's a pretty interesting in, pretty interesting comparison. Um, and then also, uh, since it is local, um, to the machine, none of your data is getting sent elsewhere. This is not using an API. It's
- 7:35
just something that's running locally in the browser with Transformers.js, um, and the Gemma 4 model. Gemma 4 is also pretty cool in the sense that you can use, um, Google AI Edge Gallery in order to analyze, um, some of the model capabilities that we have on device. You can download it and use it with Android and with iOS. Um, and it includes everything from kind of taking images and describing them to automatically transcribing audio in multiple languages, I believe Gemma
- 8:05
supports over 140 different languages, um, and then also testing it out with function calling directly on device. Uh, if you haven't had a chance to take a look at the Google AI Edge Gallery, um, we have a collection of skills as well. Um, so things like building games, doing haikus, um, asking queries about weather, being able to schedule events on your calendar for you, um, that are all available to use just with this AI Edge Gallery app completely for free and just with locally
- 8:35
installed models. Um, so if you haven't downloaded it, definitely try it out. There's also a way to use, um, the, the kind of accelerator on your local device in order to power the model and do all of the inference work as opposed to just the fallback to the CPU. Um, so if you do have something like a Pixel 10 or a higher end, um, a higher end mobile device, you're already able to use Gemma kind of locally, um, for, for all of that work, which is quite
- 9:05
cool.
- 9:07
Gemma 4, just to, to give a recap or to place how the model performance versus size shapes up, um, this is the two, uh, the two kind of largest versions of Gemma that I had mentioned before, the 31B and the 26B. Um, they're quite small in comparison to some other models that are on the market, but they're performing way above what you might expect. Um, so more than models that are in order of magnitude or larger than their
- 9:37
size. This is great because it kind of translates into a cost savings perspective. You don't have to worry about distributed inference for local models. Um, you don't have to worry about as large of a GPU footprint. Um, and it's really intended to run super, super well on a single commodity GPU as opposed to needing something a little bit more fully featured. You can also run Gemma, some of the variants, on even things like Jets and Nanos. Um, and for 12B and below, you can run it locally on your laptop. Um, 2B
- 10:07
can even fit handily on mobile devices. We've also released quantized versions. So the, the model that we saw at the very beginning that was running so, so blazingly fast, um, was using some of the QAT checkpoints that we have for Gemma. Um, and the QAT checkpoints, um, are quite small when you, uh, when you take a look, even less than a single gigabyte in size for the, the, uh, kind of 2 billion parameter version.