AI Engineer Europe 2026
Accelerating AI on Edge — Chintan Parikh and Weiyi Wang, Google DeepMind
About this talk
Google LiteRT product manager Chintan Parikh introduces Google AI Edge and Gemma 4 edge models, describing the shift toward reasoning-capable on-device agents. The presentation demonstrates local agent and robotics use cases, explains LiteRT's TensorFlow Lite foundation and cross-platform Android, iOS, and IoT deployment, and covers AI Edge Portal benchmarking plus Qualcomm and MediaTek hardware integrations. Colleague Weiyi Wang is introduced for the question-and-answer portion, which includes questions about cameras, Raspberry Pi, and Orin devices.
Chapters
- 0:00Speaker introductions and the edge AI agenda
- 1:22Gemma 4 edge models and on-device agents
- 5:43Gallery applications and local agent demonstrations
- 11:22LiteRT foundations and cross-platform deployment
- 14:35Benchmarking, hardware integrations, and robotics
- 19:30Audience questions, target devices, and open-weight models
Talk transcript
- 0:00
[upbeat music] Good afternoon, everyone.
- 0:16
We'll get started. Uh, my name is, uh, Chintan Parikh, uh, Product Manager for, uh, LiteRT, which is part of, uh, Google AI Edge. And, uh, I also have my, uh, my colleague here, Weiyi also, who will be joining us also for the, uh, Q&A part of this session.
- 0:33
Uh, quick show of hands. How many of you are working on deploying on edge or are keen on learning more about, like, what it's gonna be the big benefits?
- 0:42
Okay. And some of you are already deploying. So, um, I'm gonna go through the slides. I'll try to, you know, also leave it open to understand if you guys have any use cases or, uh, any sort of things you're working on that, you know, you'd like to discuss, uh, so we can kind of keep it open in
- 0:55
that sense. Great. So, uh, here's gonna be a quick, uh, set of agenda items, and I'm gonna go through some of the new models that are coming up. Um, also, uh, some edge use cases that would be relevant.
- 1:06
Um, I'm gonna show you some of the new capabilities in the gallery app, and then we'll go through our stack for deploying, uh, AI on edge devices. In addition, uh, we also wanna emphasize, uh, the cross-platform support we offer 'cause I think a lot of you are looking to deploy on just more than just on mobile platforms,
- 1:22
but also other platforms. So we'll kind of do some of that. Um, Google DeepMind recently launched, uh, Gemma 4, and, uh, this talk is gonna be focusing a little bit more on, like, the 2B and the 4B edge models that are, are focused on, uh, deploying on-device.
- 1:38
Uh, Google certainly has a big suite of models and even the Gemma 3 family, which are also in the smaller size, like all the way down to two hundred and seventy million parameters.
- 1:46
So if you're looking for extremely small models that you're able to fine-tune, uh, we certainly have a Hugging Face page, which I'll go over, which has more of these models.
- 1:55
Um, so yeah. So the big evolution with Gemma 4 is going to be really moving from, like, chatbot-type capabilities to more autonomous agents that also support reasoning capabilities and more sophisticated features.
- 2:06
And, uh, a lot of it is also, uh, it's gonna be great how you can deploy these on different devices. So running on edge has many benefits. I think, uh, my colleague went over some of these, uh, yesterday, and I'll just, you know, run through this, uh, really quick also.
- 2:20
So certainly, like, latency is important for those who are keen. Um, any of these, like, real-time camera use cases, like filters, where you're looking to replace backgrounds or, like, video calls, like real-time latency is king over there, so that's, you know, on-device can help with that.
- 2:33
Privacy is also critical for any use cases where, you know, you have sensitive details, you're doing some kind of summarizing, uh, documentation which is sensitive. Offline use cases where you have poor connectivity and costs.
- 2:45
Like, you know, I've seen so many presentations, uh, this, uh, uh, AI engineer event where, uh, people are complaining about the number of tokens that are getting used and this and that.
- 2:54
But I think, uh, on-device always offers this hybrid approach, where if you're looking for, like, how do I really offset, you know, uh, running things on the edge versus on cloud and where, uh, can I get the best balance?
- 3:06
I think there's, uh, there is something here.
- 3:10
Cool. So, uh, in terms of models, there's two models, and I think, like, a lot of the use cases I'll show are gonna be built on these. So the Gemma 4e-2B roughly from a, from a amount of, uh, RAM usage is roughly anywhere from one to two GB of RAM usage.
- 3:23
So it could be usable but, like, yeah, it depends on your end use cases again. Uh, certainly good for any, uh, voice interfaces or, uh, use cases for, uh, summarization or any kind of low latency, lo-local processing.
- 3:36
Uh, 4, 4B is gonna be a little more heavy-duty if you're looking to run on bigger platforms like laptops or IoT devices. It'll have, like, a higher RAM, RAM requirement.
- 3:44
Again, this is once it's been quantized, uh, to your desired size.
- 3:51
So, uh, I wanna do a quick deep dive on what's gonna be new, uh, this time on, uh, on Gemma 4e-2B and 4B capabilities in terms of the agentic capabilities.
- 4:02
So I'm gonna show you what's already there in the next couple of slides in terms of use cases, and here, uh, touch a little bit on exactly, like, what are the new capabilities are there.
- 4:10
So function calling, built-in support for tool calling, and also essentially models to interact with other local APIs. So you can certainly start the int-inferencing on edge, but you have the path to essentially, you know, call other APIs outside, and I'll, I'll show that in some of my examples.
- 4:26
But the core of the inter-in-inferencing will be on the edge. Uh, structured JSON output. So for native support for any structured JSON output is supported here, which was built into the model architecture rather than achieving through some kind of specific prompt engineering.
- 4:39
You can kind of do this too, and I'll show examples for the same. Uh, chain of thought is new. So there is a thinking mode, uh, which we will-- which can be demonstrated by an app where the thinking mode will help you understand, like, the thought process that the model's going through.
- 4:52
Our gallery app, which I will showcase, will, uh, will support that. And finally, what's gonna be great is these models are optimized for hardware native support, means you can run seamlessly across multiple platforms and multiple hardwares.
- 5:05
And really we want to give you the flexibility of deploying into, uh, various platforms.
- 5:12
So just for starters, if anyone's, like, thinking, "Okay, so where can I get these models?" Like, we have these models already ready to go. So our Hugging Face page has these models.
- 5:20
So if you're looking, you can, uh, you know, they're all Apache 2.0 licensed models, so you are able to download these and then, you know, start building with it.
- 5:28
And we'll give you some paths of how you may want to build with it. All right. So,
- 5:33
so yeah. Let's talk about some of the use cases now. So in terms of, uh, these are just some of the many possible use cases. Uh, these are use cases that are possible today.
- 5:43
Uh, this is our gallery app, which we have been demoing down on the third floor, and we also have our demo SAM session after this for anyone who's keen on learning more.
- 5:51
I do have QR codes for this too, so in the next couple of slides. Um, at present, you know, what's new is the dem-- the, the gallery app hel- allows you to demonstrate the agent skill capabilities, um- And also it has the AudioScribe capabilities or Ask Image, and also other features that were, uh, related to, uh, chat
- 6:13
experiences. Um, all of this is happening on-device. The purpose of this app that is created by Google is to help you get a playground to get a feel for what these models are capable of.
- 6:25
Uh, each of these capabilities has sample code as well that anyone can go and grab, and then you are able to also fork this app to build your own experiences.
- 6:33
But this is essentially to inspire and motivate you to build your own experiences.
- 6:39
So some of the next couple of slides I'll be focusing on is gonna be related to some of the Gemma 4 Edge use cases. Uh, these are gonna emphasize on, like, the voice agent capabilities or the local agent capabilities, and also a lot of these are going to be privacy-focused as well.
- 6:53
So there are three key pillars that are gonna be new here, or versus what was already supported in the previous slides.
- 7:02
So, let's dive into, uh, the Gemma 4 use cases. So let's see this place. So here's one use case where you're augmenting knowledge, and, you know, for example, you can build a skill to query Wikipedia, allowing the agent to query, uh, and respond to any encyclopedia question.
- 7:20
So this is a, this is a skill that's already available in our app that's, uh, that's there. So if you're looking to build something like this, this is, this is, this is possible.
- 7:29
Uh, these are new skills in addition to, like, basic on-device skills like summarization or Ask Image and things of that nature.
- 7:36
Um, this is another category-
- 7:40
Good morning, Gemma.
- 7:40
Where essentially, let's see if you can-
- 7:42
Let's call up Amy's mood journal entry with score nine and comment.
- 7:45
Same as-
- 7:45
"I got eight hours of sleep, and I'm looking forward to heading out with Amy today."
- 7:56
So here's you're gonna see creating and, uh, summarizes and display trends of hours of sleep. So you're gonna build a sleep in, a quick agent that helps you track your mood.
- 8:03
Perfect. Analyze the trend in my mood over the last seven days.
- 8:21
So you can kind of see that it's able to create this all on-device. It's taking your input, it's feeding it, it's able to un-understand. So the reasoning and the thinking capability are new in the new model, and that the, the on-device, uh, capabilities have gotten a lot more powerful.
- 8:35
Another similar ex-example is gonna be on expanding, like, the core capabilities. So for instance, you can also pair photos, and it's going to read the photo and have image understanding and generate music also all on-device.
- 8:47
So let's try this one.
- 8:48
Time for breakfast. I'm sending a photo. Can you pair this vibe with some music?
- 9:05
It's taking a few seconds, and then... [upbeat music] So a lot of these skills can be written by yourself on the app itself. So you don't even need to leave the app.
- 9:16
The app has instructions on how to do this. Uh, again, like, the whole sole focus of this is to help you understand, uh, the capabilities of the model and build something by yourself that, you know, you really like.
- 9:27
So that's useful. Uh, one last one is to really, you want to do sort of a working app that describes, let's say, vocal calls of animals. Like now, here we are trying to navigate multiple apps and users, and you can manage a more complex workflow here in this case.
- 9:43
So this is another example. And you can also change the CPU, GPU, and hopefully we have NPU support soon so you can decide the accelerator you want to use when you decide. [lion roaring]
- 9:57
So this was sound generation here from purely the prompt by the person. Again, the, the, the skill was set up and created and loaded, and the skill is running on-device, and this is for something similar. [laughing]
- 10:09
Right. So there's a lot you can do with it. There is a whole GitHub repo, uh, that we have, uh, where, where actually, uh, uh, there's like users are posting their own skills on the GitHub website and, uh, and they're able to basically share that with others, with the community.
- 10:24
So you're also welcome to explore that, or you're welcome to download one of these apps. Another option is, uh, the sample app is actually open source as well on GitHub.
- 10:33
So you're welcome to take it and fork it, and you can also make changes to that app if you like. So the GitHub link is also here, and it's an opportunity to do that.
- 10:42
And, uh, creating your own skill is also here. So these are just an examples of what folks have, have done. Uh, this QR code is instructions on how to build a skill.
- 10:51
So if you're looking to essentially get guidance on, okay, how do I do this, uh, you know, this, uh, you can kind of, uh, get started.
- 11:00
All right. Um, the next part of the talk is I'm gonna focus a little bit on deploying this now on edge devices. So we've talked about what's possible with the Gemma models.
- 11:09
Uh, the framework that is used here, right? So, so Google has, uh, an offering with LiteRT for bring your own models, essentially. So LiteRT essentially is Google's on-device framework.
- 11:22
It's built and built on the TensorFlow framework, TensorFlow Lite, if you guys are familiar with that. Is, uh, anyone ev-familiar with TensorFlow Lite? Some of you guys. Okay, awesome.
- 11:31
Okay. So that's what this is built on. Uh, it's really meant to also be built on using the same TensorFlow Lite model format, and that's what we are focusing on at this point.
- 11:40
So here, the, the point is that the underlying framework for the app that was running all these experiences is LiteRT, and then we are gonna show you that this is basically, it's been one of the most widely deployed frameworks so far.
- 11:55
It has, uh, hundred thousand plus apps. billions of active users and also lots of, uh, daily interpreter invocations. This is essentially the number of inferences that are happening every day.
- 12:06
Um, why this is interesting is to show that we are building on a trusted foundation, and when you do this, like, we also have our TFLite file format. And this is gonna be very important because if you were a developer that was building with TensorFlow Lite, your models are still gonna run on LiteRT.
- 12:22
And these models, the same model format is cross-platform. So it's not just Android, but you can run it on iOS, macOS, Linux, Win-- uh, Windows, web, and even IoT devices.
- 12:33
And I'll also show some examples on IoT that we were actually just building this morning, actually, so I'll show you that. Um,
- 12:41
cool. So from a development sign of-- uh, flow st- uh, standpoint and, uh, essentially you have your model files, so we are able to accept. The reason it's got branded to LiteRT, one of the bigger motivation for the rebranding was to also demonstrate that it's not just TensorFlow Lite models that we accept, but also PyTorch models and
- 12:59
JAX models. So if you are working with PyTorch models, you can take one of those models, convert it to the, uh, TFLite file format, and then go through this journey and deploy.
- 13:08
So we have a lot of sample apps and things like that, but I will not go through the whole journey here. But simply put, like, you have, uh, portability of the model and you have multi-framework support for the models.
- 13:22
All right. So the complete... This is a complete solution that provides one unified cross-platform architecture. Maybe a quick question. How many of you are looking to deploy on multiple Android or iOS devices?
- 13:33
Like, you want something that you build that you can test easily on many devices. 'Cause if you are trying to build apps that you're like, "Okay, I made it, but how do I know if it's gonna work on like five-year-old phones and six-year-old phones and all of that stuff?"
- 13:47
So we have options for that too. So I'll walk you through the stack. So LiteRT Torch is gonna be your conversion path. If you-- Basically, it's your bring your own model.
- 13:57
If you found a model, you like it, you wanna run it, you convert it to TensorFlow-- or TFLite format. You can quantize it if needed. Uh, if you do-- If it's LLM, you go through the LiteRT LLM path, uh, essentially, or...
- 14:09
And then if it's not, you can go through LiteRT. What's interesting here is that we also have the Model Explorer tool, which can help you explore the graph and decide which aspects of the graph you want to change and quantize.
- 14:20
So you can actually study the graph and decide how to best do mixed precision or, or basically complete, uh, convert the-- quantize the model as you like. The AI Edge Portal is a benchmarking tool which could be of interest for those who are looking to deploy broadly on Android.
- 14:35
So this is a cloud-based benchmarking service called AI Edge Portal. Um, this is available basically to help, and a lot of our, uh, uh, third-party app developers and even, like, internal developments use this tool essentially to get a good pulse check that, hey, you know, if I have a model that's so many parameters, do I need to
- 14:54
use ahead-of-time compilation or just-in-time compilation? What is gonna be my right recipe to ensure it's actually deployable across a broad fleet of devices and be reliable in that manner?
- 15:07
Um, yeah, sorry. One more thing on this slide is, uh, the next part is acceleration. So CPU and GPU are pretty universal right now. So these are, are essentially libraries that will help you run CPU and GPU.
- 15:17
I also want to emphasize back, like, with this sort of, uh, framework, you are able to deploy on multiple platforms, uh, with the CPU and GPU running. Uh, NPU acceleration is also another focus.
- 15:29
So we have completed, uh, integration with Qualcomm, MediaTek, and we are also focusing on additional integrations with other, uh, partners in the-- uh, on, on different platforms. Uh, we also offer, if you are looking...
- 15:43
I think this is gonna be a game changer for a lot of your apps who are-- or a lot of your products that you are trying to run a lot of ASR, TTS or any of these applications that are going to be requiring real-time capability or, or you're setting up some kind of AR/VR application where you want
- 16:00
to be able to have real-time improvements or, uh, updates to, uh, the camera feeds. I think the NPU is going to give you at least like three to ten X improvement in performance, and it's going to be a game changer in terms of the amount of energy you're gonna use, uh, and the amount of, uh, performance you're
- 16:16
looking for to unlock those use cases. So the flexibility part is we have the ahead of time or on device, and then for ease of use also there are many options here in terms of how we simplify, and there's a lot more, uh, documentation and support on that.
- 16:30
All right. So the next part is I'm gonna go over some performance numbers on how we are performing. So, so far we've given you an example of what you can do at the app level, uh, what's possible at the framework level, and how you can scale it.
- 16:42
Um, in terms of our coverage, right? So just with the Gemma models, like this is a-- this is the coverage right now. So a lot of times they come to the booth or, or demo booth and ask, "Hey, okay, so, you know, can you just do this on Android?"
- 16:54
But no, we've been-- we have tested this on all these platforms right now. So the models that you have on Hugging Face, you can actually test it on, on many of these platforms, so on Android, iOS, uh, Linux, Raspberry Pi as well.
- 17:07
Uh, and, and here's like, uh, just a quick, uh, demo we just did this morning. It's gonna take a while, but this is our, uh, little, uh, robot that's, uh, sitting down in our demo booth.
- 17:17
So we just made it this morning, so it's not super performant. But
- 17:21
what we did here is we showed a robot, uh, a sign to say, "Move your antenna." So now the two lights are blinking, so it's doing inferencing. Uh, and then, you know, you'll see it's running a Raspberry Pi.
- 17:32
It's running on a CPU, running on our, uh, LiteRT LLM. And you will see in a few seconds that it's going to wiggle the Sharpies [chuckles] that are its antennas.
- 17:40
And, uh, so there it is. So, so there are other questions like, "Will you marry me?" And things like that. If you're interested, you, uh, you can try it out downstairs.
- 17:48
It's there. Uh, cool. So, uh, yeah, that, that was amusing, but also, yeah, we, we do need to work on the performance because we just made it this morning.
- 17:58
Um- We also have a CLI tool. So for those who are looking to deploy and you want to have an easier way of doing it, so there is a new CLI tool that's been developed, so that is also available on our website in terms of, uh, with, uh, Python binding support and many of that.
- 18:13
So, so I'll just, uh... It's on our website, so if anyone's interested you'll find it.
- 18:17
Um, in terms of performance, so I have some quick performance numbers. Like, I just wanna emphasize, like, running on some of these, uh, new accelerators and CPUs can give you a big benefit advantage.
- 18:26
You can get up to thirteen X boost in some cases. Um, and also you have, uh, iOS performance as well, so we are supporting all of this here, so like roughly fifty-six tokens per second and so on and so forth.
- 18:40
Uh, desktop is here too. So a lot of these numbers are-- can be found, uh, on our Hugging, uh, Hugging Face page. So the quant is already listed there, so I'm just pulling from there.
- 18:49
So you will find it there. So when you download our models, you will get performance details on all these platforms and IoT. Uh, our runtime's also very performant versus Llama, at least on, like, mobile we've seen up to, like, thirty-five times faster performance.
- 19:03
On desktop it's at par, and then IoT also we have, like, three X performance. All right, so I have a minute left. So yeah, this-- In summary, yeah, LiteRT supports all these frameworks.
- 19:13
If you're looking to bring models from different fra-frameworks, you can. We have the models on Hugging, Hugging Face, and then we also have a gallery app, which is a nice playground for you to, uh, try this out.
- 19:22
Uh, and these are there. Yeah. So if there's any questions, uh, happy to take them. [audience applauding]
- 19:29
Yeah. Thank you.
- 19:30
Can you recognize faces? Like, uh-
- 19:32
Yes
- 19:33
... I'm thinking about, like, okay, the security camera at my home-
- 19:36
Yes
- 19:36
... recognizing that my son came home.
- 19:38
Yes.
- 19:39
I mean, because pushing this to the cloud would cost enormous amount of money. And since this can be done locally-
- 19:45
Right
- 19:46
... with local computer, maybe this is like a security feature of my, of my home camera.
- 19:51
Yeah. So on your phone, like on, let's say, even iPhones or Pixels, when you do face unlock, it's typically running locally already. So... And that is using the s-same framework in Google devices as well.
- 20:01
Uh, Apple has, uh, Apple's face unlock is also on device on, on CoreML. So, so that is possible, yes.
- 20:09
And then you would stream it constantly, like, um, every two seconds, for example, to recognize the face? Or h-how you would do that?
- 20:18
Typically, yeah. You have, uh, you have the camera on, and you can stream it. But, like, for, for your home camera, I think it would be a little different story, I think.
- 20:25
Yeah. So you would have to maybe do some sort of a algorithm to check the authenticity and, and kind of do that. So it'll be a little different from how it's deployed on the phone.
- 20:33
But yes, you can totally run that on device.
- 20:36
But that would consume a lot of bandwidth, right? You'd have to hook up a Pi to the camera, and only when that individual is recognized, then there's-
- 20:44
Then, then, then message your phone. Right. Right. Yeah.
- 20:46
Yeah. I'm thinking about, like, a local device-
- 20:48
Right
- 20:49
... connected with my camera.
- 20:50
Right.
- 20:51
The local device is actually checking the, the frames on the camera.
- 20:56
Yeah. Yeah. That would be a Ra- a Raspberry Pi.
- 20:58
A Raspberry Pi could do it. Right. Right.
- 21:02
Yeah.
- 21:02
You have a question?
- 21:03
Do you, um... You mentioned Orin as the target device.
- 21:08
Yeah.
- 21:09
Uh, how does LiteRT compare to TensorRT for execution on Orin?
- 21:13
Yes. Do you wanna answer that on-
- 21:15
Oh, you mean for the LiteRT? Uh, we actually use, uh, either WebGPU or just LGM, depending on the platform. That way actually it will be, uh, quite fast.
- 21:26
Do we have the numbers?
- 21:27
Uh, let me see. I don't know if I have Orin here. Uh, not here, but I think we have it on our Hugging Face. I don't think I added here.
- 21:33
Yeah.
- 21:34
Most of the time we have all the, most of the-
- 21:36
Um, have you tried any, um, exercise where you have multiple nodes which is running a model, and if any of them trigger something and they connect to a higher agent, another agent that-
- 21:51
Mm
- 21:52
... processes that input?
- 21:53
I think we have, like, dif- we have some groups that are working on, let's say, like, a speaker and a thinking agent type of architectures. Uh, I don't, um, have any examples to share, but, like, that is one way of, uh, dis-distributing, like, what...
- 22:08
And there's orchestration to decide what should run locally or not, and not, uh, and should run elsewhere. So, uh, I don't have examples here, but yes, we, there are, like, um, uh, different type of...
- 22:20
You can think of, like, a health agent, a coaching agent of sorts.
- 22:23
But it's a pretty common practice. Like, you have a classifier or train, like, a very small model just to-
- 22:29
Yes. Exactly
- 22:29
... determine whether the complexity of the request, if they're not go, go to-
- 22:32
Yeah.
- 22:33
Yeah. That's the... Actually, that's pretty common practice already. So yeah.
- 22:37
Yeah.
- 22:37
Sorry to interrupt. Uh-
- 22:38
Next one
- 22:38
Thank you. Yeah.
- 22:40
Uh, I have a mobile app which is currently using Gemma Lite APIs for audio to audio models.
- 22:47
Right.
- 22:47
Is there any recommendation on this? This seems like audio to text models are-
- 22:52
Right. Right. Right
- 22:53
... so is there any recommendation on that?
- 22:55
So if we can-- If you have open, open-weight models that you prefer, like, uh, we are able to, uh, support any of them at... I mean, hopefully, we just need to make sure we get them in the right file format and, uh, you know, if the sizes are...
- 23:08
So we should be able to, if you know which models you are thinking about, uh, uh, for your application.
- 23:15
Not these, uh, at least Gemma 4B, 4B.
- 23:19
Here, here, we've just focused on the Gemma because of the DeepMind track. But essentially our Hugging Face page has other open-weight models, and you are-- You can... So some of them we provide for ease of use.
- 23:30
Others, you are welcome to, like, convert and use yourself as well. Uh, but yeah, if there is a model that you've, uh, identified, yeah, then we can certainly discuss.
- 23:40
Thank you.
- 23:40
Yeah. Thanks. Bye. [upbeat music]