AI Engineer Summit 2025
How Deep Research Works
About this talk
Google DeepMind presenters Aarush Selvan and Mukund Sridhar explain how Gemini Deep Research combines increased inference-time compute, web research, user-visible research plans, and iterative agent planning to produce comprehensive reports. They discuss building asynchronous experiences inside a synchronous chatbot, handling failures across long-running multi-service workflows, supporting cross-platform notifications, and managing research context through recency-aware notes stored in RAG. They conclude with a vision for assistants that move beyond information aggregation toward strategic, profession-specific analysis.
Chapters
- 0:00Introducing the presenters and Gemini Deep Research
- 1:56Inference-time compute, asynchronous UX, and research planning
- 5:39Long-running agent execution, failures, notifications, and iterative research
- 11:30Recency-aware research notes and retrieval-augmented memory
- 12:12From research analyst to strategic, personalized research partner
Talk transcript
- 0:00
[on-hold music] Hey, everyone.
- 0:17
I'm Aarush. I'm a product manager here at Google.
- 0:19
Hey, I'm Mukund. I'm a software engineer at Google working on Deep Research.
- 0:22
Um, so, uh, I don't know if people have had a chance to, uh, try Deep Research on Gemini, um, or are familiar with the product, but you can try it if you go to Gemini Advanced.
- 0:34
And if you scroll past 2.0 Flash, 2.0 Flash Thinking Experimental, 2.0 Flash Thinking Experimental with apps, 2.0 Pro Experimental, you will find, uh, 1.5 Pro with Deep Research, which is what we built.
- 0:47
Um, and if you have the chance to use it and you paid the twenty bucks, uh, you will see that it's a personal research agent that can browse the web for you to, to build, uh, reports on your behalf.
- 0:58
And so our motivation and what we wanna talk about today is kind of why we built it, some of the product challenges we overcame, and some of the technical challenges you'll face of building a web research agent.
- 1:07
Um, so our motivation was really we wanted to help people get smart fast. Um, we saw that research and learning queries are some of the top use cases in Gemini.
- 1:17
But when you bring, like, really hard questions, uh, to chatbots in general, what we were finding is that it would often give you a blueprint for an answer rather than actually give you the answer itself, right?
- 1:29
So we had this query that we used to throw around of like: Tell me what does it take to get an athletic scholarship for shot put, and, like, how do I go get one?
- 1:38
And often the answers would be things like, "You should talk to coaches," "You should find out how far you should be able to throw," and, you know, uh, "You should make sure you have good grades."
- 1:46
But really what I wanna know is like, okay, what are the grade boundaries? Like, how far do I need to actually be able to throw? I want something super comprehensive, and, and that's where we saw a big opportunity.
- 1:56
Yeah, so we said, what if you remove the constraints of compute and latency at inference time? Let Gemini take as long as it wants, browse the web as much as it needs, and see if we can trade that off for a much comprehensive answer for the user.
- 2:10
But you gotta do it in five minutes, 'cause beyond that, uh, we don't have the chips. Um, so, uh,
- 2:18
this brought a bunch of product challenges for us. Um, Gemini, up to this point, is an inherently synchronous feature. It's a chatbot. Um, and so you wanted to-- we needed to figure out how do you sort of build asynchronous experiences in, in an inherently synchronous product.
- 2:33
Um, you also wanted to set expectations with users, right? Deep Research is good for, like, one very specific thing, but a lot of user queries to Gemini are things like, "What's the weather?"
- 2:40
"Write me a joke." Things like that, where waiting five minutes is not gonna get you a good answer, and we wanted to set expectations. Uh, and the last thing is our answers can be thousands of words long, and we needed to figure out how do you make it easy for users to engage with really long outputs and,
- 2:56
um, uh, in, in a chat experience. Um, so let's walk through kind of the UX and kind of think about how, how we solve some of these, right? So imagine you're a VC, uh, and everybody's talking about, you know, investing in nuclear in America.
- 3:11
And so you come with this query like, "Hey, help me learn the latest technology breakthroughs in small nuclear reactors and tell me interesting companies in the supply chain." So the first step, um, when you bring this query to Deep Research is that Gemini will actually put together a research plan for you and present it in a card.
- 3:26
And so this is the first way in which we're able to communicate with users. Like, this is different. This isn't your standard chatbot experience. Something's gonna happen. You're gonna hit Start.
- 3:35
But it's also an opportunity for us to actually show the user a research plan that they can edit and engage with. Kind of like a good analyst, right? They, they wouldn't just get to work.
- 3:42
They'd actually show you, "Okay, here's how I'm gonna approach this." And it's a way for users to, if they want, kind of engage and steer the direction of the research further.
- 3:51
Now, uh, once it-- you hit Start, we actually try and show you, um, what Gemini is doing on the-- under the hood in real time, uh, by showing you the, the websites it's browsing.
- 4:03
And this is a feature that was built before thinking models, and thoughts are also a really great way of kind of showing transparency of what the model is thinking.
- 4:11
Um, but what's really nice here is while you wait, you can sort of click through the websites, dive into any of the content. Um, but what we also inadvertently saw is people trying to game that number to see how high it could go.
- 4:22
So we definitely saw people push that number into the, into the thousands, uh, to try and, um, you know, see how many websites Deep Research could read.
- 4:30
Um, finally, we kind of get this report that's, you know, thousands of words long. And, um, we're really inspired by what, kind of what Anthropic does with, um, artifacts.
- 4:40
And so we thought that was a really great way of sort of being able to pin an artifact so that users can actually ask questions about the research while reading the material.
- 4:49
They don't have to scroll back and forth. And what's really neat about this is it means it's easy for you to engage in sort of changing the style of the report, adding sections, removing sections, asking follow-up questions, and, uh, and it sort of makes that really easy.
- 5:02
And the last part that's super important is kind of user trust and also doing right by the publishers. So we, we try and always show is all the sources we read as well as all the sources we used in the report.
- 5:12
'Cause not everything that we read is used, but it stays in context for follow-up questions. And, and also sort of these are all things that, um, carry over to Google Docs as citations and things like that if you choose to export.
- 5:29
Uh, so I thought today we can pick some of the challenges, uh, that one has to encounter while building a research agent and ta-talk through some of them. So, uh, I picked four for today.
- 5:39
So one is this, this long-running nature of tasks introduce-- This is a couple of things that we need to look into. Second is the model has to plan iteratively and spend, uh, its time and compute during this time effectively.
- 5:55
So what are those challenges there? And it has to do this, uh, while interacting with A very noisy environment that is the web. And as you do this and, uh, read through information, very quickly you can start seeing your context grow and h-how do you effectively manage context.
- 6:14
So if, if you think about a job that runs for multiple minutes and something that can make many, many d-- uh, different LLM calls and calls to different services, there are bound to be failures, right?
- 6:26
And today we are talking about, oh, of minutes, but you can very easily think in the future of, uh, these kind of research agents taking like multiple hours. So it's important to be robust to intermediate failures of these various services of various reliabilities.
- 6:41
And so being able to build a good state management solution, being able to recover from e-errors effectively so that you just don't drop the whole research, uh, task due to one failure.
- 6:53
That's one. The second aspect of doing this, what it enables us, is to enable this feature, uh, cross-platform. So we believe more and more, uh, users will start kind of registering your asks, uh, or your research tasks and just like walk away, do their thing, and then you need to get notified.
- 7:11
And this can happen now across, uh, devices and you can pick off, uh, uh, reading it, uh, uh, uh, once it's done.
- 7:21
So now what is the model doing at, uh, like through these, you know, uh, few minutes? Uh, so let's take, uh, example, right? So here, uh, we're looking for, uh, athletic scholarships, uh, for shot put.
- 7:33
There are many facets to this query, and we kind of show this in a research plan like Aarush showed. The first thing the model has to do is try to figure out which of these sub problems it can start tackling in parallel versus things that are inherently sequential, right?
- 7:49
So the model has to be able to reason to do that. And, uh, the other challenge is, here you see you're always gonna land in this state where there's partial information, so it's important to look at all the information found so far before you decide what to do next.
- 8:06
So in this instance, the model found, hey, it's-- it knows the qualifying standards, uh, for the D1 division, but in order to provide a complete report and answer the user's question, it has to go figure out what the equivalent for the D2 and D3 divisions are.
- 8:22
So this notion of being able to ground on information you find and then plan your next step is key.
- 8:31
Another example of partial information could be when you make searches. Uh, so in this case, we're trying to find the best roller coaster, uh, for kids. Uh, you might find results, uh, that provide partial information again.
- 8:44
So here, uh, you end up at a link, uh, which talks about the top ten roller coasters, but does not mention anything about them being suitable to kids. Uh, so the plan has to recognize this fact and then go ahead and in the next steps of planning, try to resolve this, uh, disambiguity.
- 9:05
Um, another example of, uh, challenges in planning is information is often not found in one place. You find facets of information spread across different sources. So here, uh, we are trying to find, uh, what would, uh, what would it take to get a certification for a scuba dive, uh, in, in, in some dive centers nearby.
- 9:26
So you see, uh, one part or one source has, uh, the kind of the structure of, uh, what, what you have to go through to get a certification, but in a completely different source, you have this notion of the pricing for this diving center.
- 9:40
So the model has to weave this together to figure out, um, you know, what the cost structure for such a certification would look like.
- 9:48
Then there's the classic, uh, entity resolution problem. So you might find mentions of the same entity across different sources, so you need to be able to reason about some information indicators to kind of figure out if they're talking about the same entity or you need to explore more to verify such, uh, disambiguities.
- 10:09
Um, yeah, I think m-most people here have worked on some notion of a web problem, and we know like it's super fragmented. So, uh, here you see two different websites, uh, talking about the same thing, uh, about music festivals in Portugal this year.
- 10:24
Uh, on the left, uh, if you end up at such a website, it's easier and you get most of your information in one go. Uh, on the right, uh, the layout is different, so having a robust, uh, browsing mechanism if you wanna navigate, uh, the web for your research tasks is another, uh, important challenge.
- 10:44
So like we saw, there is a lot of these intermediate outputs, and as you do this and you start getting streams of information during your planning, you can imagine your context size growing very quickly.
- 10:57
Um, the other challenge that, uh, about context size is your research task doesn't typically end with your first query. People have follow-ups. People can say, "Hey, can you also do the same for this other topic?"
- 11:10
So there is like this kind of a follow-up, uh, deep research and, uh, that also adds pressure on the context. Uh, we at Gemini have, uh, the liberty of really long-context models, uh, but, uh, even then you have to design, uh, some way to make sure you, you effectively manage your context.
- 11:30
And there are multiple choices here, each come with various different trade-offs. Uh, we're showing one here, uh, where we kind of have like this recency bias. So you have a lot more information about your current and your previous tasks, but as you get to older tasks, we kind of selectively pick out, uh, you know, things what we
- 11:50
call as research notes and put it in a RAG. That way, the model can still access it, but it's being selective. Uh, I'll hand it back to Aarush about, uh, to talk about what's next.
- 12:00
Yeah. So we were super excited To put this feature out in December, we weren't actually sure if anyone was gonna use it, if anyone was gonna care, um, uh, to wait five minutes, uh, for something.
- 12:12
And, uh, we were really positively surprised by the reception. Um, and, and really what we, what we saw, um, was, hey, we, we've built something that's maybe as good as, like, a McKinsey analyst, right?
- 12:23
And we give it away for 20 bucks. But, um, you know, that's, that's really great and, uh... But what it does is it just retrieves from the open web, and it's a text in, text out only system, right?
- 12:34
And so where we sort of, we sort of see a few different directions of where research agents are gonna go next, and the first one is around expertise, right?
- 12:42
So how do you go from a McKinsey analyst to a McKinsey partner or a Goldman Sachs partner or, like, a partner at a law firm, right? So that's really around not just being able to aggregate information and synthesize it, but also think through the "so what" of how do-- like, what are the implications for what we're gonna
- 12:58
do and, and what are the most interesting insights and patterns that come out of it? The, the other thing is, you know, there are plenty of domains beyond professional services, like the sciences, where you, you know, wanna get really good.
- 13:09
You know, you want something that can read many papers, form hypotheses, find really interesting patterns in, you know, what methods we used, uh, and, and come up with novel hypotheses to explore.
- 13:20
However, um, just because you build something that can be really smart doesn't mean that it's useful to someone, right? So, um, if we were thinking about a use case of running a due diligence on a company, the way you'd present that information to me would be very different to the way you'd present that information to, say, a
- 13:36
Goldman Sachs banker, right? Um, for me, you really wanna talk through, like, what, like, what is this company and how is it positioned strategically? But a banker would want to know all the financial information, actually have a DCF that they could look at, right?
- 13:50
Actually, uh, have a, have a much more, like, fine-grained, uh, sort of, uh, finan- uh, financial modeling and analysis. And, and that really should shape the way in which you browse the web, right?
- 14:00
The way you browse the web, the way you frame your answer, the kind of questions you pursue should be very personalized to kind of meeting the user where they're at.
- 14:07
I think the last part is sort of something that goes across domains of what models can do, right? So not just being able to do web research with text, but being able to combine that with abilities in coding, data science, even video generation, right?
- 14:19
So coming back to this example, if you're doing a due diligence, y- what if it could go and do, like, a lot of statistical analysis and actually build financial models to inform the research output that it gives you, right?
- 14:29
Telling you, "Hey, why is this a good company or not?" Um, I should say Google doesn't give financial advice and- [laughs] ... you know, it's not a financial advisor. Um, but yeah.
- 14:39
And so we're really excited about the potential. We think there's a ton of headroom to make research agents better, and we are really glad we didn't call this Gemini Deep Dive, which was [laughs] our best name before, uh, before launching this feature.
- 14:51
Um, that's it. Thank you so much.
- 14:54
Thank you. [clapping] [upbeat music]