AI Engineer World's Fair 2026
SOTA Generative Media Panel — Dumitru Erhan, Shane Gu & Nicole Brichtova, Google DeepMind
Read the talk
Generative Media Beyond the Demo: Representation, Realism and Real Workflows
Dumitru Erhan, Shane Gu and Nicole Brichtova discuss how multimodal models generate and edit media, why human preference can mislead evaluation, and what working creators reveal that finished training assets cannot.
From a talk by Dumitru Erhan, Shane Gu, Nicole Brichtova and swyx
At a glance
Ideas worth remembering
Text supplies useful structure and access to pretrained knowledge, while media references and video models carry sensory information that words may underspecify. The panel argues for their combination and leaves the ultimate representational limits unresolved.
Joint audiovisual generation has a clear causal motivation: visible motion and sound arise from the same event. Convincing output must also preserve scene-dependent properties such as distance and room acoustics.
A preference win does not establish realism or task success. Attractive sharpness, saturation and skin tones can mask errors, while expert inspection can reveal subtle failures or recurring artifacts such as unwanted wedding rings.
Video evaluation needs several kinds of evidence: automated checks for tractable errors, human evaluations covering thousands of items, live experiments and feedback from real workflows. Small-group comparisons are one part of that system.
Creative task trajectories reveal requirements absent from finished media. Forward deployed engineers and direct user conversations can turn decisions about pattern consistency, physical scale and exact brand colors into feedback for model development.
Faster generation changes the iteration loop
The panel opens as a chance to examine generative media beyond a mainstage launch. After introducing researchers and product work spanning video, Gemini reinforcement learning and image generation, swyx asks what developers can actually try. Nicole Brichtova describes two launches: Nano Banana 2 Lite and the Gemini Omni Flash APIs.
Brichtova presents Nano Banana 2 Lite as the fastest and cheapest image model in its family, with generation and editing quality above the original Nano Banana and approaching the larger models. Her practical emphasis is roughly 3-second latency: shorter waits make it easier to explore an idea, inspect the result and revise it. She says some outputs can also serve as production assets, so the fast model is useful beyond preliminary sketches.
The Omni Flash API launch makes video generation and editing accessible to developer workflows. swyx illustrates why that access matters with an earlier edited podcast featuring added animals and other objects: he wants similar transformations for his own videos, but needs an API to automate the work. The playful example establishes a concrete distinction between an impressive demonstration and a capability that can be incorporated into a repeatable process.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Storyboards, editing and translation
Asked for workhorse applications, Brichtova identifies two central capabilities. First, a model can accept different kinds of references and produce video: a set of images supplies a storyboard, while an audio track supplies a voice for a character. These inputs let creators specify aspects of a scene through examples. She connects this to short film production and creator workflows, while describing additional output modalities as a future direction.
Second, natural-language editing lets someone request additions, removals or cleanup in an existing video. A noisy beach vacation recording is her everyday example: the user can ask for the noise to be removed without first identifying a specialist tool. Marketing campaigns are another observed application. Educational materials extend the idea toward content adapted to a learner’s knowledge and preferred style, although that broader personalization is presented as a direction rather than a completed system.
Dumitru Erhan offers a specific image-editing example. His visiting parents needed instructions for a gadget, but the illustrated instructions were in English. He photographed the page and asked for Romanian text while keeping everything else the same. He reports that the layout remained effectively identical and qualifies the translation as more or less correct. The useful mechanism combines language understanding with text rendering inside the existing visual structure. He sees related possibilities in video localization, translated on-screen text and redubbing.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Combining symbolic reasoning with video models
swyx introduces video agents as an alternative to asking a single generation pass to accomplish everything. Shane Gu responds by focusing on cooperation between symbolic foundation models and video foundation models. Detailed language descriptions provide a useful shared representation, but his stronger hypothesis concerns spurious correlations: a predictive feature need not be a cause. Diverse training examples across interventions can help distinguish the two; conditioning on a description of what is happening may also supply information about the factors that produced a scene.
Gu then describes evaluation work on video models as zero-shot learners and reasoners. His argument is that learning to generate spatial and temporal information can support more than media production: the resulting model can attempt classical vision tasks, visual quizzes and tasks requiring physical intuition without task-specific training. He explicitly leaves substantial room for improvement. The important possibility is to combine this visual reasoning with text reasoning, rather than treating generated video solely as a final artifact.
Whether that combination lives inside one model or in an agent coordinating models remains an incremental engineering question. Gu imagines eventual consolidation, but says there is already considerable scope in connecting Gemini’s image and video understanding to Omni through an agentic workflow. That is a research direction under exploration, not evidence that a unified system has already replaced the separate components.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Why specialized models still have a role
The suggestion of one eventual model prompts a more pragmatic discussion of product boundaries. Erhan points out that a fast image model serves a different niche from a model designed for something like 4K, 30-second video. Their training and serving requirements need not fit the same checkpoint. His contrast between possible convergence in five years and continued specialization in six months is a forecast, not a product commitment: engineering, research and product tradeoffs still justify multiple models.
Brichtova says the Gemini Omni name signals an ambition for fully multimodal inputs and outputs, including eventual image generation and editing. But architectural consolidation also depends on transfer: does learning one task improve another enough to justify training them together? The panel sees a clear relationship between images and video, and a strong reason to generate audio with video. Transfer between coding and video generation, or coding and 3D representations, is less obvious. Combining tasks could help, or could consume resources without a corresponding benefit.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Language is useful structure, but not the whole world
swyx asks whether captions are the right intermediate representation for video. Describing a scene over time in English can feel inefficient, especially when some video can already be generated through code. Gu confirms that coding representations are being explored, then connects the question to a broader one: why should a model’s intermediate reasoning use natural language instead of continuous tokens or another form of additional computation?
His answer rests on the knowledge acquired during pretraining. In the recipe he describes, large-scale pretraining supplies much of the model’s intelligence, while extracting capabilities through reinforcement learning can be computationally expensive. Keeping reasoning in natural language lets intermediate steps draw directly on that pretrained knowledge. An unconstrained representation might support computation, but it does not automatically retain this convenient connection to what the model has already learned. The panel also gives a product reason for text: it is a familiar way for people to communicate their intentions.
The discussion then narrows the increasingly broad term world model. Gu uses it to mean the model in model-based reinforcement learning, while acknowledging that the term has acquired competing meanings. When swyx objects that language is a narrow, lossy channel, Gu clarifies that language alone is not the proposed solution. Video and language should work together, with video supplying a complementary foundation for aspects of the world that text does not adequately represent. His ambition extends from attractive clips toward more general intelligence, but remains a research vision.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Understanding and generation can improve each other
The migration of researchers from recognition into generation is more than reversing an image-to-text mapping. Erhan uses a cat as the example: recognizing a cat in an image is a simpler problem than turning the category cat into a particular image. Generation must resolve the many appearances and arrangements compatible with the same description. Better understanding nevertheless helps generation, including through the synthetic labels that vision systems can produce.
Gu recommends learning recognition and understanding because the ability to discriminate quality can support better generation, with reinforcement learning acting as a bridge. His own path moved through generative modeling, robotics and dexterity before language models, motivated by his expectation that symbolic intelligence would advance faster than physical intelligence. He encourages researchers to learn how different research communities frame their problems.
He sees a possible parallel with language models: early creative demonstrations became useful chatbots through instruction tuning, then more reliable reasoning systems as pretraining and post-training improved. Video models might similarly improve instruction following and reliability enough to support spatial and temporal simulation alongside textual reasoning. Erhan adds a present limitation: multimedia understanding and generation have not yet been thoroughly unified in models that excel at both. Their conceptual relationship does not by itself settle how they should be implemented.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Audio and video share an underlying event
Asked whether audio is fundamentally different from video, Erhan says the technical differences seem relatively minor from his perspective. The more consequential choice was to generate audio and visuals jointly. He describes the team’s model as producing both together, rather than generating pictures and then attaching a separate process to make the lips fit separately generated speech.
The rationale is causal: a person speaking is one underlying event that produces both visible movement and sound. Lip motion and audio therefore need to agree in time. Modeling them together gives the system a way to learn that shared structure and avoids some synchronization failures of a pipeline assembled afterward. Erhan regards joint generation as a successful design choice and says that once users experienced generated video with sound, silent output became much harder to accept.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
What a reference can convey that words struggle to specify
Gu identifies a difficulty beyond specifying spoken words: describing music, a person’s vocal tone or other sensory qualities. He compares this with taste, smell and subtle differences in skin appearance. His hypothesis is that some perceptual sensitivities are tied to primitive survival functions and are finer than ordinary vocabulary. He illustrates the vocabulary gap with a professional wine taster who borrowed language used to describe a romantic partner. This is an explanatory hypothesis and anecdote, not a demonstrated account of the limits of language.
Brichtova extends the same concern to visual style and aesthetic judgment: people can perceive distinctions they struggle to describe. Erhan is more cautious about interpreting that as a fundamental ceiling. The audio failures he encounters still look amenable to more or better data, and he does not think the frontier has been pushed far enough to establish an inherent sensory-language limit.
Direct references offer a practical response. A sample voice can communicate tone and prosody without requiring the user to master specialist vocabulary. There are also two distinct sources of difficulty: the user may lack the words, or the language model may have weak understanding of the relevant domain. If generation depends on that understanding, the weakness propagates into the output. The panel leaves open how much of the problem comes from insufficient attention and training rather than an unavoidable representational limit.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Convincing scenes need acoustics and microexpressions
Drawing on podcast production, swyx divides audio roughly into music, voice and sound effects, then examines voice alone. A large room, a small room, a car and a phone connection all change what a listener hears. He argues that uniformly studio-quality sound can betray generated video and attributes that tendency to training material. His concrete requirement is spatial consistency: someone farther away should sound softer or more diffuse. Immersive audiovisual generation therefore needs to represent how a scene affects sound, as well as what is being said.
Gu connects this to conditional generation. If a caption leaves out acoustics and other relevant details, many substantially different outputs remain compatible with the same words. He describes a modeling ideal in which a latent representation captures most of the variability, leaving the output given that representation comparatively deterministic. In that framing, richer descriptions help by specifying factors that would otherwise remain ambiguous; they do not establish that language can express every factor.
Brichtova identifies an analogous visual gap in facial expressions, skin texture and the small reactions people read during conversation. A nod or changing microexpression communicates a response to another person. She says video generation has improved substantially but still has considerable headroom in these details. Some still images already appear indistinguishable from reality to her; maintaining that impression through a person’s changing behavior is a harder remaining challenge.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Preferred does not necessarily mean realistic or useful
Erhan describes an experiment in which the team took descriptions of real videos, generated corresponding videos with Omni and asked humans to compare them. People largely preferred the generated versions. He immediately qualifies the result: the generated clip could be sharper, have a more HDR-like appearance and offer more pleasing skin tones without being more realistic or solving the user’s problem. The account supplies no sample size or numerical preference margin, so it supports a warning about the meaning of preference rather than a quantified performance claim.
Perceptual expertise also changes the judgment. Gu recounts a manga artist’s discomfort with subtly incorrect eye gaze: a small directional error can make an otherwise polished image feel unnatural. Erhan’s conclusion is that simply asking whether people like an output is an unreliable guide to what should be optimized. A broad preference score can miss exactly the distinction a skilled practitioner needs.
This is why Gu expects deliberate prompting to remain useful even as models automate more of it. Control depends on noticing a difference between the current output and the desired result, then communicating that difference. Brichtova distinguishes an ordinary viewer’s preferences from the sensitivity developed through years of design, architecture or illustration. Smooth, saturated output may win casual approval—the discussion calls this the Instagram filter effect—while an expert wants something else. Better instruction following and support for references should let users move away from the default aesthetic.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Aesthetic defaults create visible and invisible biases
Default aesthetics become consequential when many users accept them. Brichtova recalls seeing widespread Nano Banana Pro infographics and feeling that their default presentation was too cluttered. The model seemed eager to fit everything it knew about a concept into one image. Her observation is personal rather than a measured count of adoption, but the design issue is concrete: displaying more information can undermine the usefulness of an explanatory graphic.
For Omni, the team explicitly compared styles during late tuning, including muted versus saturated color and different palettes. Those choices often land with the modeling team, which raises the question of whether people with an art director’s expertise should play a larger role. Trusted testers and internal users already supply substantial feedback. One example is an optimization that seemed acceptable to the team but made the detail in a user’s grass blurry.
Another user noticed that the model tended to place wedding rings on generated hands, a pattern the developers had not recognized. Gu suggests reward hacking through a spurious preference correlation as a possible explanation. The causal diagnosis is unresolved: the panel does not establish which training signal produced the rings. The example shows how a recurring, semantically meaningful detail can pass unnoticed during development and become obvious to someone inspecting outputs with a different focus.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Evaluation combines automation, human judgment and workflow tests
The evaluation discussion distinguishes relatively objective checks from aesthetic judgments. Rendered text in an infographic is a tractable example: OCR can extract the text, and a malformed or incorrect letter may make the asset unusable. Automated evaluation of video aesthetics is much harder. Improving Gemini’s understanding can help build better evaluators, but the team still relies extensively on people examining outputs.
For close model choices, Brichtova describes rooms with 10 people watching videos side by side and discussing which they prefer. That is one decision mechanism within a larger evaluation program. When swyx interprets the coverage as only hundreds of examples, she corrects him: human evaluations cover thousands of items. Live experiments add larger-scale evidence, automated raters supply another signal, and experienced users contribute judgments grounded in daily work.
Free-form video editing makes coverage especially difficult because the space of requests is so broad. A model can add objects, change sound or perform other transformations for which no dedicated evaluation yet exists. Strong performance on one slice of human tests can coexist with a broken customer workflow. Early access programs help expose those failures before broader release, making task-level feedback a necessary complement to aggregate scores.
Gu wants human evaluation effort to become reusable through better machine understanding. Detecting errors in generated video is itself a demanding intelligence task. A recreated movie scene might look plausible while becoming semantically inconsistent; identifying the inconsistency requires more than recognizing visual polish. His proposed direction is to turn human judgments into capabilities that can detect such failures, gradually reducing the need to repeat the same manual work.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Finished media leaves out the decisions that produced it
Asked what data the team wants, Erhan gives a deliberately broad answer: high-quality media, including professionally shot material. He notes existing interest in robotics and embodied data but avoids turning the answer into a detailed account of future projects. The stated need is therefore a quality direction, not a specific procurement specification or a claim that more undifferentiated footage would solve the remaining problems.
Brichtova identifies another valuable kind of data: the trajectory of an actual creative task. A marketing workflow might start with a product photograph, produce a video ad and then adapt it into assets for several advertising formats and platforms. The intermediate requests and decisions reveal what the system must accomplish across the whole job. These trajectories are difficult to manufacture in a lab or obtain through a vendor that lacks the product surface where people do the work.
swyx argues that media workers have abundant routine work they would welcome help with. Brichtova adds that even a familiar marketing campaign involves craft: people iterate, choose one candidate over another and reject details such as incorrect eye gaze. The final asset alone does not explain those choices. A model developer who is not a marketing director may never think to ask for the distinction that determines whether an asset is usable.
Gu calls attention to knowledge that remains inside people: inspirations, conversations and the sequence of choices behind a finished paper or creative work. His assertion that 99% of information is inside people is a rhetorical estimate, not a measured statistic. The substantive distinction is between publicly visible outputs and the process that produced them. The discussion extends this to fiction, where repeated default language patterns can produce readable prose without the personal connection a reader finds in a distinctive story or character.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Bring real task failures back into model development
The closing discussion gives forward deployed engineers a role in obtaining this missing knowledge. Gu describes growing investment in engineers who work closely with users, while swyx argues that their work should feed evaluations as well as solutions. Gu frames the role as a two-way relationship: customers can improve the harness around a model, and the model team can use the same observations to improve upstream behavior. His unusually broad definition of post-training includes everything between pretraining and the final user experience; it explains his emphasis on connecting customer feedback to modeling.
Brichtova supplies concrete failures that direct conversations reveal. An interior-design user wants the same pattern reproduced across 10 different rug sizes, including a custom size, but the model does not preserve the pattern correctly. A virtual earring try-on needs the earring’s size to make sense relative to the wearer’s head. These requests test consistency and physical scale, not merely whether the output looks attractive. The team benefits from hearing them because its members do not routinely use the models for those jobs.
Brand requirements expose another level of specificity. A brand language may be expressed through images and PDFs, and a broad description such as IKEA’s blue and yellow does not capture the exact shades that matter. The example is illustrative, but the requirement is practical: a model must preserve the distinctions on which a customer’s task depends. Erhan says the goal is to build products people can use for concrete work, which requires understanding what those users care about.
swyx closes by welcoming the progress in generative media and calling for continued exploration beyond coding. The panel ends with appreciation for the work already done and an expectation that substantial development remains.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Read the complete timestamped transcript
- 0:01
[music]
- 0:12
and welcome back for those on the stream
- 0:14
and those those in person. um we take
- 0:17
tend to basically take these longer
- 0:19
sessions between uh all the sort of
- 0:21
mainstage keynotes to reflect on things
- 0:25
that um you know are particularly
- 0:27
important but like don't have like a
- 0:29
significant like sort of launch moments.
- 0:31
Today we're very lucky to have people
- 0:32
working on Omni and VO Nano Banana like
- 0:36
the you know the world's best generative
- 0:38
models here with us. Uh, Demetrio, I I I
- 0:41
first saw you when you were posting
- 0:43
about your office. [laughter]
- 0:46
Um, I think you're you're probably
- 0:48
number one uh Google Google's number one
- 0:51
office influencer at least in in San
- 0:53
Francisco. I think you like you like to
- 0:54
bike as well. You like to take photos of
- 0:56
>> bike here.
- 0:57
>> Yeah. Um, but you know, but also you
- 1:00
work on video models.
- 1:01
>> That's right.
- 1:02
>> Um, Shane, I I met you I think at like a
- 1:04
dinner.
- 1:05
>> Yeah. Um and uh and uh and I I remember
- 1:10
you were trying to get me invested in
- 1:12
like one of the companies. I forget
- 1:13
forget which one.
- 1:14
>> Forget about that. [laughter]
- 1:17
>> But now but now you're um now you're
- 1:20
working on Omni Thinking. Um and and
- 1:22
just you know a bunch of other
- 1:24
>> Gemini RL.
- 1:25
>> Yeah. Yeah. Uh and Nicole also uh the
- 1:28
rest of the gen media models uh nano
- 1:31
banana and uh all and everything you
- 1:33
just launched actually even this week.
- 1:34
Uh,
- 1:35
>> yeah. We launched some APIs.
- 1:37
>> Yeah. Yeah. Yeah.
- 1:38
>> And I haven't tried to convince you to
- 1:39
invest in anything, but maybe I should.
- 1:41
>> I mean, so I try not to be an investor.
- 1:44
People just convince me anyway. I'm like
- 1:45
just, okay, well, I'm not that rich, but
- 1:47
know like you can't not try to invest in
- 1:50
some of these things. And, you know, for
- 1:52
those of us who are not working at a
- 1:53
Frontier Lab, this is the best this
- 1:55
closest we'll ever get. Um, so yeah,
- 1:57
actually, let's kind of recap since
- 1:59
you're closest to it and we just did it,
- 2:00
like what was launched this week? What
- 2:02
should people go try out?
- 2:03
>> Yeah. Um so yesterday we had two launch
- 2:06
moments. Uh one of them we launched
- 2:08
NanoBanana 2 light uh which is our
- 2:11
fastest, cheapest um image model in the
- 2:14
nano banana model family. Um and it's
- 2:17
better than the original NanoBanana. Um
- 2:19
so really for most people um that model
- 2:21
replaces what you you know used and love
- 2:23
the original Nano Banana for across like
- 2:25
generation and editing and it gets
- 2:27
really close to the frontier quality of
- 2:30
of the kind of mainland bigger models.
- 2:32
So that that's really exciting. I think
- 2:33
if you look at some of the demos or like
- 2:35
things that people have been trying like
- 2:37
getting kind of that like 3 second
- 2:38
latency just unlocks a whole bunch of
- 2:40
things that you can do with like
- 2:41
ideation and iteration and it's just
- 2:43
really fun and the model's getting to a
- 2:45
point where like the quality is really
- 2:47
good um where um it you know you can use
- 2:50
it for iteration but you can also use
- 2:51
some of those outputs as just kind of
- 2:52
like ready um production output. So
- 2:54
that's really exciting. Um and then
- 2:56
second launch we finally um launched the
- 2:59
Gemini Omni Flash APIs um that we
- 3:01
pre-announced at IO. So thank you for
- 3:03
waiting. Um and that you know is the
- 3:08
first time that we're making the APIs
- 3:10
available for developers and it's
- 3:11
basically really exciting kind of video
- 3:12
generation and editing and we're pricing
- 3:15
it the same as Y31 fast. So we're
- 3:17
getting you kind of like really really
- 3:18
good quality for a really awesome price
- 3:20
hopefully. Um
- 3:22
>> yeah, I mean that that's incredible. I'm
- 3:24
actually really So when you guys
- 3:26
launched Omni for the first time, you
- 3:28
also did a podcast uh with Logan who
- 3:30
couldn't be here today uh and you added
- 3:32
like a sloth uh and and Ramen and all
- 3:34
these all these things. I actually
- 3:36
really want to do that to our videos. I
- 3:37
just didn't have an API for it because
- 3:39
obviously I have to automate the whole
- 3:40
thing. So thank you for the API.
- 3:41
>> Uh that is my favorite use case.
- 3:43
Everybody should do that. Um I got a cat
- 3:45
which is probably like the most boring
- 3:47
of the animals. Um if you don't know
- 3:48
what we're talking about, you should
- 3:49
look it up. It's very funny. Feurer. um
- 3:51
Furer who's um you know on on the team
- 3:54
did that.
- 3:54
>> Furer is the number one guy you should
- 3:56
follow for you should follow ideas on
- 3:58
okay what can this thing do?
- 4:00
>> Yes.
- 4:00
>> Right.
- 4:01
>> Yes. He he's he's amazing at that.
- 4:03
>> I've tried to get him for the last two
- 4:05
years to come to AIE. He hasn't made it
- 4:07
yet. He's actually come in person. He
- 4:09
just didn't want to speak because he's
- 4:10
anonymous.
- 4:11
>> I know.
- 4:11
>> I I want to say his real name but I
- 4:13
can't say his real name.
- 4:13
>> No [laughter] no we won't we won't do
- 4:15
that to him. But you should really
- 4:16
follow him. He's amazing.
- 4:17
>> He did all that work. I actually met him
- 4:19
uh in the office uh when we did the
- 4:22
podcast I think and I didn't realize it
- 4:24
was him. So his badge doesn't say
- 4:26
Popers.
- 4:27
>> Yeah,
- 4:27
>> I know.
- 4:28
>> So he used to be part of uh Replicate
- 4:30
and Replicate had this joke where like
- 4:32
everyone was Deep Fates. Deep Fates is
- 4:34
this like kind of mysterious character
- 4:35
and replicate. Replicate is very cool
- 4:37
company and both was part of it. Um, so,
- 4:40
okay, one thing I want to get on there
- 4:42
before I go into like sort of the the
- 4:43
the the sort of omniper is we added
- 4:47
cats, we added sloths, very cool, very
- 4:50
cute, very fun.
- 4:51
>> Uh, what are the, you know, inspire
- 4:53
people as to like what are the more sort
- 4:54
of workhorse use cases that maybe are
- 4:57
not just demos, you know?
- 4:59
>> Yeah. So, so obviously the hero
- 5:00
capability of the model or maybe there's
- 5:02
two like one is the ability to kind of
- 5:04
take in anything as input and then get
- 5:06
video on the other side. Obviously in
- 5:08
the future and and we've kind of talked
- 5:09
about this as a pre-announce like we
- 5:11
want to get the other output modalities
- 5:12
out as well but basically what that
- 5:14
means is you know you can take a set of
- 5:15
images that you have as maybe a
- 5:17
storyboard. You can take like an audio
- 5:19
track as a reference of you know like a
- 5:21
voice that you want a character to speak
- 5:23
and then you can get a video on the
- 5:24
other side. So like that just unlocks a
- 5:26
whole bunch of things that you can do in
- 5:27
like you know short film production or
- 5:29
you know shorts we've launched on
- 5:31
YouTube as well um to help creators kind
- 5:33
of like create um content more easily.
- 5:36
Um and then the other one is obviously
- 5:38
video editing. Like that's another thing
- 5:39
that we're really excited about that
- 5:40
we're just making easier because now you
- 5:42
can use natural language to take a
- 5:45
video, you know, add something, remove
- 5:46
something. Sloth is obviously like fun
- 5:48
example. Um, but there there's obviously
- 5:51
kind of there's consumer use cases that
- 5:53
we kind of had in mind where, you know,
- 5:54
you could take your beach vacation video
- 5:56
that was too noisy and you want to clean
- 5:58
up that noise. Maybe in the past you
- 6:00
wouldn't have because you didn't have
- 6:01
the tools or you didn't know what the
- 6:02
tools were that you needed to go to. So,
- 6:05
that's one use case that you can, you
- 6:06
know, go to. We've seen a lot of folks
- 6:08
use it for kind of marketing ad campaign
- 6:11
creation and I'm excited to see more of
- 6:13
those use cases as we launch the APIs.
- 6:16
um because obviously like we don't we
- 6:18
don't see all of it in the first party
- 6:19
products but I'm really excited for
- 6:21
people to start to explore that um in
- 6:23
the API. So those are just some of the
- 6:24
kind of like high level um things that
- 6:26
have come up. U people also use it to
- 6:29
create like education materials. Yes. Um
- 6:31
and like like that's really exciting. I
- 6:34
think we're all we've all kind of talked
- 6:35
about being excited about the future of
- 6:37
education where like everything can be
- 6:39
kind of customized to you and
- 6:40
personalized to your knowledge level and
- 6:43
the style that you prefer and and so
- 6:45
this is kind of just like a step in that
- 6:47
direction.
- 6:47
>> Yeah. I I I sort of actually used just
- 6:49
none of yesterday, but my my parents are
- 6:51
visiting and there was there was a very
- 6:52
fun sort of use case. They I bought some
- 6:55
gadget off from Amazon that they wanted
- 6:57
and the instructions to use it was were
- 6:59
only in English and there was plenty of
- 7:00
diagrams or whatever and I took a
- 7:02
picture of it and said, you know,
- 7:03
translate this into Romanian. Yes.
- 7:04
>> And keep everything else the same,
- 7:06
right? So it was amazing, right? Like it
- 7:08
was just like, yeah, it looks identical
- 7:10
and it has, you know, it's perfectly
- 7:12
translated. I mean, more or less, right?
- 7:14
But it's it's you know using Gemini
- 7:16
under the hood obviously to kind of do
- 7:17
the translation and so you can you can
- 7:19
see this use case for video as well
- 7:21
right like the the power of text
- 7:23
rendering in in in Omni is is quite next
- 7:26
level. So and you could you could you
- 7:28
could think about plenty of use cases of
- 7:29
like both text rendering translation
- 7:31
internalization all sorts of things that
- 7:33
would be actually genuinely useful to a
- 7:35
lot of different people and sort of
- 7:36
broader access to either you could like
- 7:39
redub a video or whatever it is that you
- 7:41
wanted to do. like there's plenty of
- 7:42
different things that you could you
- 7:43
could think about doing.
- 7:45
>> Yeah. Um one of the most enlightening
- 7:49
conversations I have on my podcast is
- 7:51
with uh just people researchers at the
- 7:53
frontier of these things. Um I had one
- 7:55
with um Ethan from the XAI video team,
- 7:57
the Grock video team who was basically
- 8:00
saying like you know the next trend is
- 8:02
actually not just like single model,
- 8:03
it's more like video agents. Um, and I
- 8:06
don't know if that terminology resonates
- 8:09
uh obviously for for very relevant for
- 8:11
RL. Uh, but it was it was basically kind
- 8:13
of like giving up on like trying to do
- 8:14
everything in in effectively one pass.
- 8:17
Um, do you feel that same way or is it
- 8:20
still an open research question which
- 8:22
way the trends are going? [snorts]
- 8:24
>> Yeah. So um what kind of excite me most
- 8:27
is really when the symbolic kind of
- 8:29
foundational models and this kind of
- 8:31
like video foundational model can
- 8:33
actually kind of really work together
- 8:34
and u in a way the if you look at the
- 8:37
beginning of the generative sort of like
- 8:38
image generation video generation a lot
- 8:40
of it kind of started when the language
- 8:42
model got good enough to provide a very
- 8:44
detailed captioning like from stable
- 8:46
fusion days or kind of dowi 2 days. So
- 8:49
um so basically like language is
- 8:52
extremely u helpful representation uh
- 8:55
one is that it's kind of universal but
- 8:56
the other kind of more um technical
- 8:59
thing like kind of my hypothesis is like
- 9:01
um one very difficult thing about
- 9:02
machine learning is um this sort of like
- 9:05
spirious coordination. So you don't know
- 9:07
you know if the if this kind of feature
- 9:10
right that's kind of predictive is
- 9:11
actually causal factor or not. There are
- 9:13
two ways. One is we can have really
- 9:15
diverse data training data like from
- 9:17
every intervention of the causal graph.
- 9:19
The other is you condition the causal
- 9:21
information and conditioning the
- 9:22
language is kind of like conditioning
- 9:25
like a coal information of the of the
- 9:27
kind of world. So um
- 9:29
>> which is a prompt or a concept what
- 9:32
>> yeah exactly so if you look at like you
- 9:34
know how we going to describe this video
- 9:35
how this kind of image is actually very
- 9:38
close to you know how would describe
- 9:39
this kind of causality you know behind
- 9:41
this like how this is kind of generated.
- 9:42
So one is like that can really allow for
- 9:45
very rich generalization and then uh
- 9:48
very kind of just like a good model. Um
- 9:51
the other is so eight months ago uh we
- 9:54
put the evaluation paper called video
- 9:56
models zero shot learners and reasoners.
- 9:58
>> Yes. So that was a kind of you know it's
- 10:01
it's a confirmed paper and then later on
- 10:03
actually the N banana team followed up
- 10:05
with a vision banana paper that
- 10:06
basically used n banana to do but
- 10:08
essentially the idea is uh video model
- 10:11
is extremely good sort of a foundation
- 10:13
model for space and time kind of
- 10:15
information. So um classic computer
- 10:17
vision tasks a lot of could be kind of
- 10:19
zero shorted and when you like say feed
- 10:22
in some like a visual quiz uh it can you
- 10:26
know there's definitely like a lot to
- 10:27
improve it can kind of solve and it can
- 10:30
um like robotics kind of like seeing it
- 10:32
has really good kind of physical
- 10:33
intuitions like word model uh and I
- 10:36
think the the key is really the kind of
- 10:39
mix of the visual kind of reasoning and
- 10:41
then the text kind of reasoning kind of
- 10:43
all tied together Um obviously you know
- 10:46
like whether doing it you know as kind
- 10:47
of unified model versus like just kind
- 10:49
of agent coation I think that's more
- 10:52
like uh it's going to be more kind of
- 10:54
incremental you know how it's going to I
- 10:56
imagine everything's going to go into
- 10:57
like a single model eventually
- 10:59
>> but right now there's like a lot you can
- 11:00
do if you uh basically take like really
- 11:03
good video understanding image
- 11:04
understanding Gemini agentically with
- 11:07
anomy and that's actually gonna yeah our
- 11:09
team is like exploring a lot
- 11:12
>> yeah okay that there's a there's a lot
- 11:13
in there um I I think uh one question I
- 11:17
I am increasingly starting to wonder is
- 11:18
does it all trend towards one product
- 11:20
for you guys right like now you have
- 11:22
multiple models out the naming of omni
- 11:25
does imply that eventually everything
- 11:28
will go away and it just goes into omnis
- 11:30
um is that the plan
- 11:33
>> is it [laughter] I don't know I I think
- 11:36
I think uh maybe I mean I think
- 11:40
eventually I I think there's sort of
- 11:42
different trade-offs engineering
- 11:44
research product trade-offs in like it's
- 11:47
like for the same reason like the the
- 11:50
sorry how is it called nano banana light
- 11:51
I don't know what the product name
- 11:52
>> nanob banana tite
- 11:53
>> nano banana too light yeah right it's
- 11:56
it's it's it serves a particular niche
- 11:58
right and it probably doesn't
- 12:00
necessarily fit immediately in the same
- 12:04
model literally checkpoint as uh
- 12:07
something that can do 4K you know uh 30
- 12:10
secondond videos right like they're
- 12:11
probably not like trainable in the same
- 12:14
quite way, right? Like, so I I don't
- 12:16
know. It depends on how how far into the
- 12:17
future you look like. Sure, in five
- 12:19
years from now, will they all be the
- 12:20
same model? Probably. Uh but like, you
- 12:23
know, six months from now, we'll we'll
- 12:25
probably still have, you know, multiple
- 12:26
different models doing different things
- 12:28
because kind of from pragmatically the
- 12:31
trade-offs are such that we we should
- 12:33
have multiple different kinds of models.
- 12:35
>> Yeah, I
- 12:36
>> I think that's right. And and just on
- 12:37
that note, I mean, we did call it Gemini
- 12:40
Omni because we wanted to hint at the
- 12:42
future where Gemini just becomes fully
- 12:44
multimodal in and out, right? And so so
- 12:46
it's definitely a move in that
- 12:48
direction. I think we'll probably see a
- 12:50
move in the direction where Omni also
- 12:51
generates images and edits images and
- 12:53
all those kinds of things. But Doo is
- 12:55
right that I think on the way there,
- 12:57
there's a bunch of really really useful
- 12:59
applications of some of these more
- 13:01
specialized models. And so we we will
- 13:03
probably continue to work on those as
- 13:04
well because like that serves a certain
- 13:07
need at this point in time that may not
- 13:09
exist you know a year from now. There's
- 13:10
also like a research question about like
- 13:12
just how much transfer there is between
- 13:14
different kinds of modalities, right? I
- 13:16
think you may believe that there's some
- 13:19
transfer between coding and video
- 13:20
generation and I think most people don't
- 13:23
necessarily believe that but they you
- 13:25
know you could try to think that there
- 13:26
is some some there something there or it
- 13:28
could be a waste right to put them
- 13:30
together to try to learn these both
- 13:31
tasks at the same time right so I think
- 13:32
it's it's it's interesting sort of
- 13:34
question to which extent like image and
- 13:36
video obviously kind of there's some
- 13:37
transfer like kind of not that different
- 13:40
there's value in in learning to output
- 13:42
video and audio at the same time because
- 13:44
joint audio visual is you know that's
- 13:46
how that's how it is. Um and then
- 13:48
there's you know other kind of
- 13:49
intersections of modalities that are not
- 13:51
super obvious right like 3D
- 13:52
representation coding I don't know maybe
- 13:55
uh things like that right so like I
- 13:57
think it's worth sort of exploring the
- 13:58
different corners there and we are
- 13:59
actively doing that um with a focus
- 14:02
towards like what people actually want
- 14:03
to do with these models
- 14:05
>> yeah um what one thing I feel I feel
- 14:07
like uh I'm surprised by but also I feel
- 14:11
like it's insufficiently answered is
- 14:13
what is the correct intermediate
- 14:16
representation Um, so captioning, right?
- 14:20
XI does captioning. Omni does
- 14:22
captioning. Um, and I I I understand how
- 14:26
captioning works for images. Um, and I
- 14:28
understand that you can extend it into
- 14:30
to video and and sort of guide it across
- 14:33
time. It just feels very inefficient. It
- 14:35
there's got to be I feel like there
- 14:37
should be something better. Uh maybe
- 14:38
it's code and maybe we generate you know
- 14:41
and obviously I think a lot of um ffmpeg
- 14:45
and mapplot um what's the three blue one
- 14:48
brown one manm um a lot of like video is
- 14:51
generated through code and maybe that's
- 14:53
like the optimal representation uh any
- 14:56
hypothesis as to like is is it better or
- 14:59
is just English all you need
- 15:01
>> well as so I'm in the Gemini and they
- 15:03
know we do like a lot of RL agent and of
- 15:05
course kind of coding so yeah We we're
- 15:08
definitely exploring the coding
- 15:09
representations.
- 15:10
>> Yeah.
- 15:10
>> As kind of better kind of way to
- 15:12
represent. Yeah.
- 15:13
>> But you know like do you what's your
- 15:15
probability estimate on like [laughter]
- 15:18
if we just output binaries like we just
- 15:20
you know like just it's just ones and
- 15:21
zeros.
- 15:23
>> Um I I guess maybe a kind of similar
- 15:27
discussion was like um basically is the
- 15:31
language the right representation like
- 15:33
right. So uh one kind of question for
- 15:35
example uh professor you know like some
- 15:37
ask is like you know why why does the
- 15:39
channel of thought need to be in the
- 15:41
natural language?
- 15:42
>> Yes.
- 15:42
>> Can it just be the kind of any kind of
- 15:44
like continuous tokens just any amount
- 15:46
of you know additional computations.
- 15:48
>> Um so one is like obviously the test
- 15:52
like adaptive compute is going to give
- 15:55
like you know better results. So it's
- 15:56
that but what really kind of made CH
- 15:58
thought so you know like four years ago
- 16:00
I wrote you know the larger model zero
- 16:02
sort reasoner and then self-improvement.
- 16:04
So I kind of know from the very early
- 16:05
day but the reason like it works really
- 16:08
well is um right now the recipe that
- 16:11
works is the pre-training that scales a
- 16:13
lot and then that basically like learns
- 16:15
a lot of intelligence. there are a lot
- 16:17
of you know scaling RL but those are
- 16:19
still like extremely kind of comput
- 16:21
incent intensive to extract the
- 16:23
information and um you really want to
- 16:26
rely the intelligence on that so
- 16:28
basically by tying the sort of like a
- 16:31
reasoning in the natural language you
- 16:32
basically directly use the intelligence
- 16:34
of the pre-training to it while if you
- 16:36
remove that kind of constraints then
- 16:38
you're not um and these days uh I feel
- 16:42
the a lot of advancements in the texts
- 16:45
but also doing this kind of multimodal
- 16:47
space is very driven by this uh kind of
- 16:50
text as a kind of great uh sort of
- 16:53
representation.
- 16:54
>> Yeah, it's a good backbone.
- 16:56
>> Yeah,
- 16:57
>> I I think to me it's even simpler than
- 16:58
that. It's text is is how we
- 17:00
communicate. So I think fundamentally if
- 17:02
you're building kind of products that
- 17:03
humans will be interfacing with um like
- 17:07
like that we will be using text somehow
- 17:09
if it's a text interface, right? Not not
- 17:11
for everything. So I think it's it's
- 17:13
natural to default to that.
- 17:15
>> Yeah. Obviously there's like a conf
- 17:17
discussion you know some arrow like
- 17:18
arrow maximalists is like oh we don't
- 17:20
care about you know kind of channel
- 17:22
those kind of like stuff it's just just
- 17:24
additional compute
- 17:24
>> sure but I personally yeah
- 17:26
>> ro maximalists I wonder I wonder who who
- 17:30
qualifies in that description David
- 17:31
silver
- 17:32
>> ah okay yeah I mean they they've just
- 17:35
left to to start their thing um
- 17:38
interesting okay so uh I I mean I think
- 17:41
I'm very interested in just like better
- 17:43
representations because I think that's
- 17:44
one of our themes that we're curating
- 17:46
today uh at the world fair is world
- 17:48
models. You mentioned the word world
- 17:50
models but it's not something that's
- 17:51
like super well defined. I think
- 17:52
everyone's like sort of converging on
- 17:54
some version of it that is like the
- 17:56
ideal.
- 17:57
>> Sure. Everything is a world model now.
- 17:59
It's sort of a
- 18:00
>> it's not it's not that useful, right?
- 18:02
>> So I just gave a keynote at the IER
- 18:05
world model workshop. Yeah. And then uh
- 18:07
yeah essentially uh I definitely
- 18:09
encourage to check out the definition by
- 18:10
Jandra Matic. He's like the you know OG
- 18:13
computer vision professor UC Berkeley.
- 18:15
>> Uh he has pretty you know bit of word to
- 18:17
say about world model
- 18:18
>> but also kind of Schmidt Herburver's
- 18:19
kind of how he defined the world model
- 18:21
from 2019 like 1990 sort of uh uh you
- 18:26
know like Wayne was just basically just
- 18:27
that kind of model base. Uh for me the
- 18:29
word model is basically just the model
- 18:30
in the model based RL and I feel that
- 18:32
has sufficient to describe but obviously
- 18:34
you know there are like a lot of uh FE
- 18:36
had a kind of nice blog post about what
- 18:39
about yeah this kind of broken down
- 18:41
>> um but yeah
- 18:43
>> yeah I mean so you know I I'll end this
- 18:47
part of the conversation but like I I do
- 18:49
think that language to me relying on
- 18:52
language as like the sort of like the
- 18:53
narrow pipe through which everything
- 18:54
goes through um still is like a lossy
- 18:57
compression. No, no, no. But we're not
- 18:59
seeing that, right? We're basically
- 19:00
saying the video model and the language
- 19:02
together.
- 19:02
>> So, so I think the language alone is uh
- 19:05
not sufficient. That's why we feel like
- 19:07
the video is a very complement model.
- 19:09
Right? Now the um you know kind of v
- 19:11
omni many people feel as uh you know
- 19:14
generating kind of pretty videos but I
- 19:16
think our vision it's it's much more
- 19:18
than that. It's a missing foundational
- 19:19
model that's absolutely required if you
- 19:21
want to make the AGI that match to
- 19:23
humans not just a jacked one.
- 19:25
>> Yeah. Um okay. So one one other thing
- 19:28
you know you you mentioned on the vision
- 19:30
side um and I'm kind of curious how sort
- 19:34
of uh parallel you know in terms of your
- 19:37
research careers um this development is
- 19:40
like I think basically a lot of vision
- 19:41
people have crossed over into more model
- 19:44
people um a lot of vision people also
- 19:46
become generative video and image people
- 19:50
and is it just as simple as you know
- 19:53
reversing uh image to text and then now
- 19:55
it's text to image like
- 19:57
>> [laughter]
- 19:58
>> is is that if I mean that effectively
- 20:00
was the diffusion process. Um
- 20:03
I I just you know I I just see the
- 20:06
career paths of the people that I talked
- 20:07
to and and see and I I I see this
- 20:10
overall trend of research directions and
- 20:12
I just wanted you to guys to sort of
- 20:14
reflect on on that.
- 20:16
>> I mean I certainly went that way right I
- 20:17
I started long time ago uh doing
- 20:20
computer vision sort of object detection
- 20:23
recognition things like that. Uh I think
- 20:24
just that's just simpler problem right
- 20:26
just generation is just harder like it's
- 20:28
a it's a different kind of mapping right
- 20:30
you map from the the inverse mapping is
- 20:32
not as simple as just inverting the the
- 20:34
kind of network you use right it's it's
- 20:36
a it's it's more ambiguous right to go
- 20:38
from cat to image of a cat and in some
- 20:41
ways it's also a loop because your
- 20:42
vision work creates the synthetic labels
- 20:44
that then continues
- 20:46
>> I mean sure [laughter]
- 20:48
I don't know I don't know I try to
- 20:49
validate my my sort of theories about
- 20:51
how fields develop how how careers has
- 20:53
progressed through this
- 20:55
>> I mean for like the the the better the
- 20:57
understanding side gets like we have
- 20:59
seen that the generation side also gets
- 21:02
better right so like like
- 21:03
>> it's completely bootstrapping yeah it's
- 21:05
>> and so like like like there's definitely
- 21:07
they're there to that thesis and I think
- 21:09
yeah I think a lot of people have kind
- 21:10
of like I I definitely worked with a lot
- 21:12
of um image understanding people who
- 21:14
became image generation people you know
- 21:16
and then some of them have moved on to
- 21:17
video because it's kind of like the next
- 21:19
thing where you have so many more
- 21:20
dimensions to work with so yeah I'm
- 21:22
curious about you spec as your
- 21:24
>> so I definitely like recommend start
- 21:26
with understanding recognition because
- 21:28
that's basically discriminator and then
- 21:29
that's going to lead to better
- 21:30
generation and that's what the bridge is
- 21:32
basically reinforcement learning so my
- 21:35
um my kind of journey is I initially
- 21:36
kind of worked on the algorithmic
- 21:38
research in the gent model against some
- 21:40
like you know amnest kind of generation
- 21:42
and then I worked on like RL and
- 21:44
robotics um and then like six years ago
- 21:47
I was like leading like a moonshot on
- 21:49
the dexterity it was pretty early but I
- 21:51
see now everyone's kind of doing
- 21:53
uh four years ago I basically kind of
- 21:54
figured out that this like symbolic AGI
- 21:58
is going to accelerate much faster than
- 22:00
the kind of physical AGI kind of
- 22:01
counterpart. So uh I decided to kind of
- 22:04
like language models and then those
- 22:06
things. Um and then recently kind of
- 22:08
work with Doomi and then like omni team
- 22:10
I quite enjoy kind of collaboration
- 22:12
there. the what I quite enjoy uh what I
- 22:15
recommend definitely to the researcher
- 22:17
is to uh definitely kind of explore or
- 22:20
at least like get exposure to what the
- 22:22
top people in each of the community are
- 22:24
like looking at how they kind of think
- 22:26
about problems. So when I look at the
- 22:28
video model to me it kind of reminds me
- 22:30
like pretty early on sort of like
- 22:32
language model where like very early
- 22:34
language model was a kind of creative
- 22:36
sort of demo right you kind of like try
- 22:39
to write like a story like mobile and
- 22:41
then like you know GBD2 and then those
- 22:43
kind of days like LTM kind of days right
- 22:45
and then you know uh instruction tuning
- 22:48
you actually kind of make it usable as a
- 22:50
chatbot but then at the chatbot stage it
- 22:52
still had so much hallucinations and
- 22:54
instruction for wasn't good enough so it
- 22:56
couldn't use for reasoning and when I
- 22:58
got good enough um in pre-training and
- 23:00
post- trainining for reasoning then you
- 23:02
know this kind of test time scaling the
- 23:04
RL really took off to like many of the
- 23:06
kind of best performing models and right
- 23:08
now I think the video model is as we
- 23:10
mentioned it's it is a complimentary
- 23:12
foundational model and I can imagine
- 23:13
it's going to follow a similar path it's
- 23:15
going to be very uh it's going to
- 23:17
improve a lot instruction following a
- 23:19
lot of uh this it's going to improve a
- 23:21
lot in reducing coordinations to extend
- 23:23
that it become a very reliable world
- 23:25
model so we can kind of like intermixed
- 23:27
video like space-time simulation was a
- 23:30
text simulation to solve like arbitrary
- 23:31
AGI problems. Also like I think the
- 23:34
difference still is between sort of text
- 23:36
models and like image video models is
- 23:38
that like we haven't quite unified
- 23:39
understanding and generation in in
- 23:41
multimedia I'd say yet like I mean I
- 23:44
think I think without going to the
- 23:45
details of course there's like it
- 23:46
depends on on at which level you're
- 23:48
thinking about this but generally like
- 23:51
there's not that many as far as I know
- 23:52
models sot kind of you know frontier
- 23:55
models that are genuinely
- 23:58
kind of good at both understanding and
- 24:01
generation of of let's videos, right?
- 24:04
Like it's a it's a it's an interesting
- 24:05
challenge. I'm not saying that we should
- 24:07
do this. Uh but but I think uh it kind
- 24:10
of stands to reason that like you know
- 24:12
understanding and generation are two
- 24:13
sides of the same coin. So they they
- 24:15
kind of should be in the same model in
- 24:16
some ways. Uh but we don't necessarily
- 24:18
always do that. So yeah.
- 24:20
>> Uh you mentioned audio as well, right?
- 24:22
Yeah. Uh is that as hard as video or
- 24:28
qualitatively different? If if so, in
- 24:30
what way? Uh, one of the interesting
- 24:33
directions three years ago was people
- 24:36
using um, I guess diffusion to do audio
- 24:41
uh, as in like the the sort of refusion
- 24:44
approach. I don't know if you you guys
- 24:45
saw that. Um, and I just think it's like
- 24:47
very interesting if a modality that we
- 24:51
perceive which is audio is different
- 24:52
than video actually two machines is
- 24:54
exactly the same like there's they see
- 24:57
no difference.
- 24:59
I mean I think on a technical level
- 25:01
there are some differences but I think
- 25:02
they're like relatively minor. I think
- 25:04
from my perspective audio came into into
- 25:07
my life when we shipped V3 which was I
- 25:10
believe the first model that did like a
- 25:12
joint
- 25:12
>> with the slicing of the
- 25:14
>> Yeah. Yeah. the gold bars or whatever.
- 25:16
Um it it was the first model that did
- 25:18
this sort of joint audiovisisual
- 25:19
generation. Yes. uh like in the in the I
- 25:22
mean there are there were other models
- 25:23
that did kind of you know kind of kind
- 25:25
of agentic hacking under the hood but
- 25:26
this one was truly sort of you know
- 25:29
generating everything at once and we the
- 25:32
reason we did that is because we felt
- 25:34
and I think was the right choice we felt
- 25:36
that like uh it only makes sense to
- 25:39
generate them at the same time because
- 25:40
there sort of kind of like from a
- 25:42
machine learning perspective there's one
- 25:43
latent kind of you know causal kind of
- 25:44
you know generative process right like
- 25:46
there's something that generates you
- 25:48
speaking it's not the pixels and then
- 25:50
the the audio or somehow somehow
- 25:52
generated by some other process like the
- 25:53
lips have to move in sync with with the
- 25:55
with the audio, right? So, I think that
- 25:57
that solved a lot of the issues that
- 25:59
previous models had or the way that
- 26:00
people did video generation before where
- 26:02
it was like, okay, we generate the
- 26:03
pixels and then we're going to hack
- 26:05
something on top of it that like moves
- 26:06
the lips with the audio that we
- 26:08
generate. And that's was very bad.
- 26:10
[laughter]
- 26:11
And so, I think I think that was that's
- 26:13
to me that's the the I mean after V3
- 26:16
like you know people were like what do
- 26:17
you mean like there's no audio in your
- 26:18
model? like that makes no sense like
- 26:20
once it's there like you you have to
- 26:21
have it. So I think that was that was
- 26:23
the right choice and doing it to one
- 26:25
single generative model I think was was
- 26:27
the right choice.
- 26:28
>> One thing I kind of want to also can ask
- 26:30
you guys an opinion as well once one
- 26:32
difference I find the audio and then
- 26:34
against the image and video is like the
- 26:35
audio information is less verbalized. I
- 26:38
mean of course the TTS and stuff is
- 26:40
trivial right but the when you get her
- 26:42
outside like how to describe music how
- 26:45
do you describe this like this person's
- 26:47
tone kind of pitch I feel the sort of
- 26:50
the verbalization is insufficient and
- 26:52
the interesting thing is that you kind
- 26:54
of see that in two other things like
- 26:56
taste taste sense and also uh say um
- 27:00
touch
- 27:00
>> like smell and then the another
- 27:02
interesting thing is the skin color so
- 27:05
skin color the the language is pretty
- 27:08
limited to describe the skin color and
- 27:10
the reason is that we're extremely uh
- 27:12
sensitive to the small difference
- 27:13
perturvations or not skin color because
- 27:15
that basically shows us is this person
- 27:17
going to kill me or is can I befriend
- 27:19
this person kind of those kind of
- 27:21
information and then I feel the smell
- 27:22
tastes um skin color and like sound kind
- 27:26
of stuff is very very tied into
- 27:28
primitive a like survival kind of stuff
- 27:32
and so our sort of sensory system is so
- 27:34
sensitive that it's intractable to Um,
- 27:38
so for example, I asked like one the
- 27:40
wine sort of taster and then like
- 27:42
professional and then he basically said
- 27:43
he kind of use like a language from like
- 27:45
a dating, you know, describing like a,
- 27:47
you know, partner as a way to describe
- 27:50
the taste because there's no sufficient
- 27:52
vocab to describe. Um, so I'm kind of
- 27:56
curious. Yeah. Do you guys feel that?
- 27:58
>> I think well to some extent I think the
- 28:00
same is true for visual information,
- 28:03
right? when you think about like a
- 28:06
certain style or a certain aesthetic,
- 28:08
right? Like like there are some people
- 28:10
who just have a much more kind of
- 28:11
developed like whether it's palette or
- 28:13
kind of visual taste and aesthetic,
- 28:15
right? Like I I think language just
- 28:18
tends to be a bit of a limiting factor
- 28:20
when you are trying to describe any of
- 28:22
these things that like we experience
- 28:24
with sensory information. And to your
- 28:26
point earlier, I think that is the kind
- 28:29
of the reason why we are investing in
- 28:31
world models and why we are pushing on
- 28:33
kind of the like perception and like
- 28:34
generation side of things because it it
- 28:37
is such a large part of how we as humans
- 28:40
navigate the world. It's a large part of
- 28:42
how like embodied AI navigates the
- 28:45
world. Um, and and I do I do think
- 28:48
language like does have a lot of it's
- 28:50
it's gotten us very far and it can
- 28:51
probably get us really far, but it it
- 28:53
feels limiting in a lot of these kind of
- 28:55
areas. And yeah, I don't I don't really
- 28:57
know how to describe, you know, like
- 28:58
sense and taste. Um, but yeah, I'm
- 29:01
curious to me.
- 29:03
>> Um, I I yeah, I don't know that I have
- 29:06
thought that deeply about this yet. So,
- 29:08
uh, yeah, I mean yeah, I don't have a
- 29:11
good answer about audio. I mean like I
- 29:14
don't know the limit because I'm
- 29:16
thinking about like well what is what is
- 29:17
Omni bad at in terms of audio but
- 29:19
they're all like solvable problems I
- 29:21
find uh so like with more data or better
- 29:24
data or whatever it is so I don't know
- 29:26
like that we have pushed the frontier so
- 29:28
much that like we are have hit some sort
- 29:31
of limits that are rooted in
- 29:33
evolutionary uh kind of you know limits
- 29:36
imposed by humans. I don't know. He's
- 29:38
feeling the limits of captioning which
- 29:40
is the the thing I was
- 29:41
>> Yeah, exactly. [laughter] There there's
- 29:42
a lot of information in the world and it
- 29:44
connects to basically why we do work
- 29:45
modeling you mentioned. You just need
- 29:47
srefs sref476
- 29:50
and then that's your that's what your
- 29:51
journey does, right? I guess maybe I
- 29:53
can't describe this vibe but
- 29:54
>> well well I think that that's kind of
- 29:56
the point of providing some of these
- 29:57
references, right? Because because like
- 29:59
even just describing how someone talks
- 30:02
and like their tone and and like procity
- 30:04
and all of these things like I think I
- 30:05
think some of these terms even like I
- 30:08
didn't used to know what they mean,
- 30:09
right? Well, now
- 30:10
>> yes. Dispuencuencies ex like like there
- 30:14
there's kind of an entire vocabulary
- 30:16
that even if you're not kind of steeped
- 30:17
in a domain, which is true for actually
- 30:19
like most human domains that like you
- 30:21
don't even know what it means. Um and
- 30:23
sometimes it's also a question of like
- 30:24
if we haven't focused on those things,
- 30:26
you know, with the large language models
- 30:27
that they may also have gaps in those
- 30:29
areas, right? And then we feel them on
- 30:30
the other side with generation because
- 30:32
we're like fundamentally relying on on
- 30:34
the language models understanding of the
- 30:36
world to then be able to like represent
- 30:38
it. Um, so I yeah, it all kind of goes
- 30:40
back to your question about like the the
- 30:42
language as an intermediary. Um, but
- 30:44
yeah, I think to De's point like some of
- 30:46
these might just be like focus areas and
- 30:48
things that we haven't necessarily
- 30:49
pushed on as much as we can and like as
- 30:51
we will we will discover what the actual
- 30:54
ceiling is.
- 30:55
>> Yeah, as a podcaster I think a lot about
- 30:58
sound.
- 30:59
>> Um, and and I I'll just offer a couple
- 31:02
things for discussion in case in case it
- 31:03
triggers anything with you guys. Um I
- 31:06
have three domains of rough audio which
- 31:07
is like a music voice SFX you know is
- 31:10
that rough okay covers everything and
- 31:13
then also even within voice let's just
- 31:15
let's just focus on voice forget the
- 31:16
other two um room sound like the the
- 31:18
echoiness of like big room small room in
- 31:21
person in a car over a phone all these
- 31:25
like are labelable but we experience
- 31:27
them very differently and I I often
- 31:29
think like one of the tells of a AI
- 31:31
video is that it is studio quality
- 31:33
because it was recorded in a studio
- 31:35
video because that's your training data
- 31:36
and like and and to me that's one thing
- 31:39
actually like the most interesting thing
- 31:40
is just uh when I tell this is how I
- 31:43
convince people who are kind of
- 31:44
skeptical about the need for world
- 31:46
models because you need it even for
- 31:48
audio about well I'm further away from
- 31:51
you so I should sound a little bit
- 31:52
softer or more diffused and like the the
- 31:55
video models need to pick that up
- 31:56
because if they're going to do immersive
- 31:58
video and audio you need that
- 32:01
>> I I I love that example of basically
- 32:03
like studio quality or not in a way like
- 32:05
we don't have enough language to really
- 32:07
describe like like this kind of echoing
- 32:10
or like some kind of noise kind of
- 32:12
happening we just like don't have
- 32:13
precise enough and uh if you um you know
- 32:16
basically the reason that I think it's
- 32:18
quite important to have like relatively
- 32:20
information rich like kind of captioning
- 32:22
is that we kind of rely on the natural
- 32:23
language as a representation but if you
- 32:25
basically don't have enough uh
- 32:27
representation that basically means the
- 32:29
condition on the language the generation
- 32:30
is very multimodal and if you anything
- 32:33
can learn from the BAE kind of like you
- 32:34
know very old you know GMBA kind of
- 32:36
research the idea is we really want to
- 32:38
capture most of the stoasticity in the
- 32:40
later representation and then the the X
- 32:43
given the Z should be kind of like
- 32:44
deterministic so yeah
- 32:46
>> yeah yeah um well I hope I hope there's
- 32:49
more uh progress there and I'm sure you
- 32:50
guys are doing
- 32:51
>> I even actually like facial expressions
- 32:53
right and maybe this gets to your point
- 32:55
about like things that we're very
- 32:56
sensitive to right I think you can tell
- 32:58
a lot of AI content also just by from
- 33:01
like people's facial expressions
- 33:02
stressful.
- 33:04
[laughter]
- 33:04
>> Yes.
- 33:06
>> Yes. And we try not to contribute to it,
- 33:08
but you know, um and or or like skin
- 33:11
textures, right? Like like the things
- 33:12
that kind of make things look real in
- 33:15
real life. Like I you know, I can tell
- 33:16
from the way you're nodding or from the
- 33:18
way like your micro expressions are kind
- 33:19
of changing of like how you're reacting
- 33:21
to what I'm saying. Like we haven't
- 33:23
quite crossed that chasm. I think like
- 33:26
we're we're so much better than we were
- 33:27
a year ago.
- 33:28
>> Yeah. Um, but there's so much more
- 33:30
headroom kind of in a lot of those
- 33:32
things that like we as humans are super
- 33:33
sensitive to. And like I think image
- 33:35
arguably probably is there because
- 33:38
there's there's a lot of kind of images
- 33:40
that I will see that like really do look
- 33:42
indistinguishable from reality and I
- 33:44
can't tell if they're generated or not.
- 33:46
>> Better than reality
- 33:47
>> um or well that's a different
- 33:49
>> No, I I think that one of the parad
- 33:51
[laughter] better than what I would take
- 33:52
on my vacation as a photo. Yes. One of
- 33:55
the one of the fun experiments that we
- 33:56
did a while ago in the team is is like
- 33:58
can we generate videos that are better
- 34:00
than than real videos, right? So you
- 34:01
just take the same caption from like oh
- 34:03
yeah some video and then
- 34:06
>> recycle it. Yeah. Just just try to like
- 34:08
describe a real video and then generate
- 34:10
the equivalent version with omni and
- 34:12
then do a human eval. How does how does
- 34:14
it do? And then humans largely prefer AI
- 34:17
generated
- 34:18
>> margin.
- 34:20
>> But because it's because it's the RL
- 34:21
process,
- 34:22
>> that's the RL process working.
- 34:23
>> It's however you want to rationalize it.
- 34:25
It's not necessarily the old process.
- 34:26
It's just like I think it's just I'm not
- 34:28
saying this is a good result. I'm just
- 34:30
saying is we have optimized in a way
- 34:33
that like kind of potentially sort of,
- 34:35
you know, triggers something in the
- 34:36
human brain that like, oh, it looks it
- 34:38
looks all a lot of the videos just look
- 34:41
better. Like I'm not Yeah. Yeah. Yeah.
- 34:43
on on inspection on on deeper inspection
- 34:45
they they would not actually be more
- 34:47
useful or whatever but like if you just
- 34:50
say side by side random YouTube video
- 34:52
versus
- 34:53
>> generated version of it will you will
- 34:56
just have a it will just look better
- 34:57
because it's more it's a sharper more
- 34:59
HDR uh you know the skin tone is is is
- 35:03
better it's not again it's not more
- 35:04
realistic
- 35:06
>> uh it doesn't solve your problem
- 35:07
necessarily but it it looks better
- 35:09
>> I I since also depend on the sensitivity
- 35:12
of the people. Uh I was born raised in
- 35:14
Japan and I think one thing I kind of
- 35:16
know is like they're extremely extremely
- 35:17
like sensitive about like you know
- 35:19
that's why you know like architecture
- 35:21
like food and stuff like they have.
- 35:23
>> Um so I talked to like a mangar like
- 35:25
like artist there and he's like he's
- 35:26
kind of disgusted by like the generation
- 35:28
AI and one kind of thing he mentioned is
- 35:31
like the eye gaze
- 35:32
>> eye gaze that slight difference
- 35:35
>> makes me makes him kind of feel creepy
- 35:37
about like unnatural
- 35:38
>> like if you're looking a little bit off.
- 35:40
>> Yeah. It's just uh Yeah. just like uh it
- 35:43
looks too fake. Yeah. To the point. So
- 35:45
So I think it does depend on the
- 35:46
sensitivity and
- 35:47
>> Yeah. Yeah. All I'm saying is like you
- 35:49
know human preferences are like a not
- 35:51
particularly like uh reliable barometer
- 35:54
of like what you should be optimizing
- 35:56
for like if you just ask people do you
- 35:57
like this or not you not necessarily get
- 35:59
what you wanted.
- 36:01
>> Yeah. Let let me just kind of add one
- 36:02
thing but like four years ago there was
- 36:04
a like debate that if the prompt
- 36:05
engineering is going to disappear and uh
- 36:07
my my like you know some very powerful
- 36:10
people say you know it's going to
- 36:11
disappear but I basically said like it
- 36:13
shouldn't because the prompt engineering
- 36:16
like sort of you know specifying that is
- 36:18
like the the only way you can sort of
- 36:20
control the output sort of you know when
- 36:22
you have like sort of control by the AI
- 36:25
and what allows you to prompt engineer
- 36:27
is really that sensitivity. So sure
- 36:30
maybe like right now the AI can do a lot
- 36:32
of autoprompting and that and it can
- 36:34
generate something that's sufficient but
- 36:37
uh if it's like that never be satisfied
- 36:39
like never be satisfied with the AI's
- 36:41
generated content always fine-tune your
- 36:43
sensitivity and always kind of keep
- 36:45
prompting the differences. I I think to
- 36:48
the there's also a big difference
- 36:49
between like the average human untrained
- 36:53
eye which I I would put myself in that
- 36:55
bucket you know like I have I have some
- 36:56
aesthetic sensibilities and I've done
- 36:58
this long enough that you know like I
- 37:00
have I have a preference um but you know
- 37:02
like your example of a manga artist like
- 37:05
that's somebody who has honed a craft
- 37:07
like over possibly many decades. Um, and
- 37:11
anybody who does that, whether it's like
- 37:12
design, architecture, right? Like you
- 37:14
you you just have a very different level
- 37:16
of like expertise and you see things
- 37:19
that like the average human will not
- 37:21
see. But Doom is right. Like when we
- 37:22
look at if you were to just, you know,
- 37:25
um, poll 10 people on the street, they
- 37:28
would probably prefer the like overly
- 37:31
smooth like very saturated kind of
- 37:33
>> It's called the Instagram filter.
- 37:35
>> It is. It is the Yeah. [laughter] Um,
- 37:38
and you know, and and so there's also a
- 37:39
little bit of a question of like what
- 37:41
does your default aesthetic look like if
- 37:43
you don't specify? But then to Shane's
- 37:45
point, one of the things we always try
- 37:47
to get these models better at is
- 37:49
instruction follow so that like when you
- 37:51
want to get them to a different outcome
- 37:53
like you should be able to whether
- 37:54
that's through language or whether
- 37:56
that's through references because
- 37:57
language is sometimes too limiting. Um,
- 38:00
and so like these models continue to get
- 38:02
better at it but they so much at work.
- 38:03
Do do you feel pressure as a as a
- 38:05
product director to set the default for
- 38:07
the world like I mean [laughter]
- 38:10
>> kind of
- 38:11
>> maybe I should I don't know I haven't
- 38:12
thought about this
- 38:14
>> you know you know it's like someone has
- 38:16
to have a default the default has to
- 38:18
exist
- 38:18
>> actually I will say like we have thought
- 38:20
about this um and I I think one of the
- 38:23
so for example actually like if you look
- 38:25
at nanobanana generations we had like an
- 38:27
explosion of nanobanana infographics
- 38:29
when nanobanana pro came out
- 38:31
>> I tried it yeah
- 38:32
>> um yeah
- 38:33
I think Nurb's papers were like all you
- 38:35
know so so many had like infographics
- 38:37
generated. Can you run your uh
- 38:39
watermarking on it and see how many
- 38:41
>> uh we pro we probably could we have we
- 38:44
haven't done that but I saw so like my
- 38:45
Twitter was maybe this is just also like
- 38:47
the bias of my algorithm but they were
- 38:49
everywhere um and it was actually very
- 38:51
painful because um I think our default
- 38:54
aesthetic was a little bit too it was
- 38:57
too cluttered like I think that the the
- 38:59
model is like a bit of an overeager
- 39:01
student that just like learned you know
- 39:03
it was like oh I know all these like I
- 39:05
know all this information about this
- 39:06
concept let me like shove into the same
- 39:09
image. Japanese infographics 5x that
- 39:12
[laughter]
- 39:13
>> or maybe it was you know um but it just
- 39:16
and and
- 39:17
>> wait so same prompt same content if it's
- 39:19
in Japanese it's
- 39:20
>> density density
- 39:21
>> oh wow
- 39:22
>> because that's the style in Japan
- 39:24
>> yeah some like very you know bureaucrat
- 39:27
and [laughter] there's a famous word for
- 39:29
it yeah
- 39:29
>> no but we do do go through this process
- 39:31
with Omni we did it together right like
- 39:32
where like we had like a bunch of like
- 39:34
we like at the very end okay like this
- 39:36
is we did some tuning and like okay what
- 39:38
kind of style do we prefer right like
- 39:40
you know
- 39:41
>> is it more muted more saturated
- 39:43
>> we had a lot of saturation
- 39:44
>> yeah there was there were I think Nicole
- 39:46
just has PTSD so has forgotten about it
- 39:48
but she was very much involved in this
- 39:50
of like okay which which kind of color
- 39:52
palette do we basically prefer right and
- 39:54
it's you know it's it's it's not
- 39:55
something that like you have to make a a
- 39:58
trade-off there like uh
- 40:00
>> and and and it's because it ends up
- 40:01
being us right like actually it is true
- 40:03
like it it ends up being the modeling
- 40:04
teams and you could ask the question
- 40:06
legitimately of like are we the best
- 40:07
people to do that or should we actually
- 40:10
work with someone who like has a really
- 40:12
creative point of view and is more of
- 40:14
like you know an art director and like
- 40:15
has like and we kind of go back and
- 40:17
forth on this um
- 40:19
>> we have the trusted testers I'm on
- 40:21
>> we do we have trusted testers who give
- 40:23
us a lot of feedback and we take that
- 40:24
serious
- 40:25
>> very well organized by the way to have
- 40:26
these like weekly calls and stuff like
- 40:27
it's it's amazing
- 40:29
>> um Logan's team does a lot of that so
- 40:31
kudo kuda kudos kudos to Logan um who
- 40:34
couldn't be here today um and we have a
- 40:36
lot of people actually internally at
- 40:38
Google like Fulfur who give us like a
- 40:40
ton of No, no, no. Truly like who give
- 40:42
us a ton of feedback on like when we
- 40:44
when we release new checkpoints and like
- 40:46
sometimes it will be stuff that we like
- 40:47
don't see right like we would be like oh
- 40:50
yeah this optimization seems okay and
- 40:52
then they would come back what have you
- 40:53
done like you completely ruined my grass
- 40:55
you know because now the detail is all
- 40:57
blurry.
- 40:57
>> I think he just noticed not not a super
- 40:59
secret at this point but like that our
- 41:01
model tends to put rings wedding rings
- 41:03
on on on hand. That's yeah
- 41:04
>> very strange. I had never noticed that
- 41:06
but he's like he I just saw it and
- 41:07
there's a faux fur channel basically.
- 41:09
>> Uh he posted I was like why is there
- 41:12
wedding ring in every hand? I'm like
- 41:13
that's strange.
- 41:14
>> That sounds very common reward hacking.
- 41:15
>> Yeah. Yeah. Yeah. So but you know
- 41:17
something that we would not have we
- 41:18
would not have noticed necessarily while
- 41:20
while developing this right is an oral
- 41:21
artifact or
- 41:22
>> I I don't know you do have like a lot of
- 41:25
preference based and then you know you
- 41:26
may can prefer that sperious correlation
- 41:29
reward hacking it can happen like in
- 41:31
many weird ways. Yeah
- 41:33
>> it does. It is
- 41:34
>> uh this was related to another topic
- 41:36
that again I I try to use these
- 41:38
mainstage things as introductions or
- 41:40
ties in. Uh we have the eval track we
- 41:43
have character AI and YouTube talking
- 41:45
about how they evaluate videos. Um how
- 41:48
do you evaluate videos
- 41:51
>> apart from furer [laughter]
- 41:53
>> not everyone has a fauxur but also you
- 41:55
know I think there needs to be something
- 41:56
more quantitative
- 41:57
>> well I mean it's you improve Gemini
- 42:00
to improve the evaluation for video. Um
- 42:03
that's that's no no that's that's
- 42:05
definitely one way uh it's actually very
- 42:07
hard.
- 42:08
>> It's very hard. It's very hard um to get
- 42:10
like you know audators to evaluate
- 42:13
things in a video like including
- 42:15
especially things like aesthetics right
- 42:17
like that it's like there are some
- 42:18
things that are a little bit more
- 42:20
objective like especially when we talk
- 42:22
like let's say we talk about images and
- 42:23
we look at like infographics text
- 42:24
rendering that's actually fine right
- 42:26
because like you can kind of OCR things
- 42:28
out and then you can look at like okay
- 42:31
this letter is like messed up and then
- 42:32
the whole thing is actually useless
- 42:33
because if like literally if a letter is
- 42:35
off in render text you just can't use
- 42:37
that asset. Right. So th those things
- 42:40
are like a little bit more auto ratable.
- 42:42
Um from what we found we do rely a lot
- 42:44
on humans looking at things and so we do
- 42:47
do a lot of human evals. We do a lot of
- 42:49
human evals.
- 42:50
>> Do a lot of human ev and every time Jane
- 42:53
is like um and every time we have a new
- 42:56
model we like want to do more things and
- 42:58
we want to like gem in more capabilities
- 43:00
and then we have like more emails that
- 43:02
we have to run. Um, and then at some
- 43:04
point you do get two models that are
- 43:07
like kind of close to each other and
- 43:10
then like we literally make decisions
- 43:11
based on like looking at outputs side by
- 43:13
side. Sometimes like in a room like I've
- 43:17
been in rooms where there's like 10 of
- 43:19
us and we're just like looking at video
- 43:21
side by side and we're like do you
- 43:22
prefer this or do you prefer that? like
- 43:24
oh wow it's
- 43:25
>> I mean but it is it is genuinely very
- 43:27
complicated the more capabilities you
- 43:28
add like you know even just the one
- 43:30
capability but it's like almost AGI
- 43:32
complete capabilities like video editing
- 43:33
right like think about video editing as
- 43:35
a and like editing with audio and
- 43:38
>> my editor will be very happy to hear
- 43:39
this
- 43:40
>> edit the hardest problem in g media
- 43:43
>> I mean I don't know if it's the hardest
- 43:44
but it's definitely there right like uh
- 43:47
in terms of like complexity of of
- 43:50
evaluation like free form video editing
- 43:53
is you can do anything like
- 43:55
>> yes uh and like I I spent a lot of money
- 43:57
on that and it's very hard to tell me
- 43:59
>> like adding those we don't have like add
- 44:01
a sloth eval right like uh that we
- 44:04
>> well now we should
- 44:05
>> now we should yeah yeah yeah but like
- 44:06
things like that like it's it's it's not
- 44:09
that easy to track
- 44:10
>> I think I'm just surprised at the sample
- 44:12
size that you have right like to to test
- 44:14
the entire surface of your models you
- 44:17
still rely on a magnitude of hundreds
- 44:20
>> no no no so we do like yeah well we do
- 44:22
we do a ton of human evals on like on
- 44:24
like you know thousands of things. Um I
- 44:27
I think there's also like an element of
- 44:29
you know we can talk about things like
- 44:30
live experiments right like which which
- 44:32
is also where you get signal on like
- 44:34
like some of these more minute
- 44:36
differences at like much larger scale
- 44:38
then there's auto raers which is
- 44:39
definitely kind of a more it's a very
- 44:42
well defined space I think for LLMs much
- 44:45
more nent for media models and then like
- 44:49
sometimes you still do rely on human
- 44:51
judgment and we do rely on things like
- 44:53
feedback from people who just like have
- 44:55
a very owned like aesthetic and and
- 44:59
people who just like use these models in
- 45:00
their workflows dayto-day, right?
- 45:02
Because we could also like you could
- 45:04
have a model that does really well on
- 45:05
some slice of human evals, but then it
- 45:07
like really breaks a workflow for
- 45:09
somebody. And so this is why we do like
- 45:10
early access programs and we try to get
- 45:12
feedback and then we like try to
- 45:13
incorporate it before we release
- 45:14
something more broadly. I feel like
- 45:16
Shane had a hot take based on his
- 45:19
>> expression always when we were talking
- 45:20
about this
- 45:21
>> every kind of human sort of you know
- 45:23
work should be gradually kind of
- 45:25
amortized and then the interesting thing
- 45:26
is the video understanding especially
- 45:29
like against like AI gener like
- 45:31
detecting air stuff is extremely
- 45:33
interesting uh visual task
- 45:35
>> and then like some of it kind of
- 45:37
aesthetics or this kind of visual
- 45:38
quality but for some of the kind of
- 45:40
cases like semantically doesn't make
- 45:42
sense for example you're taking like
- 45:44
some like a famous scene from a movie
- 45:46
and try to sort of um construct that and
- 45:49
then if you kind of generate it uh it
- 45:51
can generate something there but at some
- 45:53
point some of the semantic information
- 45:55
doesn't make sense like it's actually
- 45:57
inconsistent. So can the AI actually
- 46:00
detect that? So when I evaluate the AI
- 46:03
videos like oh I feel I'm so smart you
- 46:05
know like like AI is still kind of
- 46:08
behind but we should make like a lot of
- 46:10
effort. I think the video understanding
- 46:11
is extremely uh important intelligence
- 46:14
task uh beyond just the pure aesthetics
- 46:16
or the preference. Um and yeah we we
- 46:20
should always try to amatize the human
- 46:22
>> human label. Yeah.
- 46:24
>> Yeah. Um, what data do you need? A lot
- 46:29
of people I talked to wanted to get in
- 46:32
front of you actually. Uh, they I mean
- 46:34
they want to be nice about it. They have
- 46:36
a lot of video data. They have gaming
- 46:38
data. They have real world video data.
- 46:39
They have images. They have labelers.
- 46:42
What do you want?
- 46:44
>> Are you like offering?
- 46:46
>> I'm just like this is your request for
- 46:48
like Okay. Okay. We get I'm sure you get
- 46:50
a lot of pitches, right? You get a lot
- 46:52
of people want to talk to you. what's
- 46:54
like I think actually it's the signal is
- 46:57
this problem this sorting out signal
- 46:59
from noise is the main problem so
- 47:01
creating a nice API of like okay if you
- 47:04
actually do a b and c we are interested
- 47:06
in that
- 47:09
>> um
- 47:11
loaded question there so uh I don't know
- 47:13
that there's like an easy like you know
- 47:14
did you do I think we we do already have
- 47:16
a lot of data I think it's it's
- 47:19
>> hard to talk about this
- 47:21
>> you know you want to talk about the
- 47:22
public I don't want to get you in
- 47:23
trouble Yeah,
- 47:23
>> but like I think
- 47:24
>> No, no. What I just want to say is like
- 47:26
hard to talk about this in a sort of you
- 47:28
know without trying to without I have to
- 47:31
think about the what I am revealing
- 47:33
about our project and what where we're
- 47:35
going. Um generally high quality data I
- 47:38
think maybe maybe let's just put it this
- 47:39
way right it's not not the secret
- 47:40
>> embodied I'm sorry
- 47:42
>> embodied data
- 47:43
>> I mean
- 47:44
>> yeah sure I mean we have sort of
- 47:47
announced I think publicly right that we
- 47:49
we have some sort of robotics
- 47:50
collaboration right like so I think it's
- 47:51
like a like or or but you because we
- 47:55
have a robotics team at GDM so you know
- 47:57
they're always interested in things like
- 47:58
that um I mean for Omni specifically I
- 48:01
think we're just quite interested just
- 48:03
high quality data right like you know it
- 48:04
it's not some sort of not necessarily
- 48:07
like oh random YouTube video but like
- 48:09
you know some a some more professional
- 48:11
shop things like that right the things
- 48:12
that those are those are things that
- 48:15
we're always on the lookout for like uh
- 48:17
and yeah
- 48:18
>> and I think for you know maybe this is
- 48:20
easier to some extent to answer for like
- 48:23
some of the agentic work as well like
- 48:25
like like actual kind of like what are
- 48:28
the tests that people are trying to do
- 48:30
right these things are actually kind of
- 48:32
difficult to manufacture if you're doing
- 48:35
it yourself or if you're like doing it
- 48:36
with a vendor, like what is the actual
- 48:38
like if you're creating a marketing
- 48:39
campaign, like what does that look like,
- 48:41
right? Like do do you start from here's
- 48:44
like a picture of my new product and
- 48:46
then I want to turn that into a video ad
- 48:48
and I want to turn that into a bunch of
- 48:50
assets that like fit fit all these
- 48:51
different ad formats that I need to push
- 48:53
onto various platforms to promote and
- 48:55
then like so you kind of go from this to
- 48:58
that and like what is that kind of
- 48:59
trajectory of tasks that you're that
- 49:01
you're like you know experiencing along
- 49:03
the way like that is really useful and
- 49:06
that is actually kind of difficult to
- 49:08
get right u because like we don't always
- 49:11
have the right firstparty surface where
- 49:14
people are actually doing some of these
- 49:15
things or like you might work with
- 49:18
someone who's a vendor but they don't
- 49:19
also don't have that product surface
- 49:21
right like like a lot of this kind of
- 49:23
information lives in the places where
- 49:24
people are doing these tasks and so
- 49:26
that's kind of difficult to get like if
- 49:27
anyone's figured that out you should
- 49:29
reach out to us
- 49:31
>> every channel of thought yeah every
- 49:33
[laughter] thought
- 49:33
>> every thought yeah and maybe the data
- 49:36
the Chinese lab is using
- 49:38
>> yes yeah uh you know
- 49:41
as a media person myself, right? Like
- 49:43
there's so many podcasters and people in
- 49:46
in marketing departments and all these
- 49:47
like they would happy to be your data
- 49:49
like you know just like put a BCI on my
- 49:51
head
- 49:52
>> and podcast [laughter] watch my things
- 49:54
uh because like you know there's just
- 49:56
endless amount of work to do like
- 49:58
there's so much work and this is all
- 50:00
like this needs to somewhat be commodity
- 50:02
like obviously you can be an art like an
- 50:05
artisan like you can be Hollywood for
- 50:07
like the really high quality stuff but
- 50:08
actually a lot of work is commodity and
- 50:10
like should be modelable and we want you
- 50:12
to do
- 50:13
>> [laughter]
- 50:13
>> And but we we want the high quality to
- 50:16
Demi's point right like we do want we
- 50:17
want the high quality.
- 50:18
>> We want commodity. Yeah. Yes. Yes. You
- 50:20
want on both sides.
- 50:21
>> Um I I just
- 50:24
>> Thank you for the solicitation.
- 50:26
[laughter]
- 50:27
>> Uh I you know we we we also I also added
- 50:30
a data quality track. I I think that uh
- 50:32
people want to understand like what uh
- 50:35
at AI like how to raise the bar, right?
- 50:38
like like the and a lot of it is just
- 50:40
educating the market and educating
- 50:42
researchers and engineers and founders
- 50:44
on like this is where we're going a lot
- 50:47
of this is stop doing that do this do
- 50:49
this instead and I'm like people will
- 50:51
listen
- 50:53
yeah I don't know uh to that extent you
- 50:55
know
- 50:56
>> but I think to that to that point like
- 50:57
there's a lot of again just like craft
- 50:59
that goes into this right and there's a
- 51:00
lot of process like you even to the
- 51:02
marketing campaign example you don't
- 51:04
create that in like five minutes right
- 51:05
you like go you go through a process and
- 51:07
you iterate and you like pick something
- 51:10
over something else because you liked it
- 51:12
for whatever reason like maybe the eye
- 51:13
gaze was correct right like we just we
- 51:15
don't know these things right because
- 51:17
none of us are marketing directors and
- 51:19
like the models don't know these things
- 51:21
>> I even kind of say this for the natural
- 51:22
like a language as well like I I always
- 51:24
kind of say 99% of information is inside
- 51:27
people you can only extract it through
- 51:29
active dialogue and befriending them so
- 51:31
most of the stuff on the internet is
- 51:33
like sort of the outcome the output of
- 51:35
that yes but you know what are what are
- 51:37
all the trajectories you know how did
- 51:38
this person have this inspiration to
- 51:40
write this paper
- 51:41
>> what is the starting point what is the
- 51:42
inspiration what are the dialogue that
- 51:44
sparked it those kind of stuff is kind
- 51:45
of inside people so even you know those
- 51:48
kind of like even the language space is
- 51:49
kind of that I think the creative is
- 51:51
kind of similar as well there's a lot of
- 51:52
dark knowledge
- 51:53
>> yeah it's like when you write a novel
- 51:54
right like a novel speaks to you because
- 51:57
like usually there's some sort of like a
- 51:59
personal connection that you feel to
- 52:00
like the story or the trajectory or the
- 52:02
characters right like if you read most
- 52:04
of the stuff that's written by LLM's
- 52:06
today like it's, you know, it's it's it
- 52:09
starts it falls into these like default
- 52:11
par patterns and like the language
- 52:13
starts to feel really similar and all
- 52:14
the descriptions sound really similar.
- 52:16
You can kind of like quickly read it as
- 52:18
like, oh, this is not that interesting
- 52:19
because like I can't connect to it,
- 52:21
right? Um, and again, that's that's kind
- 52:23
of like a human expertise.
- 52:26
>> One nice thing recently is the Google
- 52:28
Cloud and the Google Deep Mind are kind
- 52:29
of starting to invest a lot more in the
- 52:31
FTEEs for the product engineers. And I
- 52:33
also kind of saw some uh recruiting for
- 52:35
the creative you know gem media kind of
- 52:37
space as well. So I think those are kind
- 52:39
of really the effort because we we kind
- 52:41
of feel that you know what we can kind
- 52:42
of do with a lot of public data there's
- 52:44
limits but really you know partnering
- 52:46
with that we can provide kind of better
- 52:47
models and products and yeah we kind of
- 52:49
feedback
- 52:50
>> uh we have an FD track here for the
- 52:52
first time every lab is announcing it.
- 52:54
It's it's crazy. Um, one thing I'm
- 52:56
actually very keen on doing and I push I
- 52:59
push for this at Cognition as well is to
- 53:01
turn the FDES not just into sales and
- 53:04
solutions but also to EVAL's uh eval
- 53:07
workers.
- 53:08
>> FD is not the sales FD is way way bigger
- 53:11
than that. How do you frame FDs then?
- 53:14
Because [laughter] I do think about it
- 53:16
as sales like you're you know the more
- 53:18
the more you customize the solution for
- 53:20
>> so I define post training as anything
- 53:22
between the pre-training and the final
- 53:25
user experience anything anything is a
- 53:27
post training
- 53:28
>> and to me when I first sort of you know
- 53:30
learned a lot about I mean FD kind of I
- 53:32
guess originally you know came from like
- 53:33
path here and then that so I guess the
- 53:36
kind of history is different but yeah I
- 53:38
think the key is really that um you know
- 53:40
the key is like not only to kind of work
- 53:42
uh with them and ensure that they kind
- 53:43
of know how to
- 53:44
but also to sort of code like derive
- 53:48
kind of insights that can basically kind
- 53:49
of help both parties. They can put the
- 53:51
like a lot of harness how they use the
- 53:53
model. We can improve like very
- 53:54
upstream. So how to get the customer
- 53:56
feedback to the modeling I feel is the
- 53:59
kind of more the the role I I kind of
- 54:01
want for the fds. Yeah.
- 54:03
>> Yeah. Yeah. and and even for sorry just
- 54:05
on that like if you want to talk to us
- 54:06
or at least me um I I'm not going to
- 54:10
offer up your time um but I it's really
- 54:13
helpful for us to actually talk to
- 54:14
people who are using our models and like
- 54:16
understand where they're struggling uh
- 54:18
because again that just like it's it's
- 54:20
the real world task that you're actually
- 54:22
trying to use them for right like I will
- 54:23
talk to people who do kind of interior
- 54:26
inter interior design with some of our
- 54:28
image models um you know and they will
- 54:31
say hey like I really want to take this
- 54:33
pattern pattern, but then I want to
- 54:34
scale it across like 10 different ruck
- 54:36
sizes and sometimes I have like a very
- 54:38
custom ruck size and then the model
- 54:40
fails at like replicating the pattern
- 54:42
the same way or you know I want to do a
- 54:44
try on for these earrings and then the
- 54:47
earrings have a certain size and then
- 54:48
like my head has a certain size like it
- 54:50
has to make sense if you're actually
- 54:52
trying to try things on and like the
- 54:54
models kind of fail at a bunch of these
- 54:56
things that like actually happen in the
- 54:58
real world, right? Um and so that that's
- 55:00
like useful for us because for some of
- 55:02
these things like we don't think about
- 55:04
because we don't you know we don't use
- 55:05
the models for those tasks
- 55:06
>> or like um you know I think to your
- 55:08
point about ad campaigns or whatever
- 55:10
like people have like notions of brand
- 55:11
languages or whatever like which is
- 55:13
>> yes
- 55:13
>> like a a bunch of images or PDFs saying
- 55:16
things you know it's a pretty kind of
- 55:18
you know ambiguous question as well what
- 55:20
is the IKEA brand language you know is
- 55:22
it is it blue and yellow I mean that's
- 55:25
that's not a very like
- 55:26
>> but like what shade of blue you know.
- 55:27
>> Yeah. Yeah. Yeah. So there there's like,
- 55:28
you know, and the brands are pretty
- 55:29
spec, you know, pretty, you know, like
- 55:30
they they do care about the shade of
- 55:32
blue. It's not shouldn't just be a
- 55:33
random blue and a random yellow. That's
- 55:35
not going to be IKEA, right? I'm just
- 55:36
thinking about an example. But like this
- 55:38
is the kind of stuff that, you know,
- 55:39
it's not necessarily part of our like,
- 55:41
you know, developing frontier models
- 55:42
kind of, you know, necessarily mandate,
- 55:44
but it's something that we do want to we
- 55:45
do want to fundamentally like build
- 55:47
products that people will use to solve
- 55:50
concrete tasks, not just not just
- 55:52
research artifacts, right? So I think
- 55:53
it's useful to understand what people do
- 55:55
care about. Uh well, I'm sure a lot of
- 55:58
people are very grateful for your work
- 56:00
and there's a lot more to do that you've
- 56:01
made so much progress over the last like
- 56:04
even just couple years of like Nano
- 56:06
Banana and Theo and Omni and uh I don't
- 56:09
know what else you got cooking but we're
- 56:10
very excited like you this is one of
- 56:12
those things where like I was very
- 56:14
disappointed you know when Sora shut
- 56:16
down and and I think like there needs to
- 56:18
be more general exploration of uh you
- 56:21
know generative models and not just you
- 56:24
know coding. [laughter] I think I think
- 56:25
that is
- 56:26
>> we obviously like this.
- 56:27
>> We love coding. Love coding and and uh
- 56:30
yes uh but thank you so much for your
- 56:32
time. Uh it's been a real pleasure and I
- 56:33
can't wait to see what this looks like
- 56:34
next.
- 56:35
>> Thank you for having us. Great question.
- 56:36
>> Thank you everyone. [applause]