← All AI Engineer talks

AI Engineer World's Fair 2026

SOTA Generative Media Panel — Dumitru Erhan, Shane Gu & Nicole Brichtova, Google DeepMind

Read the talk

Generative Media Beyond the Demo: Representation, Realism and Real Workflows

Dumitru Erhan, Shane Gu and Nicole Brichtova discuss how multimodal models generate and edit media, why human preference can mislead evaluation, and what working creators reveal that finished training assets cannot.

From a talk by Dumitru Erhan, Shane Gu, Nicole Brichtova and swyx

At a glance

Ideas worth remembering

  • Text supplies useful structure and access to pretrained knowledge, while media references and video models carry sensory information that words may underspecify. The panel argues for their combination and leaves the ultimate representational limits unresolved.

  • Joint audiovisual generation has a clear causal motivation: visible motion and sound arise from the same event. Convincing output must also preserve scene-dependent properties such as distance and room acoustics.

  • A preference win does not establish realism or task success. Attractive sharpness, saturation and skin tones can mask errors, while expert inspection can reveal subtle failures or recurring artifacts such as unwanted wedding rings.

  • Video evaluation needs several kinds of evidence: automated checks for tractable errors, human evaluations covering thousands of items, live experiments and feedback from real workflows. Small-group comparisons are one part of that system.

  • Creative task trajectories reveal requirements absent from finished media. Forward deployed engineers and direct user conversations can turn decisions about pattern consistency, physical scale and exact brand colors into feedback for model development.

Faster generation changes the iteration loop

The panel opens as a chance to examine generative media beyond a mainstage launch. After introducing researchers and product work spanning video, Gemini reinforcement learning and image generation, swyx asks what developers can actually try. Nicole Brichtova describes two launches: Nano Banana 2 Lite and the Gemini Omni Flash APIs.

Brichtova presents Nano Banana 2 Lite as the fastest and cheapest image model in its family, with generation and editing quality above the original Nano Banana and approaching the larger models. Her practical emphasis is roughly 3-second latency: shorter waits make it easier to explore an idea, inspect the result and revise it. She says some outputs can also serve as production assets, so the fast model is useful beyond preliminary sketches.

The Omni Flash API launch makes video generation and editing accessible to developer workflows. swyx illustrates why that access matters with an earlier edited podcast featuring added animals and other objects: he wants similar transformations for his own videos, but needs an API to automate the work. The playful example establishes a concrete distinction between an impressive demonstration and a capability that can be incorporated into a repeatable process.

0:170:21
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:01 · section reference included

Storyboards, editing and translation

Asked for workhorse applications, Brichtova identifies two central capabilities. First, a model can accept different kinds of references and produce video: a set of images supplies a storyboard, while an audio track supplies a voice for a character. These inputs let creators specify aspects of a scene through examples. She connects this to short film production and creator workflows, while describing additional output modalities as a future direction.

Second, natural-language editing lets someone request additions, removals or cleanup in an existing video. A noisy beach vacation recording is her everyday example: the user can ask for the noise to be removed without first identifying a specialist tool. Marketing campaigns are another observed application. Educational materials extend the idea toward content adapted to a learner’s knowledge and preferred style, although that broader personalization is presented as a direction rather than a completed system.

Dumitru Erhan offers a specific image-editing example. His visiting parents needed instructions for a gadget, but the illustrated instructions were in English. He photographed the page and asked for Romanian text while keeping everything else the same. He reports that the layout remained effectively identical and qualifies the translation as more or less correct. The useful mechanism combines language understanding with text rendering inside the existing visual structure. He sees related possibilities in video localization, translated on-screen text and redubbing.

4:514:54
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

4:40 · section reference included

Combining symbolic reasoning with video models

swyx introduces video agents as an alternative to asking a single generation pass to accomplish everything. Shane Gu responds by focusing on cooperation between symbolic foundation models and video foundation models. Detailed language descriptions provide a useful shared representation, but his stronger hypothesis concerns spurious correlations: a predictive feature need not be a cause. Diverse training examples across interventions can help distinguish the two; conditioning on a description of what is happening may also supply information about the factors that produced a scene.

Gu then describes evaluation work on video models as zero-shot learners and reasoners. His argument is that learning to generate spatial and temporal information can support more than media production: the resulting model can attempt classical vision tasks, visual quizzes and tasks requiring physical intuition without task-specific training. He explicitly leaves substantial room for improvement. The important possibility is to combine this visual reasoning with text reasoning, rather than treating generated video solely as a final artifact.

Whether that combination lives inside one model or in an agent coordinating models remains an incremental engineering question. Gu imagines eventual consolidation, but says there is already considerable scope in connecting Gemini’s image and video understanding to Omni through an agentic workflow. That is a research direction under exploration, not evidence that a unified system has already replaced the separate components.

8:008:03
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

7:45 · section reference included

Why specialized models still have a role

The suggestion of one eventual model prompts a more pragmatic discussion of product boundaries. Erhan points out that a fast image model serves a different niche from a model designed for something like 4K, 30-second video. Their training and serving requirements need not fit the same checkpoint. His contrast between possible convergence in five years and continued specialization in six months is a forecast, not a product commitment: engineering, research and product tradeoffs still justify multiple models.

Brichtova says the Gemini Omni name signals an ambition for fully multimodal inputs and outputs, including eventual image generation and editing. But architectural consolidation also depends on transfer: does learning one task improve another enough to justify training them together? The panel sees a clear relationship between images and video, and a strong reason to generate audio with video. Transfer between coding and video generation, or coding and 3D representations, is less obvious. Combining tasks could help, or could consume resources without a corresponding benefit.

11:1811:28
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

11:12 · section reference included

Language is useful structure, but not the whole world

swyx asks whether captions are the right intermediate representation for video. Describing a scene over time in English can feel inefficient, especially when some video can already be generated through code. Gu confirms that coding representations are being explored, then connects the question to a broader one: why should a model’s intermediate reasoning use natural language instead of continuous tokens or another form of additional computation?

His answer rests on the knowledge acquired during pretraining. In the recipe he describes, large-scale pretraining supplies much of the model’s intelligence, while extracting capabilities through reinforcement learning can be computationally expensive. Keeping reasoning in natural language lets intermediate steps draw directly on that pretrained knowledge. An unconstrained representation might support computation, but it does not automatically retain this convenient connection to what the model has already learned. The panel also gives a product reason for text: it is a familiar way for people to communicate their intentions.

The discussion then narrows the increasingly broad term world model. Gu uses it to mean the model in model-based reinforcement learning, while acknowledging that the term has acquired competing meanings. When swyx objects that language is a narrow, lossy channel, Gu clarifies that language alone is not the proposed solution. Video and language should work together, with video supplying a complementary foundation for aspects of the world that text does not adequately represent. His ambition extends from attractive clips toward more general intelligence, but remains a research vision.

14:1314:16
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

14:05 · section reference included

Understanding and generation can improve each other

The migration of researchers from recognition into generation is more than reversing an image-to-text mapping. Erhan uses a cat as the example: recognizing a cat in an image is a simpler problem than turning the category cat into a particular image. Generation must resolve the many appearances and arrangements compatible with the same description. Better understanding nevertheless helps generation, including through the synthetic labels that vision systems can produce.

Gu recommends learning recognition and understanding because the ability to discriminate quality can support better generation, with reinforcement learning acting as a bridge. His own path moved through generative modeling, robotics and dexterity before language models, motivated by his expectation that symbolic intelligence would advance faster than physical intelligence. He encourages researchers to learn how different research communities frame their problems.

He sees a possible parallel with language models: early creative demonstrations became useful chatbots through instruction tuning, then more reliable reasoning systems as pretraining and post-training improved. Video models might similarly improve instruction following and reliability enough to support spatial and temporal simulation alongside textual reasoning. Erhan adds a present limitation: multimedia understanding and generation have not yet been thoroughly unified in models that excel at both. Their conceptual relationship does not by itself settle how they should be implemented.

19:4019:46
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

19:25 · section reference included

Audio and video share an underlying event

Asked whether audio is fundamentally different from video, Erhan says the technical differences seem relatively minor from his perspective. The more consequential choice was to generate audio and visuals jointly. He describes the team’s model as producing both together, rather than generating pictures and then attaching a separate process to make the lips fit separately generated speech.

The rationale is causal: a person speaking is one underlying event that produces both visible movement and sound. Lip motion and audio therefore need to agree in time. Modeling them together gives the system a way to learn that shared structure and avoids some synchronization failures of a pipeline assembled afterward. Erhan regards joint generation as a successful design choice and says that once users experienced generated video with sound, silent output became much harder to accept.

24:2024:28
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

24:20 · section reference included

What a reference can convey that words struggle to specify

Gu identifies a difficulty beyond specifying spoken words: describing music, a person’s vocal tone or other sensory qualities. He compares this with taste, smell and subtle differences in skin appearance. His hypothesis is that some perceptual sensitivities are tied to primitive survival functions and are finer than ordinary vocabulary. He illustrates the vocabulary gap with a professional wine taster who borrowed language used to describe a romantic partner. This is an explanatory hypothesis and anecdote, not a demonstrated account of the limits of language.

Brichtova extends the same concern to visual style and aesthetic judgment: people can perceive distinctions they struggle to describe. Erhan is more cautious about interpreting that as a fundamental ceiling. The audio failures he encounters still look amenable to more or better data, and he does not think the frontier has been pushed far enough to establish an inherent sensory-language limit.

Direct references offer a practical response. A sample voice can communicate tone and prosody without requiring the user to master specialist vocabulary. There are also two distinct sources of difficulty: the user may lack the words, or the language model may have weak understanding of the relevant domain. If generation depends on that understanding, the weakness propagates into the output. The panel leaves open how much of the problem comes from insufficient attention and training rather than an unavoidable representational limit.

26:3526:42
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

26:28 · section reference included

Convincing scenes need acoustics and microexpressions

Drawing on podcast production, swyx divides audio roughly into music, voice and sound effects, then examines voice alone. A large room, a small room, a car and a phone connection all change what a listener hears. He argues that uniformly studio-quality sound can betray generated video and attributes that tendency to training material. His concrete requirement is spatial consistency: someone farther away should sound softer or more diffuse. Immersive audiovisual generation therefore needs to represent how a scene affects sound, as well as what is being said.

Gu connects this to conditional generation. If a caption leaves out acoustics and other relevant details, many substantially different outputs remain compatible with the same words. He describes a modeling ideal in which a latent representation captures most of the variability, leaving the output given that representation comparatively deterministic. In that framing, richer descriptions help by specifying factors that would otherwise remain ambiguous; they do not establish that language can express every factor.

Brichtova identifies an analogous visual gap in facial expressions, skin texture and the small reactions people read during conversation. A nod or changing microexpression communicates a response to another person. She says video generation has improved substantially but still has considerable headroom in these details. Some still images already appear indistinguishable from reality to her; maintaining that impression through a person’s changing behavior is a harder remaining challenge.

30:5531:06
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

30:55 · section reference included

Preferred does not necessarily mean realistic or useful

Erhan describes an experiment in which the team took descriptions of real videos, generated corresponding videos with Omni and asked humans to compare them. People largely preferred the generated versions. He immediately qualifies the result: the generated clip could be sharper, have a more HDR-like appearance and offer more pleasing skin tones without being more realistic or solving the user’s problem. The account supplies no sample size or numerical preference margin, so it supports a warning about the meaning of preference rather than a quantified performance claim.

Perceptual expertise also changes the judgment. Gu recounts a manga artist’s discomfort with subtly incorrect eye gaze: a small directional error can make an otherwise polished image feel unnatural. Erhan’s conclusion is that simply asking whether people like an output is an unreliable guide to what should be optimized. A broad preference score can miss exactly the distinction a skilled practitioner needs.

This is why Gu expects deliberate prompting to remain useful even as models automate more of it. Control depends on noticing a difference between the current output and the desired result, then communicating that difference. Brichtova distinguishes an ordinary viewer’s preferences from the sensitivity developed through years of design, architecture or illustration. Smooth, saturated output may win casual approval—the discussion calls this the Instagram filter effect—while an expert wants something else. Better instruction following and support for references should let users move away from the default aesthetic.

33:5534:01
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

33:55 · section reference included

Aesthetic defaults create visible and invisible biases

Default aesthetics become consequential when many users accept them. Brichtova recalls seeing widespread Nano Banana Pro infographics and feeling that their default presentation was too cluttered. The model seemed eager to fit everything it knew about a concept into one image. Her observation is personal rather than a measured count of adoption, but the design issue is concrete: displaying more information can undermine the usefulness of an explanatory graphic.

For Omni, the team explicitly compared styles during late tuning, including muted versus saturated color and different palettes. Those choices often land with the modeling team, which raises the question of whether people with an art director’s expertise should play a larger role. Trusted testers and internal users already supply substantial feedback. One example is an optimization that seemed acceptable to the team but made the detail in a user’s grass blurry.

Another user noticed that the model tended to place wedding rings on generated hands, a pattern the developers had not recognized. Gu suggests reward hacking through a spurious preference correlation as a possible explanation. The causal diagnosis is unresolved: the panel does not establish which training signal produced the rings. The example shows how a recurring, semantically meaningful detail can pass unnoticed during development and become obvious to someone inspecting outputs with a different focus.

38:0538:16
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

38:03 · section reference included

Evaluation combines automation, human judgment and workflow tests

The evaluation discussion distinguishes relatively objective checks from aesthetic judgments. Rendered text in an infographic is a tractable example: OCR can extract the text, and a malformed or incorrect letter may make the asset unusable. Automated evaluation of video aesthetics is much harder. Improving Gemini’s understanding can help build better evaluators, but the team still relies extensively on people examining outputs.

For close model choices, Brichtova describes rooms with 10 people watching videos side by side and discussing which they prefer. That is one decision mechanism within a larger evaluation program. When swyx interprets the coverage as only hundreds of examples, she corrects him: human evaluations cover thousands of items. Live experiments add larger-scale evidence, automated raters supply another signal, and experienced users contribute judgments grounded in daily work.

Free-form video editing makes coverage especially difficult because the space of requests is so broad. A model can add objects, change sound or perform other transformations for which no dedicated evaluation yet exists. Strong performance on one slice of human tests can coexist with a broken customer workflow. Early access programs help expose those failures before broader release, making task-level feedback a necessary complement to aggregate scores.

Gu wants human evaluation effort to become reusable through better machine understanding. Detecting errors in generated video is itself a demanding intelligence task. A recreated movie scene might look plausible while becoming semantically inconsistent; identifying the inconsistency requires more than recognizing visual polish. His proposed direction is to turn human judgments into capabilities that can detect such failures, gradually reducing the need to repeat the same manual work.

41:4841:57
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

41:34 · section reference included

Finished media leaves out the decisions that produced it

Asked what data the team wants, Erhan gives a deliberately broad answer: high-quality media, including professionally shot material. He notes existing interest in robotics and embodied data but avoids turning the answer into a detailed account of future projects. The stated need is therefore a quality direction, not a specific procurement specification or a claim that more undifferentiated footage would solve the remaining problems.

Brichtova identifies another valuable kind of data: the trajectory of an actual creative task. A marketing workflow might start with a product photograph, produce a video ad and then adapt it into assets for several advertising formats and platforms. The intermediate requests and decisions reveal what the system must accomplish across the whole job. These trajectories are difficult to manufacture in a lab or obtain through a vendor that lacks the product surface where people do the work.

swyx argues that media workers have abundant routine work they would welcome help with. Brichtova adds that even a familiar marketing campaign involves craft: people iterate, choose one candidate over another and reject details such as incorrect eye gaze. The final asset alone does not explain those choices. A model developer who is not a marketing director may never think to ask for the distinction that determines whether an asset is usable.

Gu calls attention to knowledge that remains inside people: inspirations, conversations and the sequence of choices behind a finished paper or creative work. His assertion that 99% of information is inside people is a rhetorical estimate, not a measured statistic. The substantive distinction is between publicly visible outputs and the process that produced them. The discussion extends this to fiction, where repeated default language patterns can produce readable prose without the personal connection a reader finds in a distinctive story or character.

46:2446:36
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

46:24 · section reference included

Bring real task failures back into model development

The closing discussion gives forward deployed engineers a role in obtaining this missing knowledge. Gu describes growing investment in engineers who work closely with users, while swyx argues that their work should feed evaluations as well as solutions. Gu frames the role as a two-way relationship: customers can improve the harness around a model, and the model team can use the same observations to improve upstream behavior. His unusually broad definition of post-training includes everything between pretraining and the final user experience; it explains his emphasis on connecting customer feedback to modeling.

Brichtova supplies concrete failures that direct conversations reveal. An interior-design user wants the same pattern reproduced across 10 different rug sizes, including a custom size, but the model does not preserve the pattern correctly. A virtual earring try-on needs the earring’s size to make sense relative to the wearer’s head. These requests test consistency and physical scale, not merely whether the output looks attractive. The team benefits from hearing them because its members do not routinely use the models for those jobs.

Brand requirements expose another level of specificity. A brand language may be expressed through images and PDFs, and a broad description such as IKEA’s blue and yellow does not capture the exact shades that matter. The example is illustrative, but the requirement is practical: a model must preserve the distinctions on which a customer’s task depends. Erhan says the goal is to build products people can use for concrete work, which requires understanding what those users care about.

swyx closes by welcoming the progress in generative media and calling for continued exploration beyond coding. The panel ends with appreciation for the work already done and an expectation that substantial development remains.

52:2652:28
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

52:26 · section reference included

Read the complete timestamped transcript
  1. 0:01

    [music]

  2. 0:12

    and welcome back for those on the stream

  3. 0:14

    and those those in person. um we take

  4. 0:17

    tend to basically take these longer

  5. 0:19

    sessions between uh all the sort of

  6. 0:21

    mainstage keynotes to reflect on things

  7. 0:25

    that um you know are particularly

  8. 0:27

    important but like don't have like a

  9. 0:29

    significant like sort of launch moments.

  10. 0:31

    Today we're very lucky to have people

  11. 0:32

    working on Omni and VO Nano Banana like

  12. 0:36

    the you know the world's best generative

  13. 0:38

    models here with us. Uh, Demetrio, I I I

  14. 0:41

    first saw you when you were posting

  15. 0:43

    about your office. [laughter]

  16. 0:46

    Um, I think you're you're probably

  17. 0:48

    number one uh Google Google's number one

  18. 0:51

    office influencer at least in in San

  19. 0:53

    Francisco. I think you like you like to

  20. 0:54

    bike as well. You like to take photos of

  21. 0:56

    >> bike here.

  22. 0:57

    >> Yeah. Um, but you know, but also you

  23. 1:00

    work on video models.

  24. 1:01

    >> That's right.

  25. 1:02

    >> Um, Shane, I I met you I think at like a

  26. 1:04

    dinner.

  27. 1:05

    >> Yeah. Um and uh and uh and I I remember

  28. 1:10

    you were trying to get me invested in

  29. 1:12

    like one of the companies. I forget

  30. 1:13

    forget which one.

  31. 1:14

    >> Forget about that. [laughter]

  32. 1:17

    >> But now but now you're um now you're

  33. 1:20

    working on Omni Thinking. Um and and

  34. 1:22

    just you know a bunch of other

  35. 1:24

    >> Gemini RL.

  36. 1:25

    >> Yeah. Yeah. Uh and Nicole also uh the

  37. 1:28

    rest of the gen media models uh nano

  38. 1:31

    banana and uh all and everything you

  39. 1:33

    just launched actually even this week.

  40. 1:34

    Uh,

  41. 1:35

    >> yeah. We launched some APIs.

  42. 1:37

    >> Yeah. Yeah. Yeah.

  43. 1:38

    >> And I haven't tried to convince you to

  44. 1:39

    invest in anything, but maybe I should.

  45. 1:41

    >> I mean, so I try not to be an investor.

  46. 1:44

    People just convince me anyway. I'm like

  47. 1:45

    just, okay, well, I'm not that rich, but

  48. 1:47

    know like you can't not try to invest in

  49. 1:50

    some of these things. And, you know, for

  50. 1:52

    those of us who are not working at a

  51. 1:53

    Frontier Lab, this is the best this

  52. 1:55

    closest we'll ever get. Um, so yeah,

  53. 1:57

    actually, let's kind of recap since

  54. 1:59

    you're closest to it and we just did it,

  55. 2:00

    like what was launched this week? What

  56. 2:02

    should people go try out?

  57. 2:03

    >> Yeah. Um so yesterday we had two launch

  58. 2:06

    moments. Uh one of them we launched

  59. 2:08

    NanoBanana 2 light uh which is our

  60. 2:11

    fastest, cheapest um image model in the

  61. 2:14

    nano banana model family. Um and it's

  62. 2:17

    better than the original NanoBanana. Um

  63. 2:19

    so really for most people um that model

  64. 2:21

    replaces what you you know used and love

  65. 2:23

    the original Nano Banana for across like

  66. 2:25

    generation and editing and it gets

  67. 2:27

    really close to the frontier quality of

  68. 2:30

    of the kind of mainland bigger models.

  69. 2:32

    So that that's really exciting. I think

  70. 2:33

    if you look at some of the demos or like

  71. 2:35

    things that people have been trying like

  72. 2:37

    getting kind of that like 3 second

  73. 2:38

    latency just unlocks a whole bunch of

  74. 2:40

    things that you can do with like

  75. 2:41

    ideation and iteration and it's just

  76. 2:43

    really fun and the model's getting to a

  77. 2:45

    point where like the quality is really

  78. 2:47

    good um where um it you know you can use

  79. 2:50

    it for iteration but you can also use

  80. 2:51

    some of those outputs as just kind of

  81. 2:52

    like ready um production output. So

  82. 2:54

    that's really exciting. Um and then

  83. 2:56

    second launch we finally um launched the

  84. 2:59

    Gemini Omni Flash APIs um that we

  85. 3:01

    pre-announced at IO. So thank you for

  86. 3:03

    waiting. Um and that you know is the

  87. 3:08

    first time that we're making the APIs

  88. 3:10

    available for developers and it's

  89. 3:11

    basically really exciting kind of video

  90. 3:12

    generation and editing and we're pricing

  91. 3:15

    it the same as Y31 fast. So we're

  92. 3:17

    getting you kind of like really really

  93. 3:18

    good quality for a really awesome price

  94. 3:20

    hopefully. Um

  95. 3:22

    >> yeah, I mean that that's incredible. I'm

  96. 3:24

    actually really So when you guys

  97. 3:26

    launched Omni for the first time, you

  98. 3:28

    also did a podcast uh with Logan who

  99. 3:30

    couldn't be here today uh and you added

  100. 3:32

    like a sloth uh and and Ramen and all

  101. 3:34

    these all these things. I actually

  102. 3:36

    really want to do that to our videos. I

  103. 3:37

    just didn't have an API for it because

  104. 3:39

    obviously I have to automate the whole

  105. 3:40

    thing. So thank you for the API.

  106. 3:41

    >> Uh that is my favorite use case.

  107. 3:43

    Everybody should do that. Um I got a cat

  108. 3:45

    which is probably like the most boring

  109. 3:47

    of the animals. Um if you don't know

  110. 3:48

    what we're talking about, you should

  111. 3:49

    look it up. It's very funny. Feurer. um

  112. 3:51

    Furer who's um you know on on the team

  113. 3:54

    did that.

  114. 3:54

    >> Furer is the number one guy you should

  115. 3:56

    follow for you should follow ideas on

  116. 3:58

    okay what can this thing do?

  117. 4:00

    >> Yes.

  118. 4:00

    >> Right.

  119. 4:01

    >> Yes. He he's he's amazing at that.

  120. 4:03

    >> I've tried to get him for the last two

  121. 4:05

    years to come to AIE. He hasn't made it

  122. 4:07

    yet. He's actually come in person. He

  123. 4:09

    just didn't want to speak because he's

  124. 4:10

    anonymous.

  125. 4:11

    >> I know.

  126. 4:11

    >> I I want to say his real name but I

  127. 4:13

    can't say his real name.

  128. 4:13

    >> No [laughter] no we won't we won't do

  129. 4:15

    that to him. But you should really

  130. 4:16

    follow him. He's amazing.

  131. 4:17

    >> He did all that work. I actually met him

  132. 4:19

    uh in the office uh when we did the

  133. 4:22

    podcast I think and I didn't realize it

  134. 4:24

    was him. So his badge doesn't say

  135. 4:26

    Popers.

  136. 4:27

    >> Yeah,

  137. 4:27

    >> I know.

  138. 4:28

    >> So he used to be part of uh Replicate

  139. 4:30

    and Replicate had this joke where like

  140. 4:32

    everyone was Deep Fates. Deep Fates is

  141. 4:34

    this like kind of mysterious character

  142. 4:35

    and replicate. Replicate is very cool

  143. 4:37

    company and both was part of it. Um, so,

  144. 4:40

    okay, one thing I want to get on there

  145. 4:42

    before I go into like sort of the the

  146. 4:43

    the the sort of omniper is we added

  147. 4:47

    cats, we added sloths, very cool, very

  148. 4:50

    cute, very fun.

  149. 4:51

    >> Uh, what are the, you know, inspire

  150. 4:53

    people as to like what are the more sort

  151. 4:54

    of workhorse use cases that maybe are

  152. 4:57

    not just demos, you know?

  153. 4:59

    >> Yeah. So, so obviously the hero

  154. 5:00

    capability of the model or maybe there's

  155. 5:02

    two like one is the ability to kind of

  156. 5:04

    take in anything as input and then get

  157. 5:06

    video on the other side. Obviously in

  158. 5:08

    the future and and we've kind of talked

  159. 5:09

    about this as a pre-announce like we

  160. 5:11

    want to get the other output modalities

  161. 5:12

    out as well but basically what that

  162. 5:14

    means is you know you can take a set of

  163. 5:15

    images that you have as maybe a

  164. 5:17

    storyboard. You can take like an audio

  165. 5:19

    track as a reference of you know like a

  166. 5:21

    voice that you want a character to speak

  167. 5:23

    and then you can get a video on the

  168. 5:24

    other side. So like that just unlocks a

  169. 5:26

    whole bunch of things that you can do in

  170. 5:27

    like you know short film production or

  171. 5:29

    you know shorts we've launched on

  172. 5:31

    YouTube as well um to help creators kind

  173. 5:33

    of like create um content more easily.

  174. 5:36

    Um and then the other one is obviously

  175. 5:38

    video editing. Like that's another thing

  176. 5:39

    that we're really excited about that

  177. 5:40

    we're just making easier because now you

  178. 5:42

    can use natural language to take a

  179. 5:45

    video, you know, add something, remove

  180. 5:46

    something. Sloth is obviously like fun

  181. 5:48

    example. Um, but there there's obviously

  182. 5:51

    kind of there's consumer use cases that

  183. 5:53

    we kind of had in mind where, you know,

  184. 5:54

    you could take your beach vacation video

  185. 5:56

    that was too noisy and you want to clean

  186. 5:58

    up that noise. Maybe in the past you

  187. 6:00

    wouldn't have because you didn't have

  188. 6:01

    the tools or you didn't know what the

  189. 6:02

    tools were that you needed to go to. So,

  190. 6:05

    that's one use case that you can, you

  191. 6:06

    know, go to. We've seen a lot of folks

  192. 6:08

    use it for kind of marketing ad campaign

  193. 6:11

    creation and I'm excited to see more of

  194. 6:13

    those use cases as we launch the APIs.

  195. 6:16

    um because obviously like we don't we

  196. 6:18

    don't see all of it in the first party

  197. 6:19

    products but I'm really excited for

  198. 6:21

    people to start to explore that um in

  199. 6:23

    the API. So those are just some of the

  200. 6:24

    kind of like high level um things that

  201. 6:26

    have come up. U people also use it to

  202. 6:29

    create like education materials. Yes. Um

  203. 6:31

    and like like that's really exciting. I

  204. 6:34

    think we're all we've all kind of talked

  205. 6:35

    about being excited about the future of

  206. 6:37

    education where like everything can be

  207. 6:39

    kind of customized to you and

  208. 6:40

    personalized to your knowledge level and

  209. 6:43

    the style that you prefer and and so

  210. 6:45

    this is kind of just like a step in that

  211. 6:47

    direction.

  212. 6:47

    >> Yeah. I I I sort of actually used just

  213. 6:49

    none of yesterday, but my my parents are

  214. 6:51

    visiting and there was there was a very

  215. 6:52

    fun sort of use case. They I bought some

  216. 6:55

    gadget off from Amazon that they wanted

  217. 6:57

    and the instructions to use it was were

  218. 6:59

    only in English and there was plenty of

  219. 7:00

    diagrams or whatever and I took a

  220. 7:02

    picture of it and said, you know,

  221. 7:03

    translate this into Romanian. Yes.

  222. 7:04

    >> And keep everything else the same,

  223. 7:06

    right? So it was amazing, right? Like it

  224. 7:08

    was just like, yeah, it looks identical

  225. 7:10

    and it has, you know, it's perfectly

  226. 7:12

    translated. I mean, more or less, right?

  227. 7:14

    But it's it's you know using Gemini

  228. 7:16

    under the hood obviously to kind of do

  229. 7:17

    the translation and so you can you can

  230. 7:19

    see this use case for video as well

  231. 7:21

    right like the the power of text

  232. 7:23

    rendering in in in Omni is is quite next

  233. 7:26

    level. So and you could you could you

  234. 7:28

    could think about plenty of use cases of

  235. 7:29

    like both text rendering translation

  236. 7:31

    internalization all sorts of things that

  237. 7:33

    would be actually genuinely useful to a

  238. 7:35

    lot of different people and sort of

  239. 7:36

    broader access to either you could like

  240. 7:39

    redub a video or whatever it is that you

  241. 7:41

    wanted to do. like there's plenty of

  242. 7:42

    different things that you could you

  243. 7:43

    could think about doing.

  244. 7:45

    >> Yeah. Um one of the most enlightening

  245. 7:49

    conversations I have on my podcast is

  246. 7:51

    with uh just people researchers at the

  247. 7:53

    frontier of these things. Um I had one

  248. 7:55

    with um Ethan from the XAI video team,

  249. 7:57

    the Grock video team who was basically

  250. 8:00

    saying like you know the next trend is

  251. 8:02

    actually not just like single model,

  252. 8:03

    it's more like video agents. Um, and I

  253. 8:06

    don't know if that terminology resonates

  254. 8:09

    uh obviously for for very relevant for

  255. 8:11

    RL. Uh, but it was it was basically kind

  256. 8:13

    of like giving up on like trying to do

  257. 8:14

    everything in in effectively one pass.

  258. 8:17

    Um, do you feel that same way or is it

  259. 8:20

    still an open research question which

  260. 8:22

    way the trends are going? [snorts]

  261. 8:24

    >> Yeah. So um what kind of excite me most

  262. 8:27

    is really when the symbolic kind of

  263. 8:29

    foundational models and this kind of

  264. 8:31

    like video foundational model can

  265. 8:33

    actually kind of really work together

  266. 8:34

    and u in a way the if you look at the

  267. 8:37

    beginning of the generative sort of like

  268. 8:38

    image generation video generation a lot

  269. 8:40

    of it kind of started when the language

  270. 8:42

    model got good enough to provide a very

  271. 8:44

    detailed captioning like from stable

  272. 8:46

    fusion days or kind of dowi 2 days. So

  273. 8:49

    um so basically like language is

  274. 8:52

    extremely u helpful representation uh

  275. 8:55

    one is that it's kind of universal but

  276. 8:56

    the other kind of more um technical

  277. 8:59

    thing like kind of my hypothesis is like

  278. 9:01

    um one very difficult thing about

  279. 9:02

    machine learning is um this sort of like

  280. 9:05

    spirious coordination. So you don't know

  281. 9:07

    you know if the if this kind of feature

  282. 9:10

    right that's kind of predictive is

  283. 9:11

    actually causal factor or not. There are

  284. 9:13

    two ways. One is we can have really

  285. 9:15

    diverse data training data like from

  286. 9:17

    every intervention of the causal graph.

  287. 9:19

    The other is you condition the causal

  288. 9:21

    information and conditioning the

  289. 9:22

    language is kind of like conditioning

  290. 9:25

    like a coal information of the of the

  291. 9:27

    kind of world. So um

  292. 9:29

    >> which is a prompt or a concept what

  293. 9:32

    >> yeah exactly so if you look at like you

  294. 9:34

    know how we going to describe this video

  295. 9:35

    how this kind of image is actually very

  296. 9:38

    close to you know how would describe

  297. 9:39

    this kind of causality you know behind

  298. 9:41

    this like how this is kind of generated.

  299. 9:42

    So one is like that can really allow for

  300. 9:45

    very rich generalization and then uh

  301. 9:48

    very kind of just like a good model. Um

  302. 9:51

    the other is so eight months ago uh we

  303. 9:54

    put the evaluation paper called video

  304. 9:56

    models zero shot learners and reasoners.

  305. 9:58

    >> Yes. So that was a kind of you know it's

  306. 10:01

    it's a confirmed paper and then later on

  307. 10:03

    actually the N banana team followed up

  308. 10:05

    with a vision banana paper that

  309. 10:06

    basically used n banana to do but

  310. 10:08

    essentially the idea is uh video model

  311. 10:11

    is extremely good sort of a foundation

  312. 10:13

    model for space and time kind of

  313. 10:15

    information. So um classic computer

  314. 10:17

    vision tasks a lot of could be kind of

  315. 10:19

    zero shorted and when you like say feed

  316. 10:22

    in some like a visual quiz uh it can you

  317. 10:26

    know there's definitely like a lot to

  318. 10:27

    improve it can kind of solve and it can

  319. 10:30

    um like robotics kind of like seeing it

  320. 10:32

    has really good kind of physical

  321. 10:33

    intuitions like word model uh and I

  322. 10:36

    think the the key is really the kind of

  323. 10:39

    mix of the visual kind of reasoning and

  324. 10:41

    then the text kind of reasoning kind of

  325. 10:43

    all tied together Um obviously you know

  326. 10:46

    like whether doing it you know as kind

  327. 10:47

    of unified model versus like just kind

  328. 10:49

    of agent coation I think that's more

  329. 10:52

    like uh it's going to be more kind of

  330. 10:54

    incremental you know how it's going to I

  331. 10:56

    imagine everything's going to go into

  332. 10:57

    like a single model eventually

  333. 10:59

    >> but right now there's like a lot you can

  334. 11:00

    do if you uh basically take like really

  335. 11:03

    good video understanding image

  336. 11:04

    understanding Gemini agentically with

  337. 11:07

    anomy and that's actually gonna yeah our

  338. 11:09

    team is like exploring a lot

  339. 11:12

    >> yeah okay that there's a there's a lot

  340. 11:13

    in there um I I think uh one question I

  341. 11:17

    I am increasingly starting to wonder is

  342. 11:18

    does it all trend towards one product

  343. 11:20

    for you guys right like now you have

  344. 11:22

    multiple models out the naming of omni

  345. 11:25

    does imply that eventually everything

  346. 11:28

    will go away and it just goes into omnis

  347. 11:30

    um is that the plan

  348. 11:33

    >> is it [laughter] I don't know I I think

  349. 11:36

    I think uh maybe I mean I think

  350. 11:40

    eventually I I think there's sort of

  351. 11:42

    different trade-offs engineering

  352. 11:44

    research product trade-offs in like it's

  353. 11:47

    like for the same reason like the the

  354. 11:50

    sorry how is it called nano banana light

  355. 11:51

    I don't know what the product name

  356. 11:52

    >> nanob banana tite

  357. 11:53

    >> nano banana too light yeah right it's

  358. 11:56

    it's it's it serves a particular niche

  359. 11:58

    right and it probably doesn't

  360. 12:00

    necessarily fit immediately in the same

  361. 12:04

    model literally checkpoint as uh

  362. 12:07

    something that can do 4K you know uh 30

  363. 12:10

    secondond videos right like they're

  364. 12:11

    probably not like trainable in the same

  365. 12:14

    quite way, right? Like, so I I don't

  366. 12:16

    know. It depends on how how far into the

  367. 12:17

    future you look like. Sure, in five

  368. 12:19

    years from now, will they all be the

  369. 12:20

    same model? Probably. Uh but like, you

  370. 12:23

    know, six months from now, we'll we'll

  371. 12:25

    probably still have, you know, multiple

  372. 12:26

    different models doing different things

  373. 12:28

    because kind of from pragmatically the

  374. 12:31

    trade-offs are such that we we should

  375. 12:33

    have multiple different kinds of models.

  376. 12:35

    >> Yeah, I

  377. 12:36

    >> I think that's right. And and just on

  378. 12:37

    that note, I mean, we did call it Gemini

  379. 12:40

    Omni because we wanted to hint at the

  380. 12:42

    future where Gemini just becomes fully

  381. 12:44

    multimodal in and out, right? And so so

  382. 12:46

    it's definitely a move in that

  383. 12:48

    direction. I think we'll probably see a

  384. 12:50

    move in the direction where Omni also

  385. 12:51

    generates images and edits images and

  386. 12:53

    all those kinds of things. But Doo is

  387. 12:55

    right that I think on the way there,

  388. 12:57

    there's a bunch of really really useful

  389. 12:59

    applications of some of these more

  390. 13:01

    specialized models. And so we we will

  391. 13:03

    probably continue to work on those as

  392. 13:04

    well because like that serves a certain

  393. 13:07

    need at this point in time that may not

  394. 13:09

    exist you know a year from now. There's

  395. 13:10

    also like a research question about like

  396. 13:12

    just how much transfer there is between

  397. 13:14

    different kinds of modalities, right? I

  398. 13:16

    think you may believe that there's some

  399. 13:19

    transfer between coding and video

  400. 13:20

    generation and I think most people don't

  401. 13:23

    necessarily believe that but they you

  402. 13:25

    know you could try to think that there

  403. 13:26

    is some some there something there or it

  404. 13:28

    could be a waste right to put them

  405. 13:30

    together to try to learn these both

  406. 13:31

    tasks at the same time right so I think

  407. 13:32

    it's it's it's interesting sort of

  408. 13:34

    question to which extent like image and

  409. 13:36

    video obviously kind of there's some

  410. 13:37

    transfer like kind of not that different

  411. 13:40

    there's value in in learning to output

  412. 13:42

    video and audio at the same time because

  413. 13:44

    joint audio visual is you know that's

  414. 13:46

    how that's how it is. Um and then

  415. 13:48

    there's you know other kind of

  416. 13:49

    intersections of modalities that are not

  417. 13:51

    super obvious right like 3D

  418. 13:52

    representation coding I don't know maybe

  419. 13:55

    uh things like that right so like I

  420. 13:57

    think it's worth sort of exploring the

  421. 13:58

    different corners there and we are

  422. 13:59

    actively doing that um with a focus

  423. 14:02

    towards like what people actually want

  424. 14:03

    to do with these models

  425. 14:05

    >> yeah um what one thing I feel I feel

  426. 14:07

    like uh I'm surprised by but also I feel

  427. 14:11

    like it's insufficiently answered is

  428. 14:13

    what is the correct intermediate

  429. 14:16

    representation Um, so captioning, right?

  430. 14:20

    XI does captioning. Omni does

  431. 14:22

    captioning. Um, and I I I understand how

  432. 14:26

    captioning works for images. Um, and I

  433. 14:28

    understand that you can extend it into

  434. 14:30

    to video and and sort of guide it across

  435. 14:33

    time. It just feels very inefficient. It

  436. 14:35

    there's got to be I feel like there

  437. 14:37

    should be something better. Uh maybe

  438. 14:38

    it's code and maybe we generate you know

  439. 14:41

    and obviously I think a lot of um ffmpeg

  440. 14:45

    and mapplot um what's the three blue one

  441. 14:48

    brown one manm um a lot of like video is

  442. 14:51

    generated through code and maybe that's

  443. 14:53

    like the optimal representation uh any

  444. 14:56

    hypothesis as to like is is it better or

  445. 14:59

    is just English all you need

  446. 15:01

    >> well as so I'm in the Gemini and they

  447. 15:03

    know we do like a lot of RL agent and of

  448. 15:05

    course kind of coding so yeah We we're

  449. 15:08

    definitely exploring the coding

  450. 15:09

    representations.

  451. 15:10

    >> Yeah.

  452. 15:10

    >> As kind of better kind of way to

  453. 15:12

    represent. Yeah.

  454. 15:13

    >> But you know like do you what's your

  455. 15:15

    probability estimate on like [laughter]

  456. 15:18

    if we just output binaries like we just

  457. 15:20

    you know like just it's just ones and

  458. 15:21

    zeros.

  459. 15:23

    >> Um I I guess maybe a kind of similar

  460. 15:27

    discussion was like um basically is the

  461. 15:31

    language the right representation like

  462. 15:33

    right. So uh one kind of question for

  463. 15:35

    example uh professor you know like some

  464. 15:37

    ask is like you know why why does the

  465. 15:39

    channel of thought need to be in the

  466. 15:41

    natural language?

  467. 15:42

    >> Yes.

  468. 15:42

    >> Can it just be the kind of any kind of

  469. 15:44

    like continuous tokens just any amount

  470. 15:46

    of you know additional computations.

  471. 15:48

    >> Um so one is like obviously the test

  472. 15:52

    like adaptive compute is going to give

  473. 15:55

    like you know better results. So it's

  474. 15:56

    that but what really kind of made CH

  475. 15:58

    thought so you know like four years ago

  476. 16:00

    I wrote you know the larger model zero

  477. 16:02

    sort reasoner and then self-improvement.

  478. 16:04

    So I kind of know from the very early

  479. 16:05

    day but the reason like it works really

  480. 16:08

    well is um right now the recipe that

  481. 16:11

    works is the pre-training that scales a

  482. 16:13

    lot and then that basically like learns

  483. 16:15

    a lot of intelligence. there are a lot

  484. 16:17

    of you know scaling RL but those are

  485. 16:19

    still like extremely kind of comput

  486. 16:21

    incent intensive to extract the

  487. 16:23

    information and um you really want to

  488. 16:26

    rely the intelligence on that so

  489. 16:28

    basically by tying the sort of like a

  490. 16:31

    reasoning in the natural language you

  491. 16:32

    basically directly use the intelligence

  492. 16:34

    of the pre-training to it while if you

  493. 16:36

    remove that kind of constraints then

  494. 16:38

    you're not um and these days uh I feel

  495. 16:42

    the a lot of advancements in the texts

  496. 16:45

    but also doing this kind of multimodal

  497. 16:47

    space is very driven by this uh kind of

  498. 16:50

    text as a kind of great uh sort of

  499. 16:53

    representation.

  500. 16:54

    >> Yeah, it's a good backbone.

  501. 16:56

    >> Yeah,

  502. 16:57

    >> I I think to me it's even simpler than

  503. 16:58

    that. It's text is is how we

  504. 17:00

    communicate. So I think fundamentally if

  505. 17:02

    you're building kind of products that

  506. 17:03

    humans will be interfacing with um like

  507. 17:07

    like that we will be using text somehow

  508. 17:09

    if it's a text interface, right? Not not

  509. 17:11

    for everything. So I think it's it's

  510. 17:13

    natural to default to that.

  511. 17:15

    >> Yeah. Obviously there's like a conf

  512. 17:17

    discussion you know some arrow like

  513. 17:18

    arrow maximalists is like oh we don't

  514. 17:20

    care about you know kind of channel

  515. 17:22

    those kind of like stuff it's just just

  516. 17:24

    additional compute

  517. 17:24

    >> sure but I personally yeah

  518. 17:26

    >> ro maximalists I wonder I wonder who who

  519. 17:30

    qualifies in that description David

  520. 17:31

    silver

  521. 17:32

    >> ah okay yeah I mean they they've just

  522. 17:35

    left to to start their thing um

  523. 17:38

    interesting okay so uh I I mean I think

  524. 17:41

    I'm very interested in just like better

  525. 17:43

    representations because I think that's

  526. 17:44

    one of our themes that we're curating

  527. 17:46

    today uh at the world fair is world

  528. 17:48

    models. You mentioned the word world

  529. 17:50

    models but it's not something that's

  530. 17:51

    like super well defined. I think

  531. 17:52

    everyone's like sort of converging on

  532. 17:54

    some version of it that is like the

  533. 17:56

    ideal.

  534. 17:57

    >> Sure. Everything is a world model now.

  535. 17:59

    It's sort of a

  536. 18:00

    >> it's not it's not that useful, right?

  537. 18:02

    >> So I just gave a keynote at the IER

  538. 18:05

    world model workshop. Yeah. And then uh

  539. 18:07

    yeah essentially uh I definitely

  540. 18:09

    encourage to check out the definition by

  541. 18:10

    Jandra Matic. He's like the you know OG

  542. 18:13

    computer vision professor UC Berkeley.

  543. 18:15

    >> Uh he has pretty you know bit of word to

  544. 18:17

    say about world model

  545. 18:18

    >> but also kind of Schmidt Herburver's

  546. 18:19

    kind of how he defined the world model

  547. 18:21

    from 2019 like 1990 sort of uh uh you

  548. 18:26

    know like Wayne was just basically just

  549. 18:27

    that kind of model base. Uh for me the

  550. 18:29

    word model is basically just the model

  551. 18:30

    in the model based RL and I feel that

  552. 18:32

    has sufficient to describe but obviously

  553. 18:34

    you know there are like a lot of uh FE

  554. 18:36

    had a kind of nice blog post about what

  555. 18:39

    about yeah this kind of broken down

  556. 18:41

    >> um but yeah

  557. 18:43

    >> yeah I mean so you know I I'll end this

  558. 18:47

    part of the conversation but like I I do

  559. 18:49

    think that language to me relying on

  560. 18:52

    language as like the sort of like the

  561. 18:53

    narrow pipe through which everything

  562. 18:54

    goes through um still is like a lossy

  563. 18:57

    compression. No, no, no. But we're not

  564. 18:59

    seeing that, right? We're basically

  565. 19:00

    saying the video model and the language

  566. 19:02

    together.

  567. 19:02

    >> So, so I think the language alone is uh

  568. 19:05

    not sufficient. That's why we feel like

  569. 19:07

    the video is a very complement model.

  570. 19:09

    Right? Now the um you know kind of v

  571. 19:11

    omni many people feel as uh you know

  572. 19:14

    generating kind of pretty videos but I

  573. 19:16

    think our vision it's it's much more

  574. 19:18

    than that. It's a missing foundational

  575. 19:19

    model that's absolutely required if you

  576. 19:21

    want to make the AGI that match to

  577. 19:23

    humans not just a jacked one.

  578. 19:25

    >> Yeah. Um okay. So one one other thing

  579. 19:28

    you know you you mentioned on the vision

  580. 19:30

    side um and I'm kind of curious how sort

  581. 19:34

    of uh parallel you know in terms of your

  582. 19:37

    research careers um this development is

  583. 19:40

    like I think basically a lot of vision

  584. 19:41

    people have crossed over into more model

  585. 19:44

    people um a lot of vision people also

  586. 19:46

    become generative video and image people

  587. 19:50

    and is it just as simple as you know

  588. 19:53

    reversing uh image to text and then now

  589. 19:55

    it's text to image like

  590. 19:57

    >> [laughter]

  591. 19:58

    >> is is that if I mean that effectively

  592. 20:00

    was the diffusion process. Um

  593. 20:03

    I I just you know I I just see the

  594. 20:06

    career paths of the people that I talked

  595. 20:07

    to and and see and I I I see this

  596. 20:10

    overall trend of research directions and

  597. 20:12

    I just wanted you to guys to sort of

  598. 20:14

    reflect on on that.

  599. 20:16

    >> I mean I certainly went that way right I

  600. 20:17

    I started long time ago uh doing

  601. 20:20

    computer vision sort of object detection

  602. 20:23

    recognition things like that. Uh I think

  603. 20:24

    just that's just simpler problem right

  604. 20:26

    just generation is just harder like it's

  605. 20:28

    a it's a different kind of mapping right

  606. 20:30

    you map from the the inverse mapping is

  607. 20:32

    not as simple as just inverting the the

  608. 20:34

    kind of network you use right it's it's

  609. 20:36

    a it's it's more ambiguous right to go

  610. 20:38

    from cat to image of a cat and in some

  611. 20:41

    ways it's also a loop because your

  612. 20:42

    vision work creates the synthetic labels

  613. 20:44

    that then continues

  614. 20:46

    >> I mean sure [laughter]

  615. 20:48

    I don't know I don't know I try to

  616. 20:49

    validate my my sort of theories about

  617. 20:51

    how fields develop how how careers has

  618. 20:53

    progressed through this

  619. 20:55

    >> I mean for like the the the better the

  620. 20:57

    understanding side gets like we have

  621. 20:59

    seen that the generation side also gets

  622. 21:02

    better right so like like

  623. 21:03

    >> it's completely bootstrapping yeah it's

  624. 21:05

    >> and so like like like there's definitely

  625. 21:07

    they're there to that thesis and I think

  626. 21:09

    yeah I think a lot of people have kind

  627. 21:10

    of like I I definitely worked with a lot

  628. 21:12

    of um image understanding people who

  629. 21:14

    became image generation people you know

  630. 21:16

    and then some of them have moved on to

  631. 21:17

    video because it's kind of like the next

  632. 21:19

    thing where you have so many more

  633. 21:20

    dimensions to work with so yeah I'm

  634. 21:22

    curious about you spec as your

  635. 21:24

    >> so I definitely like recommend start

  636. 21:26

    with understanding recognition because

  637. 21:28

    that's basically discriminator and then

  638. 21:29

    that's going to lead to better

  639. 21:30

    generation and that's what the bridge is

  640. 21:32

    basically reinforcement learning so my

  641. 21:35

    um my kind of journey is I initially

  642. 21:36

    kind of worked on the algorithmic

  643. 21:38

    research in the gent model against some

  644. 21:40

    like you know amnest kind of generation

  645. 21:42

    and then I worked on like RL and

  646. 21:44

    robotics um and then like six years ago

  647. 21:47

    I was like leading like a moonshot on

  648. 21:49

    the dexterity it was pretty early but I

  649. 21:51

    see now everyone's kind of doing

  650. 21:53

    uh four years ago I basically kind of

  651. 21:54

    figured out that this like symbolic AGI

  652. 21:58

    is going to accelerate much faster than

  653. 22:00

    the kind of physical AGI kind of

  654. 22:01

    counterpart. So uh I decided to kind of

  655. 22:04

    like language models and then those

  656. 22:06

    things. Um and then recently kind of

  657. 22:08

    work with Doomi and then like omni team

  658. 22:10

    I quite enjoy kind of collaboration

  659. 22:12

    there. the what I quite enjoy uh what I

  660. 22:15

    recommend definitely to the researcher

  661. 22:17

    is to uh definitely kind of explore or

  662. 22:20

    at least like get exposure to what the

  663. 22:22

    top people in each of the community are

  664. 22:24

    like looking at how they kind of think

  665. 22:26

    about problems. So when I look at the

  666. 22:28

    video model to me it kind of reminds me

  667. 22:30

    like pretty early on sort of like

  668. 22:32

    language model where like very early

  669. 22:34

    language model was a kind of creative

  670. 22:36

    sort of demo right you kind of like try

  671. 22:39

    to write like a story like mobile and

  672. 22:41

    then like you know GBD2 and then those

  673. 22:43

    kind of days like LTM kind of days right

  674. 22:45

    and then you know uh instruction tuning

  675. 22:48

    you actually kind of make it usable as a

  676. 22:50

    chatbot but then at the chatbot stage it

  677. 22:52

    still had so much hallucinations and

  678. 22:54

    instruction for wasn't good enough so it

  679. 22:56

    couldn't use for reasoning and when I

  680. 22:58

    got good enough um in pre-training and

  681. 23:00

    post- trainining for reasoning then you

  682. 23:02

    know this kind of test time scaling the

  683. 23:04

    RL really took off to like many of the

  684. 23:06

    kind of best performing models and right

  685. 23:08

    now I think the video model is as we

  686. 23:10

    mentioned it's it is a complimentary

  687. 23:12

    foundational model and I can imagine

  688. 23:13

    it's going to follow a similar path it's

  689. 23:15

    going to be very uh it's going to

  690. 23:17

    improve a lot instruction following a

  691. 23:19

    lot of uh this it's going to improve a

  692. 23:21

    lot in reducing coordinations to extend

  693. 23:23

    that it become a very reliable world

  694. 23:25

    model so we can kind of like intermixed

  695. 23:27

    video like space-time simulation was a

  696. 23:30

    text simulation to solve like arbitrary

  697. 23:31

    AGI problems. Also like I think the

  698. 23:34

    difference still is between sort of text

  699. 23:36

    models and like image video models is

  700. 23:38

    that like we haven't quite unified

  701. 23:39

    understanding and generation in in

  702. 23:41

    multimedia I'd say yet like I mean I

  703. 23:44

    think I think without going to the

  704. 23:45

    details of course there's like it

  705. 23:46

    depends on on at which level you're

  706. 23:48

    thinking about this but generally like

  707. 23:51

    there's not that many as far as I know

  708. 23:52

    models sot kind of you know frontier

  709. 23:55

    models that are genuinely

  710. 23:58

    kind of good at both understanding and

  711. 24:01

    generation of of let's videos, right?

  712. 24:04

    Like it's a it's a it's an interesting

  713. 24:05

    challenge. I'm not saying that we should

  714. 24:07

    do this. Uh but but I think uh it kind

  715. 24:10

    of stands to reason that like you know

  716. 24:12

    understanding and generation are two

  717. 24:13

    sides of the same coin. So they they

  718. 24:15

    kind of should be in the same model in

  719. 24:16

    some ways. Uh but we don't necessarily

  720. 24:18

    always do that. So yeah.

  721. 24:20

    >> Uh you mentioned audio as well, right?

  722. 24:22

    Yeah. Uh is that as hard as video or

  723. 24:28

    qualitatively different? If if so, in

  724. 24:30

    what way? Uh, one of the interesting

  725. 24:33

    directions three years ago was people

  726. 24:36

    using um, I guess diffusion to do audio

  727. 24:41

    uh, as in like the the sort of refusion

  728. 24:44

    approach. I don't know if you you guys

  729. 24:45

    saw that. Um, and I just think it's like

  730. 24:47

    very interesting if a modality that we

  731. 24:51

    perceive which is audio is different

  732. 24:52

    than video actually two machines is

  733. 24:54

    exactly the same like there's they see

  734. 24:57

    no difference.

  735. 24:59

    I mean I think on a technical level

  736. 25:01

    there are some differences but I think

  737. 25:02

    they're like relatively minor. I think

  738. 25:04

    from my perspective audio came into into

  739. 25:07

    my life when we shipped V3 which was I

  740. 25:10

    believe the first model that did like a

  741. 25:12

    joint

  742. 25:12

    >> with the slicing of the

  743. 25:14

    >> Yeah. Yeah. the gold bars or whatever.

  744. 25:16

    Um it it was the first model that did

  745. 25:18

    this sort of joint audiovisisual

  746. 25:19

    generation. Yes. uh like in the in the I

  747. 25:22

    mean there are there were other models

  748. 25:23

    that did kind of you know kind of kind

  749. 25:25

    of agentic hacking under the hood but

  750. 25:26

    this one was truly sort of you know

  751. 25:29

    generating everything at once and we the

  752. 25:32

    reason we did that is because we felt

  753. 25:34

    and I think was the right choice we felt

  754. 25:36

    that like uh it only makes sense to

  755. 25:39

    generate them at the same time because

  756. 25:40

    there sort of kind of like from a

  757. 25:42

    machine learning perspective there's one

  758. 25:43

    latent kind of you know causal kind of

  759. 25:44

    you know generative process right like

  760. 25:46

    there's something that generates you

  761. 25:48

    speaking it's not the pixels and then

  762. 25:50

    the the audio or somehow somehow

  763. 25:52

    generated by some other process like the

  764. 25:53

    lips have to move in sync with with the

  765. 25:55

    with the audio, right? So, I think that

  766. 25:57

    that solved a lot of the issues that

  767. 25:59

    previous models had or the way that

  768. 26:00

    people did video generation before where

  769. 26:02

    it was like, okay, we generate the

  770. 26:03

    pixels and then we're going to hack

  771. 26:05

    something on top of it that like moves

  772. 26:06

    the lips with the audio that we

  773. 26:08

    generate. And that's was very bad.

  774. 26:10

    [laughter]

  775. 26:11

    And so, I think I think that was that's

  776. 26:13

    to me that's the the I mean after V3

  777. 26:16

    like you know people were like what do

  778. 26:17

    you mean like there's no audio in your

  779. 26:18

    model? like that makes no sense like

  780. 26:20

    once it's there like you you have to

  781. 26:21

    have it. So I think that was that was

  782. 26:23

    the right choice and doing it to one

  783. 26:25

    single generative model I think was was

  784. 26:27

    the right choice.

  785. 26:28

    >> One thing I kind of want to also can ask

  786. 26:30

    you guys an opinion as well once one

  787. 26:32

    difference I find the audio and then

  788. 26:34

    against the image and video is like the

  789. 26:35

    audio information is less verbalized. I

  790. 26:38

    mean of course the TTS and stuff is

  791. 26:40

    trivial right but the when you get her

  792. 26:42

    outside like how to describe music how

  793. 26:45

    do you describe this like this person's

  794. 26:47

    tone kind of pitch I feel the sort of

  795. 26:50

    the verbalization is insufficient and

  796. 26:52

    the interesting thing is that you kind

  797. 26:54

    of see that in two other things like

  798. 26:56

    taste taste sense and also uh say um

  799. 27:00

    touch

  800. 27:00

    >> like smell and then the another

  801. 27:02

    interesting thing is the skin color so

  802. 27:05

    skin color the the language is pretty

  803. 27:08

    limited to describe the skin color and

  804. 27:10

    the reason is that we're extremely uh

  805. 27:12

    sensitive to the small difference

  806. 27:13

    perturvations or not skin color because

  807. 27:15

    that basically shows us is this person

  808. 27:17

    going to kill me or is can I befriend

  809. 27:19

    this person kind of those kind of

  810. 27:21

    information and then I feel the smell

  811. 27:22

    tastes um skin color and like sound kind

  812. 27:26

    of stuff is very very tied into

  813. 27:28

    primitive a like survival kind of stuff

  814. 27:32

    and so our sort of sensory system is so

  815. 27:34

    sensitive that it's intractable to Um,

  816. 27:38

    so for example, I asked like one the

  817. 27:40

    wine sort of taster and then like

  818. 27:42

    professional and then he basically said

  819. 27:43

    he kind of use like a language from like

  820. 27:45

    a dating, you know, describing like a,

  821. 27:47

    you know, partner as a way to describe

  822. 27:50

    the taste because there's no sufficient

  823. 27:52

    vocab to describe. Um, so I'm kind of

  824. 27:56

    curious. Yeah. Do you guys feel that?

  825. 27:58

    >> I think well to some extent I think the

  826. 28:00

    same is true for visual information,

  827. 28:03

    right? when you think about like a

  828. 28:06

    certain style or a certain aesthetic,

  829. 28:08

    right? Like like there are some people

  830. 28:10

    who just have a much more kind of

  831. 28:11

    developed like whether it's palette or

  832. 28:13

    kind of visual taste and aesthetic,

  833. 28:15

    right? Like I I think language just

  834. 28:18

    tends to be a bit of a limiting factor

  835. 28:20

    when you are trying to describe any of

  836. 28:22

    these things that like we experience

  837. 28:24

    with sensory information. And to your

  838. 28:26

    point earlier, I think that is the kind

  839. 28:29

    of the reason why we are investing in

  840. 28:31

    world models and why we are pushing on

  841. 28:33

    kind of the like perception and like

  842. 28:34

    generation side of things because it it

  843. 28:37

    is such a large part of how we as humans

  844. 28:40

    navigate the world. It's a large part of

  845. 28:42

    how like embodied AI navigates the

  846. 28:45

    world. Um, and and I do I do think

  847. 28:48

    language like does have a lot of it's

  848. 28:50

    it's gotten us very far and it can

  849. 28:51

    probably get us really far, but it it

  850. 28:53

    feels limiting in a lot of these kind of

  851. 28:55

    areas. And yeah, I don't I don't really

  852. 28:57

    know how to describe, you know, like

  853. 28:58

    sense and taste. Um, but yeah, I'm

  854. 29:01

    curious to me.

  855. 29:03

    >> Um, I I yeah, I don't know that I have

  856. 29:06

    thought that deeply about this yet. So,

  857. 29:08

    uh, yeah, I mean yeah, I don't have a

  858. 29:11

    good answer about audio. I mean like I

  859. 29:14

    don't know the limit because I'm

  860. 29:16

    thinking about like well what is what is

  861. 29:17

    Omni bad at in terms of audio but

  862. 29:19

    they're all like solvable problems I

  863. 29:21

    find uh so like with more data or better

  864. 29:24

    data or whatever it is so I don't know

  865. 29:26

    like that we have pushed the frontier so

  866. 29:28

    much that like we are have hit some sort

  867. 29:31

    of limits that are rooted in

  868. 29:33

    evolutionary uh kind of you know limits

  869. 29:36

    imposed by humans. I don't know. He's

  870. 29:38

    feeling the limits of captioning which

  871. 29:40

    is the the thing I was

  872. 29:41

    >> Yeah, exactly. [laughter] There there's

  873. 29:42

    a lot of information in the world and it

  874. 29:44

    connects to basically why we do work

  875. 29:45

    modeling you mentioned. You just need

  876. 29:47

    srefs sref476

  877. 29:50

    and then that's your that's what your

  878. 29:51

    journey does, right? I guess maybe I

  879. 29:53

    can't describe this vibe but

  880. 29:54

    >> well well I think that that's kind of

  881. 29:56

    the point of providing some of these

  882. 29:57

    references, right? Because because like

  883. 29:59

    even just describing how someone talks

  884. 30:02

    and like their tone and and like procity

  885. 30:04

    and all of these things like I think I

  886. 30:05

    think some of these terms even like I

  887. 30:08

    didn't used to know what they mean,

  888. 30:09

    right? Well, now

  889. 30:10

    >> yes. Dispuencuencies ex like like there

  890. 30:14

    there's kind of an entire vocabulary

  891. 30:16

    that even if you're not kind of steeped

  892. 30:17

    in a domain, which is true for actually

  893. 30:19

    like most human domains that like you

  894. 30:21

    don't even know what it means. Um and

  895. 30:23

    sometimes it's also a question of like

  896. 30:24

    if we haven't focused on those things,

  897. 30:26

    you know, with the large language models

  898. 30:27

    that they may also have gaps in those

  899. 30:29

    areas, right? And then we feel them on

  900. 30:30

    the other side with generation because

  901. 30:32

    we're like fundamentally relying on on

  902. 30:34

    the language models understanding of the

  903. 30:36

    world to then be able to like represent

  904. 30:38

    it. Um, so I yeah, it all kind of goes

  905. 30:40

    back to your question about like the the

  906. 30:42

    language as an intermediary. Um, but

  907. 30:44

    yeah, I think to De's point like some of

  908. 30:46

    these might just be like focus areas and

  909. 30:48

    things that we haven't necessarily

  910. 30:49

    pushed on as much as we can and like as

  911. 30:51

    we will we will discover what the actual

  912. 30:54

    ceiling is.

  913. 30:55

    >> Yeah, as a podcaster I think a lot about

  914. 30:58

    sound.

  915. 30:59

    >> Um, and and I I'll just offer a couple

  916. 31:02

    things for discussion in case in case it

  917. 31:03

    triggers anything with you guys. Um I

  918. 31:06

    have three domains of rough audio which

  919. 31:07

    is like a music voice SFX you know is

  920. 31:10

    that rough okay covers everything and

  921. 31:13

    then also even within voice let's just

  922. 31:15

    let's just focus on voice forget the

  923. 31:16

    other two um room sound like the the

  924. 31:18

    echoiness of like big room small room in

  925. 31:21

    person in a car over a phone all these

  926. 31:25

    like are labelable but we experience

  927. 31:27

    them very differently and I I often

  928. 31:29

    think like one of the tells of a AI

  929. 31:31

    video is that it is studio quality

  930. 31:33

    because it was recorded in a studio

  931. 31:35

    video because that's your training data

  932. 31:36

    and like and and to me that's one thing

  933. 31:39

    actually like the most interesting thing

  934. 31:40

    is just uh when I tell this is how I

  935. 31:43

    convince people who are kind of

  936. 31:44

    skeptical about the need for world

  937. 31:46

    models because you need it even for

  938. 31:48

    audio about well I'm further away from

  939. 31:51

    you so I should sound a little bit

  940. 31:52

    softer or more diffused and like the the

  941. 31:55

    video models need to pick that up

  942. 31:56

    because if they're going to do immersive

  943. 31:58

    video and audio you need that

  944. 32:01

    >> I I I love that example of basically

  945. 32:03

    like studio quality or not in a way like

  946. 32:05

    we don't have enough language to really

  947. 32:07

    describe like like this kind of echoing

  948. 32:10

    or like some kind of noise kind of

  949. 32:12

    happening we just like don't have

  950. 32:13

    precise enough and uh if you um you know

  951. 32:16

    basically the reason that I think it's

  952. 32:18

    quite important to have like relatively

  953. 32:20

    information rich like kind of captioning

  954. 32:22

    is that we kind of rely on the natural

  955. 32:23

    language as a representation but if you

  956. 32:25

    basically don't have enough uh

  957. 32:27

    representation that basically means the

  958. 32:29

    condition on the language the generation

  959. 32:30

    is very multimodal and if you anything

  960. 32:33

    can learn from the BAE kind of like you

  961. 32:34

    know very old you know GMBA kind of

  962. 32:36

    research the idea is we really want to

  963. 32:38

    capture most of the stoasticity in the

  964. 32:40

    later representation and then the the X

  965. 32:43

    given the Z should be kind of like

  966. 32:44

    deterministic so yeah

  967. 32:46

    >> yeah yeah um well I hope I hope there's

  968. 32:49

    more uh progress there and I'm sure you

  969. 32:50

    guys are doing

  970. 32:51

    >> I even actually like facial expressions

  971. 32:53

    right and maybe this gets to your point

  972. 32:55

    about like things that we're very

  973. 32:56

    sensitive to right I think you can tell

  974. 32:58

    a lot of AI content also just by from

  975. 33:01

    like people's facial expressions

  976. 33:02

    stressful.

  977. 33:04

    [laughter]

  978. 33:04

    >> Yes.

  979. 33:06

    >> Yes. And we try not to contribute to it,

  980. 33:08

    but you know, um and or or like skin

  981. 33:11

    textures, right? Like like the things

  982. 33:12

    that kind of make things look real in

  983. 33:15

    real life. Like I you know, I can tell

  984. 33:16

    from the way you're nodding or from the

  985. 33:18

    way like your micro expressions are kind

  986. 33:19

    of changing of like how you're reacting

  987. 33:21

    to what I'm saying. Like we haven't

  988. 33:23

    quite crossed that chasm. I think like

  989. 33:26

    we're we're so much better than we were

  990. 33:27

    a year ago.

  991. 33:28

    >> Yeah. Um, but there's so much more

  992. 33:30

    headroom kind of in a lot of those

  993. 33:32

    things that like we as humans are super

  994. 33:33

    sensitive to. And like I think image

  995. 33:35

    arguably probably is there because

  996. 33:38

    there's there's a lot of kind of images

  997. 33:40

    that I will see that like really do look

  998. 33:42

    indistinguishable from reality and I

  999. 33:44

    can't tell if they're generated or not.

  1000. 33:46

    >> Better than reality

  1001. 33:47

    >> um or well that's a different

  1002. 33:49

    >> No, I I think that one of the parad

  1003. 33:51

    [laughter] better than what I would take

  1004. 33:52

    on my vacation as a photo. Yes. One of

  1005. 33:55

    the one of the fun experiments that we

  1006. 33:56

    did a while ago in the team is is like

  1007. 33:58

    can we generate videos that are better

  1008. 34:00

    than than real videos, right? So you

  1009. 34:01

    just take the same caption from like oh

  1010. 34:03

    yeah some video and then

  1011. 34:06

    >> recycle it. Yeah. Just just try to like

  1012. 34:08

    describe a real video and then generate

  1013. 34:10

    the equivalent version with omni and

  1014. 34:12

    then do a human eval. How does how does

  1015. 34:14

    it do? And then humans largely prefer AI

  1016. 34:17

    generated

  1017. 34:18

    >> margin.

  1018. 34:20

    >> But because it's because it's the RL

  1019. 34:21

    process,

  1020. 34:22

    >> that's the RL process working.

  1021. 34:23

    >> It's however you want to rationalize it.

  1022. 34:25

    It's not necessarily the old process.

  1023. 34:26

    It's just like I think it's just I'm not

  1024. 34:28

    saying this is a good result. I'm just

  1025. 34:30

    saying is we have optimized in a way

  1026. 34:33

    that like kind of potentially sort of,

  1027. 34:35

    you know, triggers something in the

  1028. 34:36

    human brain that like, oh, it looks it

  1029. 34:38

    looks all a lot of the videos just look

  1030. 34:41

    better. Like I'm not Yeah. Yeah. Yeah.

  1031. 34:43

    on on inspection on on deeper inspection

  1032. 34:45

    they they would not actually be more

  1033. 34:47

    useful or whatever but like if you just

  1034. 34:50

    say side by side random YouTube video

  1035. 34:52

    versus

  1036. 34:53

    >> generated version of it will you will

  1037. 34:56

    just have a it will just look better

  1038. 34:57

    because it's more it's a sharper more

  1039. 34:59

    HDR uh you know the skin tone is is is

  1040. 35:03

    better it's not again it's not more

  1041. 35:04

    realistic

  1042. 35:06

    >> uh it doesn't solve your problem

  1043. 35:07

    necessarily but it it looks better

  1044. 35:09

    >> I I since also depend on the sensitivity

  1045. 35:12

    of the people. Uh I was born raised in

  1046. 35:14

    Japan and I think one thing I kind of

  1047. 35:16

    know is like they're extremely extremely

  1048. 35:17

    like sensitive about like you know

  1049. 35:19

    that's why you know like architecture

  1050. 35:21

    like food and stuff like they have.

  1051. 35:23

    >> Um so I talked to like a mangar like

  1052. 35:25

    like artist there and he's like he's

  1053. 35:26

    kind of disgusted by like the generation

  1054. 35:28

    AI and one kind of thing he mentioned is

  1055. 35:31

    like the eye gaze

  1056. 35:32

    >> eye gaze that slight difference

  1057. 35:35

    >> makes me makes him kind of feel creepy

  1058. 35:37

    about like unnatural

  1059. 35:38

    >> like if you're looking a little bit off.

  1060. 35:40

    >> Yeah. It's just uh Yeah. just like uh it

  1061. 35:43

    looks too fake. Yeah. To the point. So

  1062. 35:45

    So I think it does depend on the

  1063. 35:46

    sensitivity and

  1064. 35:47

    >> Yeah. Yeah. All I'm saying is like you

  1065. 35:49

    know human preferences are like a not

  1066. 35:51

    particularly like uh reliable barometer

  1067. 35:54

    of like what you should be optimizing

  1068. 35:56

    for like if you just ask people do you

  1069. 35:57

    like this or not you not necessarily get

  1070. 35:59

    what you wanted.

  1071. 36:01

    >> Yeah. Let let me just kind of add one

  1072. 36:02

    thing but like four years ago there was

  1073. 36:04

    a like debate that if the prompt

  1074. 36:05

    engineering is going to disappear and uh

  1075. 36:07

    my my like you know some very powerful

  1076. 36:10

    people say you know it's going to

  1077. 36:11

    disappear but I basically said like it

  1078. 36:13

    shouldn't because the prompt engineering

  1079. 36:16

    like sort of you know specifying that is

  1080. 36:18

    like the the only way you can sort of

  1081. 36:20

    control the output sort of you know when

  1082. 36:22

    you have like sort of control by the AI

  1083. 36:25

    and what allows you to prompt engineer

  1084. 36:27

    is really that sensitivity. So sure

  1085. 36:30

    maybe like right now the AI can do a lot

  1086. 36:32

    of autoprompting and that and it can

  1087. 36:34

    generate something that's sufficient but

  1088. 36:37

    uh if it's like that never be satisfied

  1089. 36:39

    like never be satisfied with the AI's

  1090. 36:41

    generated content always fine-tune your

  1091. 36:43

    sensitivity and always kind of keep

  1092. 36:45

    prompting the differences. I I think to

  1093. 36:48

    the there's also a big difference

  1094. 36:49

    between like the average human untrained

  1095. 36:53

    eye which I I would put myself in that

  1096. 36:55

    bucket you know like I have I have some

  1097. 36:56

    aesthetic sensibilities and I've done

  1098. 36:58

    this long enough that you know like I

  1099. 37:00

    have I have a preference um but you know

  1100. 37:02

    like your example of a manga artist like

  1101. 37:05

    that's somebody who has honed a craft

  1102. 37:07

    like over possibly many decades. Um, and

  1103. 37:11

    anybody who does that, whether it's like

  1104. 37:12

    design, architecture, right? Like you

  1105. 37:14

    you you just have a very different level

  1106. 37:16

    of like expertise and you see things

  1107. 37:19

    that like the average human will not

  1108. 37:21

    see. But Doom is right. Like when we

  1109. 37:22

    look at if you were to just, you know,

  1110. 37:25

    um, poll 10 people on the street, they

  1111. 37:28

    would probably prefer the like overly

  1112. 37:31

    smooth like very saturated kind of

  1113. 37:33

    >> It's called the Instagram filter.

  1114. 37:35

    >> It is. It is the Yeah. [laughter] Um,

  1115. 37:38

    and you know, and and so there's also a

  1116. 37:39

    little bit of a question of like what

  1117. 37:41

    does your default aesthetic look like if

  1118. 37:43

    you don't specify? But then to Shane's

  1119. 37:45

    point, one of the things we always try

  1120. 37:47

    to get these models better at is

  1121. 37:49

    instruction follow so that like when you

  1122. 37:51

    want to get them to a different outcome

  1123. 37:53

    like you should be able to whether

  1124. 37:54

    that's through language or whether

  1125. 37:56

    that's through references because

  1126. 37:57

    language is sometimes too limiting. Um,

  1127. 38:00

    and so like these models continue to get

  1128. 38:02

    better at it but they so much at work.

  1129. 38:03

    Do do you feel pressure as a as a

  1130. 38:05

    product director to set the default for

  1131. 38:07

    the world like I mean [laughter]

  1132. 38:10

    >> kind of

  1133. 38:11

    >> maybe I should I don't know I haven't

  1134. 38:12

    thought about this

  1135. 38:14

    >> you know you know it's like someone has

  1136. 38:16

    to have a default the default has to

  1137. 38:18

    exist

  1138. 38:18

    >> actually I will say like we have thought

  1139. 38:20

    about this um and I I think one of the

  1140. 38:23

    so for example actually like if you look

  1141. 38:25

    at nanobanana generations we had like an

  1142. 38:27

    explosion of nanobanana infographics

  1143. 38:29

    when nanobanana pro came out

  1144. 38:31

    >> I tried it yeah

  1145. 38:32

    >> um yeah

  1146. 38:33

    I think Nurb's papers were like all you

  1147. 38:35

    know so so many had like infographics

  1148. 38:37

    generated. Can you run your uh

  1149. 38:39

    watermarking on it and see how many

  1150. 38:41

    >> uh we pro we probably could we have we

  1151. 38:44

    haven't done that but I saw so like my

  1152. 38:45

    Twitter was maybe this is just also like

  1153. 38:47

    the bias of my algorithm but they were

  1154. 38:49

    everywhere um and it was actually very

  1155. 38:51

    painful because um I think our default

  1156. 38:54

    aesthetic was a little bit too it was

  1157. 38:57

    too cluttered like I think that the the

  1158. 38:59

    model is like a bit of an overeager

  1159. 39:01

    student that just like learned you know

  1160. 39:03

    it was like oh I know all these like I

  1161. 39:05

    know all this information about this

  1162. 39:06

    concept let me like shove into the same

  1163. 39:09

    image. Japanese infographics 5x that

  1164. 39:12

    [laughter]

  1165. 39:13

    >> or maybe it was you know um but it just

  1166. 39:16

    and and

  1167. 39:17

    >> wait so same prompt same content if it's

  1168. 39:19

    in Japanese it's

  1169. 39:20

    >> density density

  1170. 39:21

    >> oh wow

  1171. 39:22

    >> because that's the style in Japan

  1172. 39:24

    >> yeah some like very you know bureaucrat

  1173. 39:27

    and [laughter] there's a famous word for

  1174. 39:29

    it yeah

  1175. 39:29

    >> no but we do do go through this process

  1176. 39:31

    with Omni we did it together right like

  1177. 39:32

    where like we had like a bunch of like

  1178. 39:34

    we like at the very end okay like this

  1179. 39:36

    is we did some tuning and like okay what

  1180. 39:38

    kind of style do we prefer right like

  1181. 39:40

    you know

  1182. 39:41

    >> is it more muted more saturated

  1183. 39:43

    >> we had a lot of saturation

  1184. 39:44

    >> yeah there was there were I think Nicole

  1185. 39:46

    just has PTSD so has forgotten about it

  1186. 39:48

    but she was very much involved in this

  1187. 39:50

    of like okay which which kind of color

  1188. 39:52

    palette do we basically prefer right and

  1189. 39:54

    it's you know it's it's it's not

  1190. 39:55

    something that like you have to make a a

  1191. 39:58

    trade-off there like uh

  1192. 40:00

    >> and and and it's because it ends up

  1193. 40:01

    being us right like actually it is true

  1194. 40:03

    like it it ends up being the modeling

  1195. 40:04

    teams and you could ask the question

  1196. 40:06

    legitimately of like are we the best

  1197. 40:07

    people to do that or should we actually

  1198. 40:10

    work with someone who like has a really

  1199. 40:12

    creative point of view and is more of

  1200. 40:14

    like you know an art director and like

  1201. 40:15

    has like and we kind of go back and

  1202. 40:17

    forth on this um

  1203. 40:19

    >> we have the trusted testers I'm on

  1204. 40:21

    >> we do we have trusted testers who give

  1205. 40:23

    us a lot of feedback and we take that

  1206. 40:24

    serious

  1207. 40:25

    >> very well organized by the way to have

  1208. 40:26

    these like weekly calls and stuff like

  1209. 40:27

    it's it's amazing

  1210. 40:29

    >> um Logan's team does a lot of that so

  1211. 40:31

    kudo kuda kudos kudos to Logan um who

  1212. 40:34

    couldn't be here today um and we have a

  1213. 40:36

    lot of people actually internally at

  1214. 40:38

    Google like Fulfur who give us like a

  1215. 40:40

    ton of No, no, no. Truly like who give

  1216. 40:42

    us a ton of feedback on like when we

  1217. 40:44

    when we release new checkpoints and like

  1218. 40:46

    sometimes it will be stuff that we like

  1219. 40:47

    don't see right like we would be like oh

  1220. 40:50

    yeah this optimization seems okay and

  1221. 40:52

    then they would come back what have you

  1222. 40:53

    done like you completely ruined my grass

  1223. 40:55

    you know because now the detail is all

  1224. 40:57

    blurry.

  1225. 40:57

    >> I think he just noticed not not a super

  1226. 40:59

    secret at this point but like that our

  1227. 41:01

    model tends to put rings wedding rings

  1228. 41:03

    on on on hand. That's yeah

  1229. 41:04

    >> very strange. I had never noticed that

  1230. 41:06

    but he's like he I just saw it and

  1231. 41:07

    there's a faux fur channel basically.

  1232. 41:09

    >> Uh he posted I was like why is there

  1233. 41:12

    wedding ring in every hand? I'm like

  1234. 41:13

    that's strange.

  1235. 41:14

    >> That sounds very common reward hacking.

  1236. 41:15

    >> Yeah. Yeah. Yeah. So but you know

  1237. 41:17

    something that we would not have we

  1238. 41:18

    would not have noticed necessarily while

  1239. 41:20

    while developing this right is an oral

  1240. 41:21

    artifact or

  1241. 41:22

    >> I I don't know you do have like a lot of

  1242. 41:25

    preference based and then you know you

  1243. 41:26

    may can prefer that sperious correlation

  1244. 41:29

    reward hacking it can happen like in

  1245. 41:31

    many weird ways. Yeah

  1246. 41:33

    >> it does. It is

  1247. 41:34

    >> uh this was related to another topic

  1248. 41:36

    that again I I try to use these

  1249. 41:38

    mainstage things as introductions or

  1250. 41:40

    ties in. Uh we have the eval track we

  1251. 41:43

    have character AI and YouTube talking

  1252. 41:45

    about how they evaluate videos. Um how

  1253. 41:48

    do you evaluate videos

  1254. 41:51

    >> apart from furer [laughter]

  1255. 41:53

    >> not everyone has a fauxur but also you

  1256. 41:55

    know I think there needs to be something

  1257. 41:56

    more quantitative

  1258. 41:57

    >> well I mean it's you improve Gemini

  1259. 42:00

    to improve the evaluation for video. Um

  1260. 42:03

    that's that's no no that's that's

  1261. 42:05

    definitely one way uh it's actually very

  1262. 42:07

    hard.

  1263. 42:08

    >> It's very hard. It's very hard um to get

  1264. 42:10

    like you know audators to evaluate

  1265. 42:13

    things in a video like including

  1266. 42:15

    especially things like aesthetics right

  1267. 42:17

    like that it's like there are some

  1268. 42:18

    things that are a little bit more

  1269. 42:20

    objective like especially when we talk

  1270. 42:22

    like let's say we talk about images and

  1271. 42:23

    we look at like infographics text

  1272. 42:24

    rendering that's actually fine right

  1273. 42:26

    because like you can kind of OCR things

  1274. 42:28

    out and then you can look at like okay

  1275. 42:31

    this letter is like messed up and then

  1276. 42:32

    the whole thing is actually useless

  1277. 42:33

    because if like literally if a letter is

  1278. 42:35

    off in render text you just can't use

  1279. 42:37

    that asset. Right. So th those things

  1280. 42:40

    are like a little bit more auto ratable.

  1281. 42:42

    Um from what we found we do rely a lot

  1282. 42:44

    on humans looking at things and so we do

  1283. 42:47

    do a lot of human evals. We do a lot of

  1284. 42:49

    human evals.

  1285. 42:50

    >> Do a lot of human ev and every time Jane

  1286. 42:53

    is like um and every time we have a new

  1287. 42:56

    model we like want to do more things and

  1288. 42:58

    we want to like gem in more capabilities

  1289. 43:00

    and then we have like more emails that

  1290. 43:02

    we have to run. Um, and then at some

  1291. 43:04

    point you do get two models that are

  1292. 43:07

    like kind of close to each other and

  1293. 43:10

    then like we literally make decisions

  1294. 43:11

    based on like looking at outputs side by

  1295. 43:13

    side. Sometimes like in a room like I've

  1296. 43:17

    been in rooms where there's like 10 of

  1297. 43:19

    us and we're just like looking at video

  1298. 43:21

    side by side and we're like do you

  1299. 43:22

    prefer this or do you prefer that? like

  1300. 43:24

    oh wow it's

  1301. 43:25

    >> I mean but it is it is genuinely very

  1302. 43:27

    complicated the more capabilities you

  1303. 43:28

    add like you know even just the one

  1304. 43:30

    capability but it's like almost AGI

  1305. 43:32

    complete capabilities like video editing

  1306. 43:33

    right like think about video editing as

  1307. 43:35

    a and like editing with audio and

  1308. 43:38

    >> my editor will be very happy to hear

  1309. 43:39

    this

  1310. 43:40

    >> edit the hardest problem in g media

  1311. 43:43

    >> I mean I don't know if it's the hardest

  1312. 43:44

    but it's definitely there right like uh

  1313. 43:47

    in terms of like complexity of of

  1314. 43:50

    evaluation like free form video editing

  1315. 43:53

    is you can do anything like

  1316. 43:55

    >> yes uh and like I I spent a lot of money

  1317. 43:57

    on that and it's very hard to tell me

  1318. 43:59

    >> like adding those we don't have like add

  1319. 44:01

    a sloth eval right like uh that we

  1320. 44:04

    >> well now we should

  1321. 44:05

    >> now we should yeah yeah yeah but like

  1322. 44:06

    things like that like it's it's it's not

  1323. 44:09

    that easy to track

  1324. 44:10

    >> I think I'm just surprised at the sample

  1325. 44:12

    size that you have right like to to test

  1326. 44:14

    the entire surface of your models you

  1327. 44:17

    still rely on a magnitude of hundreds

  1328. 44:20

    >> no no no so we do like yeah well we do

  1329. 44:22

    we do a ton of human evals on like on

  1330. 44:24

    like you know thousands of things. Um I

  1331. 44:27

    I think there's also like an element of

  1332. 44:29

    you know we can talk about things like

  1333. 44:30

    live experiments right like which which

  1334. 44:32

    is also where you get signal on like

  1335. 44:34

    like some of these more minute

  1336. 44:36

    differences at like much larger scale

  1337. 44:38

    then there's auto raers which is

  1338. 44:39

    definitely kind of a more it's a very

  1339. 44:42

    well defined space I think for LLMs much

  1340. 44:45

    more nent for media models and then like

  1341. 44:49

    sometimes you still do rely on human

  1342. 44:51

    judgment and we do rely on things like

  1343. 44:53

    feedback from people who just like have

  1344. 44:55

    a very owned like aesthetic and and

  1345. 44:59

    people who just like use these models in

  1346. 45:00

    their workflows dayto-day, right?

  1347. 45:02

    Because we could also like you could

  1348. 45:04

    have a model that does really well on

  1349. 45:05

    some slice of human evals, but then it

  1350. 45:07

    like really breaks a workflow for

  1351. 45:09

    somebody. And so this is why we do like

  1352. 45:10

    early access programs and we try to get

  1353. 45:12

    feedback and then we like try to

  1354. 45:13

    incorporate it before we release

  1355. 45:14

    something more broadly. I feel like

  1356. 45:16

    Shane had a hot take based on his

  1357. 45:19

    >> expression always when we were talking

  1358. 45:20

    about this

  1359. 45:21

    >> every kind of human sort of you know

  1360. 45:23

    work should be gradually kind of

  1361. 45:25

    amortized and then the interesting thing

  1362. 45:26

    is the video understanding especially

  1363. 45:29

    like against like AI gener like

  1364. 45:31

    detecting air stuff is extremely

  1365. 45:33

    interesting uh visual task

  1366. 45:35

    >> and then like some of it kind of

  1367. 45:37

    aesthetics or this kind of visual

  1368. 45:38

    quality but for some of the kind of

  1369. 45:40

    cases like semantically doesn't make

  1370. 45:42

    sense for example you're taking like

  1371. 45:44

    some like a famous scene from a movie

  1372. 45:46

    and try to sort of um construct that and

  1373. 45:49

    then if you kind of generate it uh it

  1374. 45:51

    can generate something there but at some

  1375. 45:53

    point some of the semantic information

  1376. 45:55

    doesn't make sense like it's actually

  1377. 45:57

    inconsistent. So can the AI actually

  1378. 46:00

    detect that? So when I evaluate the AI

  1379. 46:03

    videos like oh I feel I'm so smart you

  1380. 46:05

    know like like AI is still kind of

  1381. 46:08

    behind but we should make like a lot of

  1382. 46:10

    effort. I think the video understanding

  1383. 46:11

    is extremely uh important intelligence

  1384. 46:14

    task uh beyond just the pure aesthetics

  1385. 46:16

    or the preference. Um and yeah we we

  1386. 46:20

    should always try to amatize the human

  1387. 46:22

    >> human label. Yeah.

  1388. 46:24

    >> Yeah. Um, what data do you need? A lot

  1389. 46:29

    of people I talked to wanted to get in

  1390. 46:32

    front of you actually. Uh, they I mean

  1391. 46:34

    they want to be nice about it. They have

  1392. 46:36

    a lot of video data. They have gaming

  1393. 46:38

    data. They have real world video data.

  1394. 46:39

    They have images. They have labelers.

  1395. 46:42

    What do you want?

  1396. 46:44

    >> Are you like offering?

  1397. 46:46

    >> I'm just like this is your request for

  1398. 46:48

    like Okay. Okay. We get I'm sure you get

  1399. 46:50

    a lot of pitches, right? You get a lot

  1400. 46:52

    of people want to talk to you. what's

  1401. 46:54

    like I think actually it's the signal is

  1402. 46:57

    this problem this sorting out signal

  1403. 46:59

    from noise is the main problem so

  1404. 47:01

    creating a nice API of like okay if you

  1405. 47:04

    actually do a b and c we are interested

  1406. 47:06

    in that

  1407. 47:09

    >> um

  1408. 47:11

    loaded question there so uh I don't know

  1409. 47:13

    that there's like an easy like you know

  1410. 47:14

    did you do I think we we do already have

  1411. 47:16

    a lot of data I think it's it's

  1412. 47:19

    >> hard to talk about this

  1413. 47:21

    >> you know you want to talk about the

  1414. 47:22

    public I don't want to get you in

  1415. 47:23

    trouble Yeah,

  1416. 47:23

    >> but like I think

  1417. 47:24

    >> No, no. What I just want to say is like

  1418. 47:26

    hard to talk about this in a sort of you

  1419. 47:28

    know without trying to without I have to

  1420. 47:31

    think about the what I am revealing

  1421. 47:33

    about our project and what where we're

  1422. 47:35

    going. Um generally high quality data I

  1423. 47:38

    think maybe maybe let's just put it this

  1424. 47:39

    way right it's not not the secret

  1425. 47:40

    >> embodied I'm sorry

  1426. 47:42

    >> embodied data

  1427. 47:43

    >> I mean

  1428. 47:44

    >> yeah sure I mean we have sort of

  1429. 47:47

    announced I think publicly right that we

  1430. 47:49

    we have some sort of robotics

  1431. 47:50

    collaboration right like so I think it's

  1432. 47:51

    like a like or or but you because we

  1433. 47:55

    have a robotics team at GDM so you know

  1434. 47:57

    they're always interested in things like

  1435. 47:58

    that um I mean for Omni specifically I

  1436. 48:01

    think we're just quite interested just

  1437. 48:03

    high quality data right like you know it

  1438. 48:04

    it's not some sort of not necessarily

  1439. 48:07

    like oh random YouTube video but like

  1440. 48:09

    you know some a some more professional

  1441. 48:11

    shop things like that right the things

  1442. 48:12

    that those are those are things that

  1443. 48:15

    we're always on the lookout for like uh

  1444. 48:17

    and yeah

  1445. 48:18

    >> and I think for you know maybe this is

  1446. 48:20

    easier to some extent to answer for like

  1447. 48:23

    some of the agentic work as well like

  1448. 48:25

    like like actual kind of like what are

  1449. 48:28

    the tests that people are trying to do

  1450. 48:30

    right these things are actually kind of

  1451. 48:32

    difficult to manufacture if you're doing

  1452. 48:35

    it yourself or if you're like doing it

  1453. 48:36

    with a vendor, like what is the actual

  1454. 48:38

    like if you're creating a marketing

  1455. 48:39

    campaign, like what does that look like,

  1456. 48:41

    right? Like do do you start from here's

  1457. 48:44

    like a picture of my new product and

  1458. 48:46

    then I want to turn that into a video ad

  1459. 48:48

    and I want to turn that into a bunch of

  1460. 48:50

    assets that like fit fit all these

  1461. 48:51

    different ad formats that I need to push

  1462. 48:53

    onto various platforms to promote and

  1463. 48:55

    then like so you kind of go from this to

  1464. 48:58

    that and like what is that kind of

  1465. 48:59

    trajectory of tasks that you're that

  1466. 49:01

    you're like you know experiencing along

  1467. 49:03

    the way like that is really useful and

  1468. 49:06

    that is actually kind of difficult to

  1469. 49:08

    get right u because like we don't always

  1470. 49:11

    have the right firstparty surface where

  1471. 49:14

    people are actually doing some of these

  1472. 49:15

    things or like you might work with

  1473. 49:18

    someone who's a vendor but they don't

  1474. 49:19

    also don't have that product surface

  1475. 49:21

    right like like a lot of this kind of

  1476. 49:23

    information lives in the places where

  1477. 49:24

    people are doing these tasks and so

  1478. 49:26

    that's kind of difficult to get like if

  1479. 49:27

    anyone's figured that out you should

  1480. 49:29

    reach out to us

  1481. 49:31

    >> every channel of thought yeah every

  1482. 49:33

    [laughter] thought

  1483. 49:33

    >> every thought yeah and maybe the data

  1484. 49:36

    the Chinese lab is using

  1485. 49:38

    >> yes yeah uh you know

  1486. 49:41

    as a media person myself, right? Like

  1487. 49:43

    there's so many podcasters and people in

  1488. 49:46

    in marketing departments and all these

  1489. 49:47

    like they would happy to be your data

  1490. 49:49

    like you know just like put a BCI on my

  1491. 49:51

    head

  1492. 49:52

    >> and podcast [laughter] watch my things

  1493. 49:54

    uh because like you know there's just

  1494. 49:56

    endless amount of work to do like

  1495. 49:58

    there's so much work and this is all

  1496. 50:00

    like this needs to somewhat be commodity

  1497. 50:02

    like obviously you can be an art like an

  1498. 50:05

    artisan like you can be Hollywood for

  1499. 50:07

    like the really high quality stuff but

  1500. 50:08

    actually a lot of work is commodity and

  1501. 50:10

    like should be modelable and we want you

  1502. 50:12

    to do

  1503. 50:13

    >> [laughter]

  1504. 50:13

    >> And but we we want the high quality to

  1505. 50:16

    Demi's point right like we do want we

  1506. 50:17

    want the high quality.

  1507. 50:18

    >> We want commodity. Yeah. Yes. Yes. You

  1508. 50:20

    want on both sides.

  1509. 50:21

    >> Um I I just

  1510. 50:24

    >> Thank you for the solicitation.

  1511. 50:26

    [laughter]

  1512. 50:27

    >> Uh I you know we we we also I also added

  1513. 50:30

    a data quality track. I I think that uh

  1514. 50:32

    people want to understand like what uh

  1515. 50:35

    at AI like how to raise the bar, right?

  1516. 50:38

    like like the and a lot of it is just

  1517. 50:40

    educating the market and educating

  1518. 50:42

    researchers and engineers and founders

  1519. 50:44

    on like this is where we're going a lot

  1520. 50:47

    of this is stop doing that do this do

  1521. 50:49

    this instead and I'm like people will

  1522. 50:51

    listen

  1523. 50:53

    yeah I don't know uh to that extent you

  1524. 50:55

    know

  1525. 50:56

    >> but I think to that to that point like

  1526. 50:57

    there's a lot of again just like craft

  1527. 50:59

    that goes into this right and there's a

  1528. 51:00

    lot of process like you even to the

  1529. 51:02

    marketing campaign example you don't

  1530. 51:04

    create that in like five minutes right

  1531. 51:05

    you like go you go through a process and

  1532. 51:07

    you iterate and you like pick something

  1533. 51:10

    over something else because you liked it

  1534. 51:12

    for whatever reason like maybe the eye

  1535. 51:13

    gaze was correct right like we just we

  1536. 51:15

    don't know these things right because

  1537. 51:17

    none of us are marketing directors and

  1538. 51:19

    like the models don't know these things

  1539. 51:21

    >> I even kind of say this for the natural

  1540. 51:22

    like a language as well like I I always

  1541. 51:24

    kind of say 99% of information is inside

  1542. 51:27

    people you can only extract it through

  1543. 51:29

    active dialogue and befriending them so

  1544. 51:31

    most of the stuff on the internet is

  1545. 51:33

    like sort of the outcome the output of

  1546. 51:35

    that yes but you know what are what are

  1547. 51:37

    all the trajectories you know how did

  1548. 51:38

    this person have this inspiration to

  1549. 51:40

    write this paper

  1550. 51:41

    >> what is the starting point what is the

  1551. 51:42

    inspiration what are the dialogue that

  1552. 51:44

    sparked it those kind of stuff is kind

  1553. 51:45

    of inside people so even you know those

  1554. 51:48

    kind of like even the language space is

  1555. 51:49

    kind of that I think the creative is

  1556. 51:51

    kind of similar as well there's a lot of

  1557. 51:52

    dark knowledge

  1558. 51:53

    >> yeah it's like when you write a novel

  1559. 51:54

    right like a novel speaks to you because

  1560. 51:57

    like usually there's some sort of like a

  1561. 51:59

    personal connection that you feel to

  1562. 52:00

    like the story or the trajectory or the

  1563. 52:02

    characters right like if you read most

  1564. 52:04

    of the stuff that's written by LLM's

  1565. 52:06

    today like it's, you know, it's it's it

  1566. 52:09

    starts it falls into these like default

  1567. 52:11

    par patterns and like the language

  1568. 52:13

    starts to feel really similar and all

  1569. 52:14

    the descriptions sound really similar.

  1570. 52:16

    You can kind of like quickly read it as

  1571. 52:18

    like, oh, this is not that interesting

  1572. 52:19

    because like I can't connect to it,

  1573. 52:21

    right? Um, and again, that's that's kind

  1574. 52:23

    of like a human expertise.

  1575. 52:26

    >> One nice thing recently is the Google

  1576. 52:28

    Cloud and the Google Deep Mind are kind

  1577. 52:29

    of starting to invest a lot more in the

  1578. 52:31

    FTEEs for the product engineers. And I

  1579. 52:33

    also kind of saw some uh recruiting for

  1580. 52:35

    the creative you know gem media kind of

  1581. 52:37

    space as well. So I think those are kind

  1582. 52:39

    of really the effort because we we kind

  1583. 52:41

    of feel that you know what we can kind

  1584. 52:42

    of do with a lot of public data there's

  1585. 52:44

    limits but really you know partnering

  1586. 52:46

    with that we can provide kind of better

  1587. 52:47

    models and products and yeah we kind of

  1588. 52:49

    feedback

  1589. 52:50

    >> uh we have an FD track here for the

  1590. 52:52

    first time every lab is announcing it.

  1591. 52:54

    It's it's crazy. Um, one thing I'm

  1592. 52:56

    actually very keen on doing and I push I

  1593. 52:59

    push for this at Cognition as well is to

  1594. 53:01

    turn the FDES not just into sales and

  1595. 53:04

    solutions but also to EVAL's uh eval

  1596. 53:07

    workers.

  1597. 53:08

    >> FD is not the sales FD is way way bigger

  1598. 53:11

    than that. How do you frame FDs then?

  1599. 53:14

    Because [laughter] I do think about it

  1600. 53:16

    as sales like you're you know the more

  1601. 53:18

    the more you customize the solution for

  1602. 53:20

    >> so I define post training as anything

  1603. 53:22

    between the pre-training and the final

  1604. 53:25

    user experience anything anything is a

  1605. 53:27

    post training

  1606. 53:28

    >> and to me when I first sort of you know

  1607. 53:30

    learned a lot about I mean FD kind of I

  1608. 53:32

    guess originally you know came from like

  1609. 53:33

    path here and then that so I guess the

  1610. 53:36

    kind of history is different but yeah I

  1611. 53:38

    think the key is really that um you know

  1612. 53:40

    the key is like not only to kind of work

  1613. 53:42

    uh with them and ensure that they kind

  1614. 53:43

    of know how to

  1615. 53:44

    but also to sort of code like derive

  1616. 53:48

    kind of insights that can basically kind

  1617. 53:49

    of help both parties. They can put the

  1618. 53:51

    like a lot of harness how they use the

  1619. 53:53

    model. We can improve like very

  1620. 53:54

    upstream. So how to get the customer

  1621. 53:56

    feedback to the modeling I feel is the

  1622. 53:59

    kind of more the the role I I kind of

  1623. 54:01

    want for the fds. Yeah.

  1624. 54:03

    >> Yeah. Yeah. and and even for sorry just

  1625. 54:05

    on that like if you want to talk to us

  1626. 54:06

    or at least me um I I'm not going to

  1627. 54:10

    offer up your time um but I it's really

  1628. 54:13

    helpful for us to actually talk to

  1629. 54:14

    people who are using our models and like

  1630. 54:16

    understand where they're struggling uh

  1631. 54:18

    because again that just like it's it's

  1632. 54:20

    the real world task that you're actually

  1633. 54:22

    trying to use them for right like I will

  1634. 54:23

    talk to people who do kind of interior

  1635. 54:26

    inter interior design with some of our

  1636. 54:28

    image models um you know and they will

  1637. 54:31

    say hey like I really want to take this

  1638. 54:33

    pattern pattern, but then I want to

  1639. 54:34

    scale it across like 10 different ruck

  1640. 54:36

    sizes and sometimes I have like a very

  1641. 54:38

    custom ruck size and then the model

  1642. 54:40

    fails at like replicating the pattern

  1643. 54:42

    the same way or you know I want to do a

  1644. 54:44

    try on for these earrings and then the

  1645. 54:47

    earrings have a certain size and then

  1646. 54:48

    like my head has a certain size like it

  1647. 54:50

    has to make sense if you're actually

  1648. 54:52

    trying to try things on and like the

  1649. 54:54

    models kind of fail at a bunch of these

  1650. 54:56

    things that like actually happen in the

  1651. 54:58

    real world, right? Um and so that that's

  1652. 55:00

    like useful for us because for some of

  1653. 55:02

    these things like we don't think about

  1654. 55:04

    because we don't you know we don't use

  1655. 55:05

    the models for those tasks

  1656. 55:06

    >> or like um you know I think to your

  1657. 55:08

    point about ad campaigns or whatever

  1658. 55:10

    like people have like notions of brand

  1659. 55:11

    languages or whatever like which is

  1660. 55:13

    >> yes

  1661. 55:13

    >> like a a bunch of images or PDFs saying

  1662. 55:16

    things you know it's a pretty kind of

  1663. 55:18

    you know ambiguous question as well what

  1664. 55:20

    is the IKEA brand language you know is

  1665. 55:22

    it is it blue and yellow I mean that's

  1666. 55:25

    that's not a very like

  1667. 55:26

    >> but like what shade of blue you know.

  1668. 55:27

    >> Yeah. Yeah. Yeah. So there there's like,

  1669. 55:28

    you know, and the brands are pretty

  1670. 55:29

    spec, you know, pretty, you know, like

  1671. 55:30

    they they do care about the shade of

  1672. 55:32

    blue. It's not shouldn't just be a

  1673. 55:33

    random blue and a random yellow. That's

  1674. 55:35

    not going to be IKEA, right? I'm just

  1675. 55:36

    thinking about an example. But like this

  1676. 55:38

    is the kind of stuff that, you know,

  1677. 55:39

    it's not necessarily part of our like,

  1678. 55:41

    you know, developing frontier models

  1679. 55:42

    kind of, you know, necessarily mandate,

  1680. 55:44

    but it's something that we do want to we

  1681. 55:45

    do want to fundamentally like build

  1682. 55:47

    products that people will use to solve

  1683. 55:50

    concrete tasks, not just not just

  1684. 55:52

    research artifacts, right? So I think

  1685. 55:53

    it's useful to understand what people do

  1686. 55:55

    care about. Uh well, I'm sure a lot of

  1687. 55:58

    people are very grateful for your work

  1688. 56:00

    and there's a lot more to do that you've

  1689. 56:01

    made so much progress over the last like

  1690. 56:04

    even just couple years of like Nano

  1691. 56:06

    Banana and Theo and Omni and uh I don't

  1692. 56:09

    know what else you got cooking but we're

  1693. 56:10

    very excited like you this is one of

  1694. 56:12

    those things where like I was very

  1695. 56:14

    disappointed you know when Sora shut

  1696. 56:16

    down and and I think like there needs to

  1697. 56:18

    be more general exploration of uh you

  1698. 56:21

    know generative models and not just you

  1699. 56:24

    know coding. [laughter] I think I think

  1700. 56:25

    that is

  1701. 56:26

    >> we obviously like this.

  1702. 56:27

    >> We love coding. Love coding and and uh

  1703. 56:30

    yes uh but thank you so much for your

  1704. 56:32

    time. Uh it's been a real pleasure and I

  1705. 56:33

    can't wait to see what this looks like

  1706. 56:34

    next.

  1707. 56:35

    >> Thank you for having us. Great question.

  1708. 56:36

    >> Thank you everyone. [applause]