← All AI Engineer talks

AI Engineer World's Fair 2025

Veo 3 for developers

About this talk

Google DeepMind developer-relations engineering lead Paige Bailey demonstrates Veo 3 and related generative-media models, contrasting Veo 2 editing and API workflows with video generation that incorporates audio, dialogue, lip synchronization, and image-to-video animation. She discusses Imagen 4, Lyria 2, Music AI Sandbox, MusicFX, SynthID watermarking, earlier video models, and a practical Gemini-orchestrated workflow that segments prompts and adds generated music.

Chapters

  1. 0:00Introduction and Google DeepMind's generative-media models
  2. 1:48Veo 2 APIs, Flow, editing, and synchronized dialogue
  3. 8:25SynthID, image generation, and AI music tools
  4. 11:57Earlier video models, image-to-video, and richer prompting
  5. 18:25Gemini-assisted video reconstruction and Veo 3 takeaways

Talk transcript

  1. 0:00

    [on-hold electronic music] Thank you so much for having me.

  2. 0:17

    Thank you all for being here today and wanting to learn more about generative media. Um, I'm going to keep this pretty quick. There's a lot to show, especially for some of the new features that we've released in Veo 3.

  3. 0:27

    Um, but there's also a lot, uh, to discuss in the, in the frame of how does this revolutionize the way that people build, um, things? How do people create ads?

  4. 0:36

    Uh, and how do people replicate, uh, some of the experiences that you might see every day? Um, so as mentioned, I'm Paige. I am the eng lead for our DevRel team at Google DeepMind.

  5. 0:45

    Um, but I'm here today on behalf of our generative media team, um, who are all brilliant. Uh, they are wonderful. Uh, Tom Hume, Dmitri Erhart, many of them, uh, could not be here today.

  6. 0:57

    Um, but, uh, but this is all their work. So I just want to send like a thank you to the heavens for, for everything that they've been building. Um, today, we're going to be talking about three different models.

  7. 1:07

    So Veo 3, which is our new video and audio generation model, Imagen 4, which can generate these static images, and also Lyria 2, which is a music generation model.

  8. 1:17

    Um, and stay tuned. There will be more of all of the above coming shortly. Um, Veo is, uh, kind of magical in the sense that, you know, you can create videos of things that you've never seen before, and we've saw-- we've seen some of these examples already.

  9. 1:31

    But it also has the potential, um, to really revolutionize everything that we build and create as humans. And I loved this quote from Andrej Karpathy, um, just recently, that video has the potential to be an incredible surface for communication, um, but also for education and for human creativity.

  10. 1:48

    So all of these things are designed with that in mind. Um, Veo 2, uh, just want to touch on it briefly before we get into the Veo 3 capabilities.

  11. 1:57

    Um, we released a whole bunch of additional things around creative control, um, just recently. So think on the order of about a month ago, um, or not even that.

  12. 2:06

    Um, so things like reference-powered videos, um, outpainting, the ability to add and remove objects, character control and consistency, which I think you'll, uh, y'all were talking about just a little while ago, um, and also the ability to interpolate across first and last frames.

  13. 2:20

    And let's take a look at what that means, um, because for people who are not necessarily filmmakers, um, it's much, much easier to show and not tell. Um, so reference-powered videos are things like you have, uh, you have a person, you have an environment, um, and you're able to kind of put one within the other or compose

  14. 2:38

    them together into something that, um, feels very, uh, very stylized, but also like really, really well, um, well-crafted. Um, so here you can see a reference-powered video, um, with Veo 2.

  15. 2:50

    This is available via the API and via some of our tools like, uh, like Flow today.

  16. 2:56

    Uh, and another example of reference-powered video, so a really, really cute, uh, little monster in a variety of environments that you just control by describing. Um, it's performing pretty well on benchmarks.

  17. 3:09

    So you can see here, um, the green is, uh, Veo. So, um, uh, compared to things like Runway, Gen-4, and Kling, um, the more green, the better. Um, and for reference-powered video, most human raters, um, uh, uh, selected for Veo for some of these side-by-side comparisons.

  18. 3:27

    You can also match styles, um, so upload a reference image and then have, uh, different styles composed together.

  19. 3:36

    Another example of styles being preserved, and then also camera controls the same that you might have if you were a filmmaker. So things like being able to move back, move right, um, rotate up, zoom in, um, and to be able to precisely control all of these camera movements, again, just via natural language and through some of the

  20. 3:54

    tools that are available in the APIs. Um, when I saw all of this, I was blown away because I don't think that we have nearly enough code samples demonstrating some of these capabilities.

  21. 4:04

    Um, but these are all things that you can do with the Veo 2 models today, um, with the APIs. We also have the ability to do outpainting. This was important for, uh, a recent project with the sphere around "Wizard of Oz."

  22. 4:16

    So being able to take a sc-- uh, like a sc- a scene or, or an individual frame of a video, um, and imagine what the rest of the scene might look like.

  23. 4:26

    So even if you only have a view into a small portion, um, being able to create something that, uh, that looks real, um, or that looks consistent, uh, across the outpainting.

  24. 4:39

    Um, adding objects or removing objects from scenes. So you can see here a few examples as well. Um, and again, all of these available for you to test and to try today.

  25. 4:49

    Um, these are all, these are all things that exist, um, uh, that our, that our research team has kind of gotten into the, to the API designs. Um, some more examples of removing objects.

  26. 5:01

    Character control is quite nice. You might have seen some of these demonstrated into your favorite products, um, for, um, kind of controlling mouth movements, controlling reference face movements for particular, um, for particular characters.

  27. 5:14

    We also give the ability to add a script, um, and to add kind of a voice tone and to have the character, um, map the lips to producing that sound and producing it, um, in a way that feels consistent with the, with the location.

  28. 5:29

    Um, these are some of the, the motion examples, um, with Veo 2 and with Veo 3. So being able to have an input image and controlling them or changing the design, um, across, across the scene.

  29. 5:45

    Um, more benchmarks, and then the first and last frame. Um, so you can have an input image and an output image, and Veo is able to interpolate across them, um, to, to kind of make those, make those images, um, stitch together into a video.

  30. 6:03

    And those are just another couple of examples. I feel like generative media presentations are very gratifying. Like, these are certainly the most beautiful things, um, that we, uh, that we get to see at developer conferences.

  31. 6:16

    So Veo 3, everything that you just saw, Veo 2. Veo 3, right? Like, like blown away. So Veo 3 is, um, video but coupled together with audio. And so all of the tokens compose together natively, um, not audio being pulled in as a tool, um, but the model actually able to compose together all of these tokens across

  32. 6:37

    multiple modalities. This is similar to what you see with Gemini's native audio output. Um, in addition to being able to output text and code, you can also output images, edit images, and also edit audio, compose audio, et cetera.

  33. 6:51

    Um, so Veo 3, our latest state-of-the-art, uh, video generation model. Um, it's, uh, you know, it, it has these things around prompt adherence and native audio generation. Um, but again, so much cooler to show, um, and not tell.

  34. 7:05

    The little llama. Then another one that... [upbeat music] So, so interestingly, you're able to do not just, uh, not just, uh, background noises, but also things like, uh, things like music.

  35. 7:40

    Including s- very, very subtle sounds and the like.

  36. 7:47

    So let's, uh, let's go to the next.

  37. 7:51

    There we go. So, so this is hard, right? Like, like, uh, it looks very cool. It's very hard to capture the nuances in an, of an input prompt. Um, and it's also been really historically very hard to, to preserve visual consistency.

  38. 8:05

    Um, so characters often, like, jump from one, um, um, from one frame to another. There might be backgrounds, uh, and then suddenly walls disappear and you're able to see behind them.

  39. 8:15

    Um, this is one of the reasons why Veo 3, uh, feels like a leap forward, is because the stylistic consistency and then also the contextual consistency, um, is much, much better.

  40. 8:25

    Um, built on years of research, um, so things like GQN, um, Walt, et cetera. Um, and it has responsibility at its core. So you can see little, uh, human visible watermarks as well as SynthID watermarks, um, for synthetically generated images and video.

  41. 8:41

    We've also been partnering really closely with many, many, many, uh, artists along the way. So, uh, Darren Aronofsky, um, also, uh, musicians for our Lyria models, artists for the Imagen models, um, and we'll take a look at a couple of these as well.

  42. 8:58

    Um, so Imagen is image generation. Um, uh, you know, able to kind of preserve realism, um, everything from humans to whales, uh, you know, uh, cute puppies. Um, I've heard that the mo- the more cute puppies that you have in a presentation, the better it always is.

  43. 9:16

    Um, and then also being able to preserve detail across all of these images as well, um, including diverse styles, uh, and even things like typography. So I love these stamps of Alamo Square and the Mission.

  44. 9:28

    I really, really wish that we just had these as, um, uh, like swag ideas that, uh, uh, stickers for laptops or stickers just in general. Um, and, uh, another example of an artist that the team has been really closely collaborating with, Ross Lovegrove, um, on, uh, on some of his, uh, some of his designs.

  45. 9:51

    So, uh, Lyria 2, um, also very exciting. It's high-fidelity music and professional-grade audio. Um, they also gives you a very, very granular creative control, so the ability to steer the inputs and outputs and to steer the tones and the styles of the music, um, along the way.

  46. 10:09

    Um, Music AI Sandbox is one of the products that's been created as a visual for this. Um, if folks are familiar with Ableton, um, or things like it, uh, this, this probably looks, uh, looks very similar.

  47. 10:21

    Um, and then there's also MusicFX, which is a project from our labs team, uh, that allows you to, to kind of, uh, compose together beats just via natural language.

  48. 10:31

    Um, and we use that for the demo later on today. Uh, Lyria Real-Time, uh, has also been a deep collaboration with many musicians, um, both Jacob Collier, uh, who's a legend, and, uh, Toro y Moi.

  49. 10:44

    Um, and me circa, like, uh, me circa college was, like, blown away that, uh, Toro y Moi was at, uh, Google I/O. Um, huge fan.

  50. 10:56

    I do think that a big part of teaching music is giving people a chance to play music and play with music.

  51. 11:02

    Yeah.

  52. 11:03

    However, people don't have access to the whole of music from day one.

  53. 11:07

    Yeah.

  54. 11:07

    What I think this offers an interesting perspective into is the whole of music, mathematics, physics, history, geography, the human body, language, spelling, syntax. One thing I've come to realize is that a lot of the same forces that make music work are the forces that make life work.

  55. 11:27

    Was that Veo 3?

  56. 11:28

    Oh, that was not Veo 3. [audience laughing] Some parts, but it's hard to tell, right? Like it's-

  57. 11:34

    The veins were like-

  58. 11:35

    Yeah, yeah. Yeah, you know. Parts of it, parts of it, um, might have included, uh, visuals generated by Veo 3 though. Um, so Lyria, again, built in collaboration with the creative industry, um, not outside it.

  59. 11:47

    Uh, and then also incorporates many of these techniques like SynthID to make sure that you have some sort of digital watermarking for the, for the assets themselves. Um, so now we're gonna get into it.

  60. 11:57

    Uh, these are some examples that I thought might be fun to share to show just how far we've come in the last couple of years. Because I think being here, we get very sucked into the Bay Area bubble, and we don't really kind of take a step back and appreciate how far, um, you know, the world has

  61. 12:12

    changed in just a matter of months. Um, so this is an example of one of the, one of the papers that was produced around 2023. Um, so released in 2023, research happening around 2022, um, for text to video.

  62. 12:25

    A raccoon wearing a black jacket dancing in slow motion in front of the pyramids. So just have that in your brain when you see it. Um, this is Walt, uh, circa 2023.

  63. 12:35

    Um, so very, very choppy, like really, really hard to have, um, even-- Like, that's not even eight seconds worth of a frame. Um, LTX video from 2024. This is one of the ones that were available on Hugging Face.

  64. 12:49

    Kling 2.0, which I did with Fal. Um, uh, heck yeah. Yep. Yeah. So, uh, which I did with Fal, um, released in 2025. Uh, and then Veo 2, 2024.

  65. 13:00

    Um, a very cute little raccoon, but I'm not sure how well he's dancing. Um, but this is just kind of a splattering of how the world has changed in the space, um, of just a couple of years.

  66. 13:10

    And then when you put the same prompt, um, through Veo 3, um, you get this. [whooshing]

  67. 13:20

    Very stylish raccoon. Yeah. There we go. So image to video, um, transforming static images into dynamic video content. So you can see here an image, um, of a woman and her puppy.

  68. 13:32

    Um, uh, a woman I, um, can only assume is in Texas, um, walking very slowly forward on the way to a gunfight. Um, and also different, uh, stylized images of a person in a single frame, um, being applied to, to different scenarios.

  69. 13:48

    So running towards the camera, um, tractor beam taking off, being lifted into the sky, again, all steered via natural language. Um, prompt rewriting is something that we've also released in Veo 3.

  70. 13:59

    So the ability to take that very, very simple, um, sentence before of a raccoon wearing a black jacket dancing in slow motion in front of the pyramids, um, and turn it into something a bit more fully formed, um, uh, that the, that Veo is, is much better able or better equipped to understand.

  71. 14:17

    Um, and so have that in your brain, um, as well as the, the concept of sound generation, so both music, sound effects, and background noises. Um, and we'll take a look at what that simple prompt, um, is now. [whooshing]

  72. 14:37

    Wow. I still, like, get blown away about how much detail there is, like being-- You can almost see the reflections in the eyes, um, for the, for the things as they, as they walk forward.

  73. 14:49

    Um, and then this is from our team in Paris. Uh, and they are very, very excited to have it shared today. Um, does anybody know [REDACTED:origin] in the audience?

  74. 15:00

    Okay, amazing. The rest of us will have a translation in a second. Um. [upbeat music]

  75. 15:09

    It's like Daft Punk. I can't believe this new Veo model. It is amazing.

  76. 15:18

    Oh, yeah. [laughs] [shouting]

  77. 15:18

    Yeah. [laughs]

  78. 15:49

    You can see the hair cut at the top, like, floating around. And for anyone who was curious, the translation is, um, so Veo means I see, I see, I see in Spanish.

  79. 16:09

    Um, uh, and then the guy responded with, "Yeah, I don't understand what you're saying because I'm [REDACTED:origin], actually." Um, artificial intelligence, artificial intelligence. Um, and then, uh, also the-- that's the interview.

  80. 16:22

    Um, and, uh, "Boss, the humans are about to create AGI. Should we contact them?" Um, "They're not ready yet." Uh, so amazing.

  81. 16:32

    Cool. Or actually, I'm gonna zoom to the next, zoom to the next one. Uh, so how do you access it? Big question. Um, right now we have a few different ways to access the Veo 3 models.

  82. 16:44

    One is through the Google AI Ultra plan, um, which is available in many different countries, um, including the UK just recently. Um, the Google AI Pro, uh, subscribers, so being able to access via the Gemini mobile app for a limited number of uses.

  83. 17:00

    And then also Veo 3 is available in private preview currently in Vertex AI. Veo 2 also available via Vertex AI for the Gemini APIs. Um, hoping to bring them to AI Studio.

  84. 17:10

    Um, crossing fingers, but, but we'll see. Um, and then you can also fill out a form for early access if you would like it. QR code coming shortly, so take a picture of that if you would like, um, with the form to, to go, uh, to go submit and test it out.

  85. 17:26

    Um, this is also, while you're looking, a code sample of how easy it is to use, um, with, uh, just your output bucket where you would like the, the video to be deposited, um, if you have an input image that you want to use as a starter, um, some things around aspect ratios, um, uh, and the like.

  86. 17:44

    So, um, toggling between different models or being able to specify some of these is just a handful of lines of code, which is pretty magical. Um, so this was intended to be a live demo.

  87. 17:55

    I'm not sure if I can tempt the demo gods, and also I am probably, like, close to being over time. Um, but I wanted to see how well I could replicate, uh, a commercial, um, with Veo that seemed pretty simple.

  88. 18:08

    Um, so the commercial is this one which took my name.

  89. 18:13

    Hey, my name's Paige, and what makes the Chick-fil-A chicken sandwich-

  90. 18:15

    Took my name. That's not me. [laughs]

  91. 18:16

    ... original to me is the crispiness of the breading and the tenderness of the filet. It's tasty, it's warm, it's total satisfaction.

  92. 18:25

    Hey, my name's Paige, and what makes the Chick-

  93. 18:27

    So that was the, that was the input video. Um, the process for replicating it with Veo 2 is you give Gemini the original video, have it create a really, really detailed plan, um, segmented into prompts be-- to like handle the eight-second limitation, um, used video-- uh, MusicFX to create the, the background track, which was a combination of

  94. 18:46

    Down Home Farm and Slow Guitar. Um, and then, uh, put it all in Camtasia, stitch it together with transitions. Whole process, uh, was relatively quick, but it also took a lot of thinking and a lot of work to get to that final assembly stage.

  95. 19:00

    Uh, and it looked like this. So this is, uh... [gentle music]

  96. 19:07

    Hey, my name's Paige, and what makes a Chick-fil-A Chicken Sandwich original to me is the crispiness of the breading and the tenderness of the filet. It's tasty. It's warm.

  97. 19:17

    It's total satisfaction.

  98. 19:20

    And that was, uh, again, like using, uh, using Veo 2, using Gemini text-to-speech, um, and stitching it all together myself. And I actually like that one better than the, the original commercial, but your mileage may vary.

  99. 19:33

    Process with Veo 3, um, you have the original video, you generate the description, you give it to Veo 3, and you see how well it does. Um, uh, this took, uh, just the span of like submitting the prompt to Veo 3's inflow.

  100. 19:46

    Hey, my name's Paige, and what makes the Chick-fil-A Chicken Sandwich original to me is the crispiness of the breading and the tenderness of the filet.

  101. 19:54

    So again, one prompt, and it was able to produce this.

  102. 19:59

    Hey, my name's Paige, and what makes the Chick-fil-A Chicken Sandwich original to me is the crispiness of the breading and the tenderness of the filet.

  103. 20:08

    Incredible. Um, so takeaways, Veo 3, pretty magical. We're committed to expanding it, expanding access as quick as we can, um, and around, uh, adding controls around steerability. Um, and thank you so much.

  104. 20:24

    That is not me actually waving at the camera. That is a static photo of me that has been Veo 3 animated. Excellent. Thank you. [audience applauding] [upbeat music]