← All AI Engineer talks

AI Engineer World's Fair 2025

Veo 3 for developers

Read the talk

Veo 3 for developers: from creative controls to audiovisual generation

A tour of reference images, scene editing, native audio and music tools leads to a practical comparison: rebuilding a commercial with separate generators versus one Veo 3 prompt.

From a talk by Paige Bailey

Before you start: Basic familiarity with API requests and Python is helpful for the code example; no filmmaking or music-production background is required.

What should it take to create an ad?

How do you create an advertisement—or recreate an experience you see every day—when the interface is a description rather than a production timeline? Paige Bailey, introducing herself as engineering lead for Google DeepMind’s DevRel team, approaches that question through the generative media team’s models. Veo 3, Imagen 4 and Lyria 2 cover video with audio, static images and music, respectively.

The broader possibility is to make video a useful medium for expressing ideas, not merely a finished entertainment product. Bailey invokes Andrej Karpathy’s framing of video as a surface for communication, education and human creativity. Before introducing native audiovisual generation, she starts with the creative controls in Veo 2: the tools for deciding what a scene contains and how it moves.

Slide quoting Andrej Karpathy on video, diagrams and animation as media for communication and creativity.
Video as a surface for AI–human communication and human creativity.
0:270:36
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:27 · section reference included

Keep the character, change the environment

Reference-powered video separates what you want to preserve from what you want to change. Supply a person and an environment as references, then compose the person into that setting. Bailey describes this as a way to produce a stylized, deliberately composed result, and says reference-powered generation was available through APIs and Flow at presentation time.

The little-monster example makes the operation concrete. A purple-and-yellow reference character appears in several environments, with descriptions controlling where it goes. The reference supplies a recurring subject; the prompt supplies a changing setting. The slide places the reference above four results, making character consistency across scenes directly comparable.

Veo 2 reference-powered video slide showing a purple and yellow monster reference above four scenes featuring the character.
A reference character appears in four different environments.

Bailey reports that most human raters preferred Veo in some reference-video side-by-side comparisons with Runway Gen-4 and Kling. The presentation does not establish a numerical score or a complete rating protocol here. The useful distinction is that this comparison concerns reference-powered generation: producing an appealing clip while following supplied visual references.

References can also guide style. Upload an image to preserve its visual treatment or compose styles together, then direct the camera through natural language and API controls. Bailey names moving back, moving right, rotating up and zooming in. These are separate creative decisions: a reference constrains appearance, while a camera instruction constrains how the viewer encounters the scene. She also notes a practical developer gap: too few code samples demonstrating the available controls.

1:572:06
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

1:57 · section reference included

Extend a scene, edit its contents, direct its motion

Outpainting starts with a limited view and generates what could exist beyond its boundaries. Bailey connects it to the Sphere’s Wizard of Oz project: take a scene or an individual video frame and imagine the surrounding world while keeping the extension visually consistent with the visible portion. The problem is not simply making the picture larger; it is making the newly visible area belong to the same scene.

Object editing changes the contents rather than the boundaries. Bailey shows adding objects and removing objects, including further removal examples, and describes these Veo 2 capabilities as available to test through APIs. Together, outpainting and object editing let a creator revise an existing scene instead of requesting a completely new composition.

Character controls move from scene composition to performance. Reference facial and mouth movements can guide a character’s animation. A script and voice tone can guide lip movements aligned with the sound and appropriate to the location. Bailey then shows input-image motion and design-change examples across Veo 2 and Veo 3: the starting image anchors the appearance while the generated sequence changes its motion or presentation.

First-and-last-frame generation provides a different constraint: specify the beginning and ending images, then have Veo generate the transition between them. After another benchmark slide, Bailey shows endpoint examples stitched into continuous video. These controls address distinct production needs:

ControlWhat the creator specifiesWhat generation supplies
Character performanceFace motion, script, voice toneA corresponding animated performance
Input-image motionStarting appearance and actionMovement from that visual starting point
First and last framesBoth visual endpointsThe intervening sequence

The endpoint approach is especially concrete: the destination is already an image, rather than something the model must infer entirely from prose.

4:044:16
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

4:04 · section reference included

Generate sound and pictures together

Veo 3’s central distinction is native audiovisual generation. Bailey describes the model as composing tokens across video and audio together, rather than calling a separate tool to add sound afterward. She compares this with Gemini’s native audio and broader multimodal capabilities: output need not stop at text or code.

Prompt adherence matters across both modalities. The demonstrations begin with a little llama, then show that generated sound can include music as well as background noise and very subtle effects. The target is therefore more than a moving picture with an arbitrary soundtrack: it is a scene whose visual events and audible details are generated together.

That coherence is difficult even before adding sound. A prompt may contain nuances that generation misses; a character may jump between frames; a wall may disappear and expose a background that should remain hidden. Bailey describes Veo 3 as improving both stylistic consistency and contextual consistency—the persistence of how a scene looks and what exists within it.

She places this progress in a research lineage that includes GQN and WALT. GQN concerns scene representation and rendering, so this is a broad lineage rather than a disclosed Veo architecture. Alongside generation, Bailey describes human-visible watermarks and SynthID watermarks for synthetic images and video. She also emphasizes collaboration with artists, including Darren Aronofsky, alongside musicians working with Lyria and artists working with Imagen.

6:166:37
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

6:16 · section reference included

Image detail includes typography

Imagen 4 addresses the static-image part of the workflow. Bailey’s examples range from humans to whales and puppies, then move from realism to detail across different styles. Typography is part of that detail: the Alamo Square and Mission stamp designs combine an illustrated treatment with readable place names. Bailey imagines the designs as laptop stickers, a small but concrete use for generated imagery beyond a gallery of samples.

The image section closes with designs from a collaboration with Ross Lovegrove. As with the video examples, the emphasis is on creators shaping an output’s appearance, rather than treating image generation as a single undirected request.

8:589:16
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

8:58 · section reference included

Music generation as a workspace for exploration

Bailey presents Lyria 2 as a high-fidelity music model with granular control over musical inputs, outputs, tones and styles. Music AI Sandbox gives those capabilities a visual workspace. Its editor, with colored waveforms, text fields and a clip list, resembles the kind of music-production environment an Ableton user would recognize.

Music AI Sandbox slide with an editor screenshot showing colored audio waveforms, text fields and a right-hand clip list.
Music AI Sandbox presents a visual workspace for musical exploration.

MusicFX, from Google Labs, offers natural-language composition of beats and supplies the background music for Bailey’s later commercial experiment. Lyria RealTime brings another emphasis: musical interaction developed with musicians, including Jacob Collier and Toro y Moi.

The musician clip that follows argues for learning by playing music and playing with music. Beginners do not have access to the whole musical landscape on day one; exploratory tools can broaden what they can try. The discussion connects music to mathematics, physics, history, geography, the human body and language, treating musical experimentation as a way into a wider set of patterns and relationships.

An audience member asks whether that clip was Veo 3 and comments on visible veins. Bailey initially says it was not, then qualifies that some visuals might have been generated with Veo 3. She does not identify which portions. Returning to Lyria, she emphasizes development with the creative industry and SynthID watermarking for the resulting music assets.

9:5110:09
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

9:51 · section reference included

One raccoon prompt across model generations

To make progress visible, Bailey holds the requested scene constant: a raccoon wearing a black jacket, dancing in slow motion in front of the pyramids. She starts with WALT, dating the example to 2023 and the underlying research activity to around 2022. Its output is visibly choppy in her account, and she describes the short sequence as not even reaching eight seconds.

The sequence then moves through LTX-Video, which she dates to 2024 and accessed through Hugging Face; Kling 2.0, dated to 2025 and accessed through fal; and Veo 2, dated to 2024. For Veo 2, Bailey likes the raccoon but questions how well it is actually dancing. That distinction matters: a convincing subject is only part of satisfying a prompt that also specifies an action.

Finally, the same prompt goes through Veo 3. The sequence is a qualitative comparison of examples, not a controlled numerical ranking. Its value is the repeated task: the subject, clothing, action, speed and location remain recognizable requirements as the models change.

11:5712:12
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

11:57 · section reference included

Animate a still and expand a short prompt

Image-to-video generation begins with a visual anchor instead of asking language to define everything. Bailey shows a woman with a puppy, then a western-style woman walking slowly forward toward what she describes as a gunfight. Other examples take a stylized person from a single frame and direct actions such as running toward the camera or being lifted into the sky by a tractor beam. Natural language specifies what should happen to the supplied subject.

Prompt rewriting works on the language side of that interface. The short raccoon sentence is expanded into a fuller description that Veo can more readily use. This adds a preparation step between the user’s concise intent and the generation request: elaborate the description, then generate from that richer specification.

Bailey asks the audience to watch the resulting example with music, sound effects and background noise in mind. She highlights the visual detail, including apparent reflections in the subjects’ eyes. The demonstration brings the two threads together: more fully specified visual direction and sound generated alongside the image sequence.

13:2013:32
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

13:20 · section reference included

Music, multilingual dialogue and visual oddities

The Paris team’s sequence broadens the demonstrations to music and dialogue. Bailey compares the musical style to Daft Punk, then points out a haircut floating at the top of a head. The examples can be expressive while still containing conspicuous visual oddities.

Her explanation of the dialogue supplies the jokes for viewers who do not speak French or Spanish: Spanish veo means I see, while another character responds that he does not understand because he is French. The sequence includes repeated references to artificial intelligence and an interview, then an alien exchange about whether to contact humans who are about to create AGI. The response is that the humans are not ready yet. These are demonstrations of generated audiovisual scenes carrying timing, language and humor, not just recognizable objects.

14:4915:00
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

14:49 · section reference included

Access routes and a small generation request

The access discussion describes the presentation’s rollout state, rather than a current availability matrix:

  • Google AI Ultra: access in multiple countries, with the UK described as recently added.
  • Google AI Pro: a limited number of uses through the Gemini mobile app.
  • Vertex AI: Veo 3 in private preview.
  • Veo 2: available through Vertex AI or the Gemini APIs.

Bailey expresses hope for AI Studio access and offers an early-access form through a QR code. That hope should be read in the Veo 3 context: the earlier Veo 2 developer release had already announced Veo 2 in AI Studio.

The code walkthrough exposes a small set of production choices: which model to call, where to deposit the generated video, whether to start from an input image, and which aspect ratio to request. An output bucket is the destination; a starting image is a generation input. Keeping those roles separate makes the request easy to reason about.

A Python request in the historical Gemini API style illustrates the model-and-configuration boundary. Here the repeated raccoon prompt is the text input, and aspect_ratio controls the output shape:

python

from google import genai
from google.genai import types

client = genai.Client()

operation = client.models.generate_videos(
    model="veo-2.0-generate-001",
    prompt=(
        "A raccoon wearing a black jacket dancing in slow motion "
        "in front of the pyramids."
    ),
    config=types.GenerateVideosConfig(
        aspect_ratio="16:9",
        number_of_videos=1,
    ),
)

print(operation.name)

This submits a generation operation; it does not make a completed video immediately available. The historical Veo 2 example illustrates the small request surface Bailey emphasizes, while her onstage walkthrough also discusses output-bucket and optional-image configuration.

16:3216:44
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

16:32 · section reference included

Rebuild a commercial with separate generators

The final experiment returns to advertising. Bailey had intended a live demo, but time and uncertainty about the demo lead to a walkthrough. The source is a Chick-fil-A commercial whose speaker also happens to be named Paige—Bailey explicitly says it is not her. The copy praises crispy breading, a tender filet, warmth and satisfaction. Recreating it requires more than matching a sandwich image: it requires shots, speech, music and a finished sequence.

The Veo 2 workflow begins by turning that finished source into a production plan:

  1. Give the original video to Gemini and ask for a detailed reconstruction plan.
  2. Segment the plan into generation prompts that fit the workflow’s eight-second Veo 2 clip limit.
  3. Use MusicFX to create a background track combining Down Home Farm and Slow Guitar.
  4. Bring the generated material into Camtasia and stitch it together with transitions.

Bailey describes the process as relatively quick, but still requiring substantial thought and work to reach final assembly.

She plays the assembled version, including the recreated introduction and sandwich description, then identifies Gemini text-to-speech as the speech source. The complete production therefore combines Veo 2 video, separately generated music, separately generated speech and manual editing. Bailey prefers her assembled result to the original commercial, but the operational point is how many outputs she had to coordinate to get there.

Process (Veo 2) flowchart with blue analysis and scripting, yellow parallel generation tracks, and purple final assembly regions connected by arrows.
The Veo 2 workflow separates analysis, generation and final assembly.
17:5518:08
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

17:55 · section reference included

One prompt, then an animated farewell

With Veo 3, Bailey describes a shorter path: take the original video, generate a description, and submit that description to Veo 3 in Flow. The shown output includes the Paige introduction and the copy about crispy breading and a tender filet. Bailey attributes the shown Veo 3 commercial result to one prompt submitted in Flow. Her account describes the reduction in manual work; it does not provide measured generation latency or repeatability.

The repeated playback makes the practical consequence of native audio concrete. In this example, the video and spoken performance arrive together, removing the separate speech-and-video assembly step that occupied the Veo 2 workflow. The result is still judged by how well it recreates the intended scene and delivery, but the creator has fewer independently generated pieces to coordinate.

Bailey closes with a commitment to broaden access and add more steerability controls, without a delivery schedule. The farewell provides one final demonstration: the Paige waving on screen is not footage of her waving. It is a static photograph animated with Veo 3—the same image-to-video mechanism now applied to the presenter herself.

19:3319:46
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

19:33 · section reference included

Resources

From the talk

Read the complete timestamped transcript
  1. 0:00

    [on-hold electronic music] Thank you so much for having me.

  2. 0:17

    Thank you all for being here today and wanting to learn more about generative media. Um, I'm going to keep this pretty quick. There's a lot to show, especially for some of the new features that we've released in Veo 3.

  3. 0:27

    Um, but there's also a lot, uh, to discuss in the, in the frame of how does this revolutionize the way that people build, um, things? How do people create ads?

  4. 0:36

    Uh, and how do people replicate, uh, some of the experiences that you might see every day? Um, so as mentioned, I'm Paige. I am the eng lead for our DevRel team at Google DeepMind.

  5. 0:45

    Um, but I'm here today on behalf of our generative media team, um, who are all brilliant. Uh, they are wonderful. Uh, Tom Hume, Dmitri Erhart, many of them, uh, could not be here today.

  6. 0:57

    Um, but, uh, but this is all their work. So I just want to send like a thank you to the heavens for, for everything that they've been building. Um, today, we're going to be talking about three different models.

  7. 1:07

    So Veo 3, which is our new video and audio generation model, Imagen 4, which can generate these static images, and also Lyria 2, which is a music generation model.

  8. 1:17

    Um, and stay tuned. There will be more of all of the above coming shortly. Um, Veo is, uh, kind of magical in the sense that, you know, you can create videos of things that you've never seen before, and we've saw-- we've seen some of these examples already.

  9. 1:31

    But it also has the potential, um, to really revolutionize everything that we build and create as humans. And I loved this quote from Andrej Karpathy, um, just recently, that video has the potential to be an incredible surface for communication, um, but also for education and for human creativity.

  10. 1:48

    So all of these things are designed with that in mind. Um, Veo 2, uh, just want to touch on it briefly before we get into the Veo 3 capabilities.

  11. 1:57

    Um, we released a whole bunch of additional things around creative control, um, just recently. So think on the order of about a month ago, um, or not even that.

  12. 2:06

    Um, so things like reference-powered videos, um, outpainting, the ability to add and remove objects, character control and consistency, which I think you'll, uh, y'all were talking about just a little while ago, um, and also the ability to interpolate across first and last frames.

  13. 2:20

    And let's take a look at what that means, um, because for people who are not necessarily filmmakers, um, it's much, much easier to show and not tell. Um, so reference-powered videos are things like you have, uh, you have a person, you have an environment, um, and you're able to kind of put one within the other or compose

  14. 2:38

    them together into something that, um, feels very, uh, very stylized, but also like really, really well, um, well-crafted. Um, so here you can see a reference-powered video, um, with Veo 2.

  15. 2:50

    This is available via the API and via some of our tools like, uh, like Flow today.

  16. 2:56

    Uh, and another example of reference-powered video, so a really, really cute, uh, little monster in a variety of environments that you just control by describing. Um, it's performing pretty well on benchmarks.

  17. 3:09

    So you can see here, um, the green is, uh, Veo. So, um, uh, compared to things like Runway, Gen-4, and Kling, um, the more green, the better. Um, and for reference-powered video, most human raters, um, uh, uh, selected for Veo for some of these side-by-side comparisons.

  18. 3:27

    You can also match styles, um, so upload a reference image and then have, uh, different styles composed together.

  19. 3:36

    Another example of styles being preserved, and then also camera controls the same that you might have if you were a filmmaker. So things like being able to move back, move right, um, rotate up, zoom in, um, and to be able to precisely control all of these camera movements, again, just via natural language and through some of the

  20. 3:54

    tools that are available in the APIs. Um, when I saw all of this, I was blown away because I don't think that we have nearly enough code samples demonstrating some of these capabilities.

  21. 4:04

    Um, but these are all things that you can do with the Veo 2 models today, um, with the APIs. We also have the ability to do outpainting. This was important for, uh, a recent project with the sphere around "Wizard of Oz."

  22. 4:16

    So being able to take a sc-- uh, like a sc- a scene or, or an individual frame of a video, um, and imagine what the rest of the scene might look like.

  23. 4:26

    So even if you only have a view into a small portion, um, being able to create something that, uh, that looks real, um, or that looks consistent, uh, across the outpainting.

  24. 4:39

    Um, adding objects or removing objects from scenes. So you can see here a few examples as well. Um, and again, all of these available for you to test and to try today.

  25. 4:49

    Um, these are all, these are all things that exist, um, uh, that our, that our research team has kind of gotten into the, to the API designs. Um, some more examples of removing objects.

  26. 5:01

    Character control is quite nice. You might have seen some of these demonstrated into your favorite products, um, for, um, kind of controlling mouth movements, controlling reference face movements for particular, um, for particular characters.

  27. 5:14

    We also give the ability to add a script, um, and to add kind of a voice tone and to have the character, um, map the lips to producing that sound and producing it, um, in a way that feels consistent with the, with the location.

  28. 5:29

    Um, these are some of the, the motion examples, um, with Veo 2 and with Veo 3. So being able to have an input image and controlling them or changing the design, um, across, across the scene.

  29. 5:45

    Um, more benchmarks, and then the first and last frame. Um, so you can have an input image and an output image, and Veo is able to interpolate across them, um, to, to kind of make those, make those images, um, stitch together into a video.

  30. 6:03

    And those are just another couple of examples. I feel like generative media presentations are very gratifying. Like, these are certainly the most beautiful things, um, that we, uh, that we get to see at developer conferences.

  31. 6:16

    So Veo 3, everything that you just saw, Veo 2. Veo 3, right? Like, like blown away. So Veo 3 is, um, video but coupled together with audio. And so all of the tokens compose together natively, um, not audio being pulled in as a tool, um, but the model actually able to compose together all of these tokens across

  32. 6:37

    multiple modalities. This is similar to what you see with Gemini's native audio output. Um, in addition to being able to output text and code, you can also output images, edit images, and also edit audio, compose audio, et cetera.

  33. 6:51

    Um, so Veo 3, our latest state-of-the-art, uh, video generation model. Um, it's, uh, you know, it, it has these things around prompt adherence and native audio generation. Um, but again, so much cooler to show, um, and not tell.

  34. 7:05

    The little llama. Then another one that... [upbeat music] So, so interestingly, you're able to do not just, uh, not just, uh, background noises, but also things like, uh, things like music.

  35. 7:40

    Including s- very, very subtle sounds and the like.

  36. 7:47

    So let's, uh, let's go to the next.

  37. 7:51

    There we go. So, so this is hard, right? Like, like, uh, it looks very cool. It's very hard to capture the nuances in an, of an input prompt. Um, and it's also been really historically very hard to, to preserve visual consistency.

  38. 8:05

    Um, so characters often, like, jump from one, um, um, from one frame to another. There might be backgrounds, uh, and then suddenly walls disappear and you're able to see behind them.

  39. 8:15

    Um, this is one of the reasons why Veo 3, uh, feels like a leap forward, is because the stylistic consistency and then also the contextual consistency, um, is much, much better.

  40. 8:25

    Um, built on years of research, um, so things like GQN, um, Walt, et cetera. Um, and it has responsibility at its core. So you can see little, uh, human visible watermarks as well as SynthID watermarks, um, for synthetically generated images and video.

  41. 8:41

    We've also been partnering really closely with many, many, many, uh, artists along the way. So, uh, Darren Aronofsky, um, also, uh, musicians for our Lyria models, artists for the Imagen models, um, and we'll take a look at a couple of these as well.

  42. 8:58

    Um, so Imagen is image generation. Um, uh, you know, able to kind of preserve realism, um, everything from humans to whales, uh, you know, uh, cute puppies. Um, I've heard that the mo- the more cute puppies that you have in a presentation, the better it always is.

  43. 9:16

    Um, and then also being able to preserve detail across all of these images as well, um, including diverse styles, uh, and even things like typography. So I love these stamps of Alamo Square and the Mission.

  44. 9:28

    I really, really wish that we just had these as, um, uh, like swag ideas that, uh, uh, stickers for laptops or stickers just in general. Um, and, uh, another example of an artist that the team has been really closely collaborating with, Ross Lovegrove, um, on, uh, on some of his, uh, some of his designs.

  45. 9:51

    So, uh, Lyria 2, um, also very exciting. It's high-fidelity music and professional-grade audio. Um, they also gives you a very, very granular creative control, so the ability to steer the inputs and outputs and to steer the tones and the styles of the music, um, along the way.

  46. 10:09

    Um, Music AI Sandbox is one of the products that's been created as a visual for this. Um, if folks are familiar with Ableton, um, or things like it, uh, this, this probably looks, uh, looks very similar.

  47. 10:21

    Um, and then there's also MusicFX, which is a project from our labs team, uh, that allows you to, to kind of, uh, compose together beats just via natural language.

  48. 10:31

    Um, and we use that for the demo later on today. Uh, Lyria Real-Time, uh, has also been a deep collaboration with many musicians, um, both Jacob Collier, uh, who's a legend, and, uh, Toro y Moi.

  49. 10:44

    Um, and me circa, like, uh, me circa college was, like, blown away that, uh, Toro y Moi was at, uh, Google I/O. Um, huge fan.

  50. 10:56

    I do think that a big part of teaching music is giving people a chance to play music and play with music.

  51. 11:02

    Yeah.

  52. 11:03

    However, people don't have access to the whole of music from day one.

  53. 11:07

    Yeah.

  54. 11:07

    What I think this offers an interesting perspective into is the whole of music, mathematics, physics, history, geography, the human body, language, spelling, syntax. One thing I've come to realize is that a lot of the same forces that make music work are the forces that make life work.

  55. 11:27

    Was that Veo 3?

  56. 11:28

    Oh, that was not Veo 3. [audience laughing] Some parts, but it's hard to tell, right? Like it's-

  57. 11:34

    The veins were like-

  58. 11:35

    Yeah, yeah. Yeah, you know. Parts of it, parts of it, um, might have included, uh, visuals generated by Veo 3 though. Um, so Lyria, again, built in collaboration with the creative industry, um, not outside it.

  59. 11:47

    Uh, and then also incorporates many of these techniques like SynthID to make sure that you have some sort of digital watermarking for the, for the assets themselves. Um, so now we're gonna get into it.

  60. 11:57

    Uh, these are some examples that I thought might be fun to share to show just how far we've come in the last couple of years. Because I think being here, we get very sucked into the Bay Area bubble, and we don't really kind of take a step back and appreciate how far, um, you know, the world has

  61. 12:12

    changed in just a matter of months. Um, so this is an example of one of the, one of the papers that was produced around 2023. Um, so released in 2023, research happening around 2022, um, for text to video.

  62. 12:25

    A raccoon wearing a black jacket dancing in slow motion in front of the pyramids. So just have that in your brain when you see it. Um, this is Walt, uh, circa 2023.

  63. 12:35

    Um, so very, very choppy, like really, really hard to have, um, even-- Like, that's not even eight seconds worth of a frame. Um, LTX video from 2024. This is one of the ones that were available on Hugging Face.

  64. 12:49

    Kling 2.0, which I did with Fal. Um, uh, heck yeah. Yep. Yeah. So, uh, which I did with Fal, um, released in 2025. Uh, and then Veo 2, 2024.

  65. 13:00

    Um, a very cute little raccoon, but I'm not sure how well he's dancing. Um, but this is just kind of a splattering of how the world has changed in the space, um, of just a couple of years.

  66. 13:10

    And then when you put the same prompt, um, through Veo 3, um, you get this. [whooshing]

  67. 13:20

    Very stylish raccoon. Yeah. There we go. So image to video, um, transforming static images into dynamic video content. So you can see here an image, um, of a woman and her puppy.

  68. 13:32

    Um, uh, a woman I, um, can only assume is in Texas, um, walking very slowly forward on the way to a gunfight. Um, and also different, uh, stylized images of a person in a single frame, um, being applied to, to different scenarios.

  69. 13:48

    So running towards the camera, um, tractor beam taking off, being lifted into the sky, again, all steered via natural language. Um, prompt rewriting is something that we've also released in Veo 3.

  70. 13:59

    So the ability to take that very, very simple, um, sentence before of a raccoon wearing a black jacket dancing in slow motion in front of the pyramids, um, and turn it into something a bit more fully formed, um, uh, that the, that Veo is, is much better able or better equipped to understand.

  71. 14:17

    Um, and so have that in your brain, um, as well as the, the concept of sound generation, so both music, sound effects, and background noises. Um, and we'll take a look at what that simple prompt, um, is now. [whooshing]

  72. 14:37

    Wow. I still, like, get blown away about how much detail there is, like being-- You can almost see the reflections in the eyes, um, for the, for the things as they, as they walk forward.

  73. 14:49

    Um, and then this is from our team in Paris. Uh, and they are very, very excited to have it shared today. Um, does anybody know [REDACTED:origin] in the audience?

  74. 15:00

    Okay, amazing. The rest of us will have a translation in a second. Um. [upbeat music]

  75. 15:09

    It's like Daft Punk. I can't believe this new Veo model. It is amazing.

  76. 15:18

    Oh, yeah. [laughs] [shouting]

  77. 15:18

    Yeah. [laughs]

  78. 15:49

    You can see the hair cut at the top, like, floating around. And for anyone who was curious, the translation is, um, so Veo means I see, I see, I see in Spanish.

  79. 16:09

    Um, uh, and then the guy responded with, "Yeah, I don't understand what you're saying because I'm [REDACTED:origin], actually." Um, artificial intelligence, artificial intelligence. Um, and then, uh, also the-- that's the interview.

  80. 16:22

    Um, and, uh, "Boss, the humans are about to create AGI. Should we contact them?" Um, "They're not ready yet." Uh, so amazing.

  81. 16:32

    Cool. Or actually, I'm gonna zoom to the next, zoom to the next one. Uh, so how do you access it? Big question. Um, right now we have a few different ways to access the Veo 3 models.

  82. 16:44

    One is through the Google AI Ultra plan, um, which is available in many different countries, um, including the UK just recently. Um, the Google AI Pro, uh, subscribers, so being able to access via the Gemini mobile app for a limited number of uses.

  83. 17:00

    And then also Veo 3 is available in private preview currently in Vertex AI. Veo 2 also available via Vertex AI for the Gemini APIs. Um, hoping to bring them to AI Studio.

  84. 17:10

    Um, crossing fingers, but, but we'll see. Um, and then you can also fill out a form for early access if you would like it. QR code coming shortly, so take a picture of that if you would like, um, with the form to, to go, uh, to go submit and test it out.

  85. 17:26

    Um, this is also, while you're looking, a code sample of how easy it is to use, um, with, uh, just your output bucket where you would like the, the video to be deposited, um, if you have an input image that you want to use as a starter, um, some things around aspect ratios, um, uh, and the like.

  86. 17:44

    So, um, toggling between different models or being able to specify some of these is just a handful of lines of code, which is pretty magical. Um, so this was intended to be a live demo.

  87. 17:55

    I'm not sure if I can tempt the demo gods, and also I am probably, like, close to being over time. Um, but I wanted to see how well I could replicate, uh, a commercial, um, with Veo that seemed pretty simple.

  88. 18:08

    Um, so the commercial is this one which took my name.

  89. 18:13

    Hey, my name's Paige, and what makes the Chick-fil-A chicken sandwich-

  90. 18:15

    Took my name. That's not me. [laughs]

  91. 18:16

    ... original to me is the crispiness of the breading and the tenderness of the filet. It's tasty, it's warm, it's total satisfaction.

  92. 18:25

    Hey, my name's Paige, and what makes the Chick-

  93. 18:27

    So that was the, that was the input video. Um, the process for replicating it with Veo 2 is you give Gemini the original video, have it create a really, really detailed plan, um, segmented into prompts be-- to like handle the eight-second limitation, um, used video-- uh, MusicFX to create the, the background track, which was a combination of

  94. 18:46

    Down Home Farm and Slow Guitar. Um, and then, uh, put it all in Camtasia, stitch it together with transitions. Whole process, uh, was relatively quick, but it also took a lot of thinking and a lot of work to get to that final assembly stage.

  95. 19:00

    Uh, and it looked like this. So this is, uh... [gentle music]

  96. 19:07

    Hey, my name's Paige, and what makes a Chick-fil-A Chicken Sandwich original to me is the crispiness of the breading and the tenderness of the filet. It's tasty. It's warm.

  97. 19:17

    It's total satisfaction.

  98. 19:20

    And that was, uh, again, like using, uh, using Veo 2, using Gemini text-to-speech, um, and stitching it all together myself. And I actually like that one better than the, the original commercial, but your mileage may vary.

  99. 19:33

    Process with Veo 3, um, you have the original video, you generate the description, you give it to Veo 3, and you see how well it does. Um, uh, this took, uh, just the span of like submitting the prompt to Veo 3's inflow.

  100. 19:46

    Hey, my name's Paige, and what makes the Chick-fil-A Chicken Sandwich original to me is the crispiness of the breading and the tenderness of the filet.

  101. 19:54

    So again, one prompt, and it was able to produce this.

  102. 19:59

    Hey, my name's Paige, and what makes the Chick-fil-A Chicken Sandwich original to me is the crispiness of the breading and the tenderness of the filet.

  103. 20:08

    Incredible. Um, so takeaways, Veo 3, pretty magical. We're committed to expanding it, expanding access as quick as we can, um, and around, uh, adding controls around steerability. Um, and thank you so much.

  104. 20:24

    That is not me actually waving at the camera. That is a static photo of me that has been Veo 3 animated. Excellent. Thank you. [audience applauding] [upbeat music]