AI Engineer World's Fair 2026
Generative Video at the Speed of Light
About this talk
uRun founder Keegan McCallum discusses how generative video is advancing in efficiency and continuous, real-time output alongside visual quality. Using Helios and its Wan 2.1 14B lineage, he explores long-horizon generation, declining usage costs, interactive avatars, accessibility, distributed GPU deployment, WebRTC networking, and agent-oriented application development through a CLI or MCP server.
Chapters
- 0:00Introduction and the shift from video quality to efficiency
- 1:16Helios and real-time long-horizon video generation
- 2:57Interactive video models and generation economics
- 4:23Accessibility and visual human-computer interaction
- 6:01Global GPUs and real-time networking infrastructure
- 7:39Agent-facing developer tools and closing invitation
Talk transcript
- 0:00
[on-hold jingle] I am Keegan. I'm the founder of uRun, um, a new kind of inference provider
- 0:20
focused around, uh, interactive media. And I'm here to talk about generative video. So we hear a lot about generative video improving along the quality axis at the frontier. We have the classic Will Smith eating spaghetti from twenty-twenty-three.
- 0:36
It is nightmare fuel and not something you would ever mistake for reality.
- 0:42
In twenty-twenty-four, we got Sora, and it gets a little better. It'll-- still has a bit of, you know, an AI feel to it, but it, it's getting there. And Sora 2, you know, even better.
- 0:54
But Cdance this year, um, absolutely incredible. So photorealistic and it's, it's no wonder that we talk a lot about quality. But I'm here to talk about another axis which models are improving along, which is efficiency and, uh, h- the long-horizon generations. [sighs]
- 1:16
So what you're watching here is a demo for a model called Helios that we serve at uRun. Um, the generation in the bottom right corner, you'll see, is a long, continuous generation, and the other video, um, is a bunch of clips, um, that have been generated faster than you can consume them.
- 1:35
Uh, and they're about at the same quality as the frontier models were last year. [inhales] They're-- Uh, Helios is a distill of Wan two point one fourteen B, um, and I'll talk a bit about the techniques that are used in the various models that are hitting the scene right now.
- 1:56
But there's been an explosion in just the last year, uh, in terms of efficiency and capabilities.
- 2:03
Um, so like looking at this, I kinda ruined it with the last clip, but you can guess which one is real-time, um, and which one was generated in, uh, a number of minutes.
- 2:15
And the one on the right, uh, is in-- arguably a bit better. It's got better motion, um, and it was generated for about a one-hundredth of the cost.
- 2:28
And these are just some of the charts showing, um, the quality bar, uh, for both long and short video generation. Uh, Helios came out in March, and it's, it's pretty incredible to see how fast these are improving.
- 2:42
But these are techniques that are being applied, uh, all over the place, not just to one model. Uh, there's world models which can keep consistency over long horizons, and you can control, uh, in a fine-grained way, uh, the camera and the viewport.
- 2:57
Uh, there's avatar models like we just talked about with Lemon Slice, um, and there's video-to-video models that can, can transform, uh, what you're seeing in, in real-time, almost like a, a magic mirror.
- 3:10
Uh, there's actually been an explosion of innovation. There's been, uh, at least forty models, uh, with real-time capabilities, uh, and long-horizon generation capabilities released this year. Uh, show of hands, who ha-- here has burned ten or even fifty dollars worth of tokens in an hour with CloudCode?
- 3:33
A lot of people. Um, and so we're at a place right now where ten dollars can get you three hours worth of generated video continuously with most of these models, and fifty dollars would give you an entire day interacting with an AI in a visual medium.
- 3:49
Fifteen hours. And so I wanna talk a little bit about the different things this enables in terms of the way that we interact with computers, um, a-and I'll talk a little bit about what we're doing at uRun to try and make it easier for folks to experiment and build out applications like this.
- 4:07
So one such use case would be a magic mirror. You could have, uh, your webcam, and you could ask to see yourself in any outfit. You could ask to, uh, see yourself in a car you like or with a haircut you're considering.
- 4:23
Um, a lot of different, uh, possibilities because these are open-ended models that can transform, um, what they're seeing on a webcam in real-time. I also think about accessibility a lot with these models.
- 4:34
Um, you know, working with AI involves a lot of reading and a lot of text. Uh, for some people, that's more difficult. Uh, for some people, they just don't think, um, in, in text.
- 4:47
They think visually and learn better that way. Uh, so there's more opportunities to have companions or visual mediums that are gonna allow more people to experience, uh, the things a lot of us have with coding models.
- 5:02
And I'm excited about content creation. Um, so far, we very much had a slot machine type approach where you're setting up a prompt and maybe some keyframes and spending about ten dollars a minute, uh, to try and get the shot that you want.
- 5:17
But with these models, you can actually steer them in real-time, um, in under a second while they're generating and get the actual shots that you want. Maybe you're piloting an agent that it-- you're able to look over its shoulder and see what it's generating in real-time.
- 5:33
Um, but you're able to more granularly control the content you're generating. And with modern models like Google, uh, Gemini Omni, you can actually render these out as a more full-fidelity clip.
- 5:47
And of course, we all are thinking about world models, uh, but I wanna take the, the focus off of just kind of the, the, the basic, uh, world models that we talk a lot about and just try to expand, uh, the horizons of what we can do with this technology.
- 6:01
And so what does it look like to actually build an application like this? So you're gonna need GPUs all over the world, potentially, if you've got a global audience that's gonna be using these.
- 6:14
You're gonna need to think about, uh, where you're connecting the, uh, users to, what GPUs you're gonna use to serve them. You're gonna need to set up probably WebRTC and ICE and TURN.
- 6:27
Um, and for the most interesting use cases, you're gonna wanna model, um, wire multiple models together in continuous streaming workflows, um, building those real-time harnesses. And you're gonna want, um, things synchronized with your controls that you're providing to your end users, um, with every frame, uh, and continually providing a smooth streaming experience.
- 6:54
And so our idea is what if there was just a React component that you could drop into your application, uh, to make it easy to provide video interactively inside your applications with any model.
- 7:13
And behind the scenes, there's a programmable Python runtime that lets you easily build these complex pipelines generating asynchronously, um, so that you can build avatar models, you can build these video-to-video transformation models.
- 7:32
You can experiment and, and build whatever you can really imagine on top of these.
- 7:39
And I argue that in 2026, we don't just need platforms, we need software factories and ways for our agents to interact with these. And so we've actually built one that will let folks hook into a CLI or an MCP server and build these kinds of applications.
- 7:58
And so the models are here, and the frontier is really in how we serve them. Uh, I
- 8:09
went way over-- I went way under time. Um, [laughs] but we are looking for design partners who wanna push the boundaries of human c- human-computer interaction, and we're hiring at uRun.
- 8:20
Um, so come see me after the talk if, uh, if you're interested in chatting more. [upbeat music]