← All AI Engineer talks

AI Engineer World's Fair 2026

The Next Medium: Why Real-Time Interactive Video Changes Everything — Ahmed Ahres, Reactor

About this talk

Reactor's Ahmed Ahres argues that world models are best understood as real-time, interactive video rather than static or batch-generated media. He contrasts continuous, controllable generation with conventional video models; discusses applications in advertising, immersive education, audience-directed livestreaming, and video editing; and introduces Reactor's developer platform and Helios model. Audience questions address interactive generation frame rates and ongoing research.

Chapters

  1. 0:00World models and the case for interactive video
  2. 1:56Ahmed Ahres, Reactor, and lessons from real-time media
  3. 4:48Continuous video generation and creative control
  4. 8:33Education, interactive livestreams, and video editing
  5. 10:59Reactor's API, Helios, and previsualization
  6. 14:28Audience questions on frame rates and research

Talk transcript

  1. 0:00

    [upbeat music] Hi, everyone. Uh, welcome to the talk. First of all, thank you all for making the time.

  2. 0:16

    I know it's the last talk of the day probably, or I think the last one is at, uh, three forty-five. But, uh, yeah, thank you all for your time.

  3. 0:22

    I, I know you're all probably very busy. Um, today I'm gonna be talking about something that is a little bit slightly futuristic, though not for San Francisco, and that's World models.

  4. 0:31

    Um, I know here it's written real-time interactive video, but the way we think about World models is really in the real-time interactive video, and I'll explain why. Um, and in today's world, I think World models is a little bit of a marketing term that people think about it from a Gaussian splatting standpoint or others from video.

  5. 0:47

    Uh, but the way we define World models is really real-time interactive video, and I have strong evidence or, like, beliefs that this will actually change everything in how we produce and consume content.

  6. 0:59

    Oops. What happened? Sorry. [clears throat] Sorry about that. Cool. [clears throat] So if you think about video, video has always been something passive. Before, people used to produce videos and movies, and then we would watch it.

  7. 1:19

    And in today's world, we have these models, the Veo 3, the CDense 2, and what they do is you prompt them, you get back a file, you watch, and good luck.

  8. 1:28

    It's a slot machine. You cannot change it. You cannot do anything about it. So there's a question that, um, we like to think about in our company is what happens when video becomes programmable like software?

  9. 1:39

    And what happens when pixels can be generated in real time? Once this happens, it actually changes completely how we think about consuming content and even producing content, and I'll be talking about how in history we've seen that real time has always been the future and the co- kind of applications that it unfolded and how you can get

  10. 1:56

    started today. Quick, quick background here. Um, so my name is Ahmed. I'm the head of go-to-market at Reactor. My background is in computer vision and machine learning. Um, and I state machine learning because this is the time where we used to actually train our models.

  11. 2:12

    Um, I built and shipped games on iOS, uh, on iOS and Android for fun. Um, I was a founder, um, and today, I'm the head of go-to-market at Reactor.

  12. 2:21

    And who we are, quickly, is we're a Series A company building the platform for these real-time World models. So, so far, main-- most of them have been still models, but we are building the infrastructure and the developer platform to make them usable so that anyone can integrate these real-time interactive video, and we can democratize access to this

  13. 2:40

    technology. I'll start with a problem. We can today generate pretty much anything, but we just can't change it, right? A generated video, as I mentioned earlier, is still a recording.

  14. 2:52

    You get it back. You can't do anything about it. And real-time changes what the medium is. It doesn't just make it faster.

  15. 2:59

    And I'm gonna talk about two examples that actually show us what real time has unlocked in the past. Before, in the nineteen fifties and even before, we used to look at a map to know where we are.

  16. 3:10

    Someone produces a map, you look at where you are, and that's it. You cannot do anything about it. Then GPS came. GPS made it real time. Suddenly, I can know where, where I am at the instant, and I can, you know, track it.

  17. 3:21

    Now, you think GPS has just made it a bit faster to know where I am, but actually, Uber would not exist if we did not have the GPS. Another example, which is even a bit more powerful, I'd say.

  18. 3:34

    Before we used to d- to use the film to produce content, right? We had a film. Someone shoots something. They can't see what they're shooting. They go somewhere. They produce that film, and then you can see it.

  19. 3:45

    And then it became digital. You can start seeing what you're shooting. If today you pick up your iPhone and you start recording a video, then you can see what's going on, and that's why we can produce high-quality content.

  20. 3:56

    It's because you are able to see what's going on in the screen and adapt accordingly. That gave rise to In- Instagram and TikTok. Instagram and TikTok would not exist if we could not produce high-quality content, and the only reason why we're able to produce high-quality content, among many reasons, is because we can see in real time what's

  21. 4:14

    happening. It's not a slot machine. We can actually just edit and see, and that's what unlocks all of these new use cases.

  22. 4:24

    So when video becomes programmable, you can address it, you can condition it, you can change it. I can show you on the screen whatever I wanna show you on the screen, and it becomes programmable a little bit like software or like anything that is programmable in the world.

  23. 4:39

    And in, in the market today, we are seeing three kinds of models that do this. Some of them you will be familiar with, others you will not be familiar with.

  24. 4:48

    The first one is these are Veo-- think about Veo or Sora, but real-time and interactive, meaning that, first of all, they're infinite. So they don't stop after five, ten, or thirty seconds.

  25. 5:00

    They actually continue forever. They're interactive, meaning you can change what's happening on the screen,

  26. 5:06

    and they're in real time, so you don't need to wait to see what's going on. And assuming this works. So this is an example of a, of a video that I passed an image with a dog, and this was all generated in real time.

  27. 5:17

    And at some point, I'm gonna prompt, "A cat shows up," and you will see that a cat showed up in the, in the video. This would not be possible in the existing batch regular video generation models because you would get back the video, and you cannot do anything about it.

  28. 5:30

    I could have added anything. I could have gone on to create an entire story with it. I could have said the ca-- the dog starts running, starts jumping. A ca-- a dragon shows up.

  29. 5:38

    It goes to, I don't know, to the World Cup. All of this would have happened in front of you.

  30. 5:43

    These type, these types of models pro-- like, unlock a few things. First of all, control. So if you think about generative media today, the big problem that content creators all have is I don't have the control I need.

  31. 5:55

    Like, yes, it's great to use CDense 2 or Veo 3 to generate videos, but I just don't have the control. And this is always the thing that any filmmaker or movie producer or any content creator will tell you.

  32. 6:07

    And so real-time actually ends the slot machine type of mentality and actually gives you the cr- the control that you need. And a big saying I like to say is, "Instant feedback is the ultimate level of control," and you-- we will never be able to have this level of control if we don't have real-time.

  33. 6:24

    The second thing it unlocks which, you know, among other things, which is a field I'm not particularly fond of, but I think is gonna be big, is advertising. If I can know what you looked for a minute ago, why can't I insert the logo of whatever you've been looking for?

  34. 6:39

    Why can't I produce an ad in real time in front of you? We don't need to produce, pre-produce anything. Now granted, this is gonna take some time because, you know, brands afraid of AI, afraid of, uh, you know, if their logo has one pixel that is white instead of dark.

  35. 6:55

    But it will happen eventually. And so-- And I think at the moment that happens, we will not be to-- need to produce any ads anymore. Everything will be happening in front of you in real-time.

  36. 7:06

    The second type of model is the one that probably you're most familiar with. It's the Genie 3 like from Google. These are models where you can pass an image and a text typically, and you can control a character.

  37. 7:18

    They're fun. The first thing that you think about when you think about this is games, right? It's a character, it's a world, you can generate anything. But actually--

  38. 7:28

    It actually goes way far, way beyond games.

  39. 7:32

    It creates entire new interactive experiences where, combined with the first types of models that I talked about, we've already been seeing people in our community building a mix of games and, and movies.

  40. 7:42

    If you've ever watched Bandersnatch from Netflix, which is the ga-- the movie that you can pick your next scene, this is one of those things that gets, uh, that, that, that becomes possible, that you can control a character, you can control what's happening and create entire new types of interactive experiences that were not possible.

  41. 8:00

    The second thing is robotics. So because you can simulate and you can control, you can actually create as much training data as, as you want. And robotics-- World models in robotics is actually a giganormous market today.

  42. 8:12

    Um, I cannot tell you the number of robotics labs that, um, are training and building these models. But because you can control whatever you wanna control in any environment, this creates a new opportunity to generate infinite amount of data for robotics.

  43. 8:27

    And finally, something I like to think about, this is more maybe a passionate thing that I have, is education. Um,

  44. 8:33

    because you can step into anything, who-- With today's world in AI, I don't actually believe that edu- in the future of education is LLM-based or textbook-based. If you can put any, any kid in the situation, for example, in a history lesson, that enables entire new types of edu-- experiences that, uh, can be educational.

  45. 8:56

    The third type of model is probably the type of model that, you know, is more, let's say, um, co- something that we've been seeing before, which is avatars, but live and interactive.

  46. 9:06

    The thing with avatars though is it hasn't actually been cracked. They're still all kinda weird. If you speak to an avatar in any customer support or anything, it's still kinda off, right?

  47. 9:16

    Um, and these types of models, and we're seeing a rise of these live and interactive avatar models in research preview, that combined with Model one and Model two is actually gonna be, I believe, a, a big change in what we've been seeing so far.

  48. 9:30

    And this will be applied to things like customer support, training, sales, gaming, streaming services, et cetera.

  49. 9:39

    And so ju-just to give you a glimpse of what our users are building today at Reactor with these types of models, some of them are building interactive livestream, right?

  50. 9:47

    A livestream where people are watching, and then the users can type what happens next, and then they vote. Why? Because pixels can be generated in real-time. So there's no reason why I cannot put a livestream on X, YouTube, or Twitch and enable users to pick what happens next.

  51. 10:02

    Something that was surprising to me is a little bit on, on the medical simulation. So we've seen users create applications where you generate a world, and then you simulate what happens next.

  52. 10:12

    What if I put this medicine? What if I remove this medicine, right? And this creates-- can be a training playground for people wanting to become doctors.

  53. 10:21

    The third one, which is also kinda surprising, is cooking simulation. So people are building applications where you can simulate cooking and what happens if you put this ingredient.

  54. 10:32

    And finally, video editing. With video-to-video models, video editing becomes very interesting because I'm able to just add visual effects in real-time, and we've seen people build entire video edit-editing platforms.

  55. 10:46

    Now granted, they're not very good yet, just because of the quality of the models, but it's a new paradigm once you're able to v-- edit videos just via prompting or via talking to it or via clicking.

  56. 10:59

    And now for the final part, how do you actually do all of this? And this is why I like to say the world behind an API.

  57. 11:07

    At, at Reactor, we have four types of models today. First one is called Helios, which is the interactive video model that I talked about. This one is from ByteDance.

  58. 11:15

    LingBot, which is a world model like Genie 3, trained by Alibaba.

  59. 11:20

    LongLive 2 from NVIDIA, which is multi-shot film. You can prompt things in advance and create a consistent story over time.

  60. 11:30

    And SANA-Streaming, also from NVIDIA, which is a model that does video-to-video editing. So people are already using this, for example, by shooting something or creating something on Seedens2, uploading it, and then adding visual effects, removing people, adding background.

  61. 11:44

    And this be-- gets very interesting in previsualization, for example, for Hollywood movies.

  62. 11:51

    And under the hood, when we talk about infrastructure,

  63. 11:54

    the thing that I think I like to drive home is building infrastructure for regular video generation models is very different from real-time. Because in regular video generation models, you're talking about requests.

  64. 12:06

    You just send a request, a job gets run in the cloud, and I'm oversimplifying here, but you know, a job runs in the cloud, and it gives you back a file.

  65. 12:14

    With real-time, it's a different ballgame. Uh, you cannot just take what works for batch inference and apply to real-time inference. For example, you need to r-think about streaming, right?

  66. 12:25

    Once you need to think about streaming, um, pixels, uh, from a server to the client, um, it adds entire new-- entire complexities that batch generation does not have to think about.

  67. 12:37

    The second one is that everything is a live session, so everything runs constantly, and there's memory to be kept into account. Now, granted, one of the things that live vi-- live real-time models struggle with is memory.

  68. 12:49

    If you've seen demos from Genie-3, for example, we've all seen that the character can look back and then not remember what-what's going on. So there is a lot of work that needs to be going into maintaining that context window so that you can remember what happened if you turned your character left and right.

  69. 13:06

    And finally, global scale. If you th-if you think about real time, it needs to be sub one hundred millisecond latency anywhere you are. And if you're deploying applications in the world, then someone based in India or someone based in Japan should be routed to a GPU that is based in India or Japan, or as close as possible

  70. 13:25

    to it. If not, if you don't have the compute worldwide, then the experiences are not real time anymore, and it breaks completely the medium.

  71. 13:35

    And with Reactor, this is as easy as it gets to integrate these real-time models. Okay, I kind of maybe oversimplified it a little bit here, but it really is maybe ten lines of code.

  72. 13:44

    Um, and we have docs to do all of this. But essentially, you can just load the model with an API key and just start integrating it in whatever video, image, plugin, anything that you're building.

  73. 13:57

    And if you wanna get started, there's a QR code there, and I've added a promo code, AIE2026, which will give you seventy-five dollars off. Well, seventy-five dollars worth of credits, which is, um, significant amount of compute, uh, in our case because we make the models extremely cheap.

  74. 14:15

    Thank you very much. [audience clapping] I'm happy to take questions if anybody has a question.

  75. 14:25

    Oh. Hi.

  76. 14:28

    So I saw the sixteen FPS was the stated, uh, frame per second on the video game generator. What would it take to get that to something more like doable thirty FPS?

  77. 14:39

    Uh, well, one thing that actually we do is multi-GPUs. So we use multiple GPUs, um, optim-optimizing the model weights, applying quantization techniques. So there are ways, it's just a matter of priorities, but it is, uh, there are ways around it.

  78. 14:56

    Yeah. Good question, though. Hi. Yeah.

  79. 15:08

    Uh, I'm, I'm, I'm wondering if you are gonna be showcasing something at IBC in Amsterdam?

  80. 15:15

    Sorry?

  81. 15:16

    Are you aware of what the IBC is?

  82. 15:18

    Uh, no.

  83. 15:19

    The International Broadcasting Convention.

  84. 15:21

    Okay.

  85. 15:21

    There is, uh, I was wondering if you're gonna be presenting some showcasing regards of the, uh, next generation.

  86. 15:27

    I wasn't, but now I'm gonna look at it. Yeah. [laughing]

  87. 15:31

    Yeah. Yeah.

  88. 15:35

    Yeah. How do you feel like about comparing, uh, some deterministic engines too inside-

  89. 15:41

    Yeah

  90. 15:42

    ... of the traditional 3D Russell? Can we get a lack of feedback maybe?

  91. 15:45

    Yeah.

  92. 15:46

    Do, do you guys experiment with that? Can you share something?

  93. 15:49

    When you say deterministic engines, what do you mean exactly?

  94. 15:53

    Oh, any, any kind of, like, deterministic rule sets that you check against maybe, so the simulation stays, uh, grounded. I would, that's how I would describe it.

  95. 16:03

    Hmm. So we don't do any of that today. Um, the reason why also we build a developer platform, and but we've seen people build that on top of us.

  96. 16:11

    Um, building the infra for this is already a lot of work. Um, and I think we, we were seeing already developers building it and then open sourcing it, and then we reuse it.

  97. 16:20

    Um, but so the community is doing it for us, which is even better. Yeah.

  98. 16:25

    Is that a question? Yeah.

  99. 16:28

    How do you measure and evaluate the visual

  100. 16:32

    consistency?

  101. 16:33

    Uh, you're asking a question that the entire research community in world models has not answered. There are evals for real time and consistencies, uh, and fidelity. Well, fidelity is easy.

  102. 16:43

    It's just like pixels, right? But, um, evaluation for these real-time models is an unsolved problem. So today, it's literally just look at it and

  103. 16:54

    human judgment. That's what it is today. And this is including, by the way, G-DeepMind and everything. Nobody has solved this problem yet.

  104. 17:03

    So what are you guys working on?

  105. 17:04

    Uh, we're working on it. We have a research team. Yes. [laughing]

  106. 17:11

    Awesome. Cool. Well, thank you everyone. Thanks for your time. [outro jingle]