AI Engineer World's Fair 2025
The State of Generative Media Today
Read the talk
Generative Media: From Surprising Images to Interactive Experiences
Gorkem Yurtseven traces the market for generated images, audio and video from fal.ai’s perspective, connecting broader model access with personalized advertising, visual shopping and interactive video.
From a talk by Gorkem Yurtseven
When extraordinary images became accessible
In 2022, Gorkem Yurtseven was sitting on the floor at home, watching Sam Altman reply to people’s requests with DALL·E 2 images. The pictures seemed extraordinary. Even with experience in the industry, he thought OpenAI had established a lead that would be very difficult for anyone else to close. Looking back, those same images already seemed low quality by the time of this talk. His perspective comes from fal.ai, a generative media platform whose inference engine serves image, audio and video models, alongside partnerships with closed-source providers.
There had been earlier waves of excitement: GANs, Google’s DeepDream, and consumer applications such as Prisma, which transformed uploaded selfies into stylized portraits. Their applications were narrower, but the ambition to make art with computers was much older still. Yurtseven shows a recreation of Harold Cohen’s work, with a large machine drawing on a huge canvas. Computer graphics and generative graphics belong to that longer history of using computation to produce visual art.
The apparently unassailable lead in image generation did not last. Yurtseven’s release sequence begins with DALL·E 2 on April 6, 2022, followed by Midjourney through a Discord bot and the open release of Stable Diffusion. That last step changed who could participate: people could run comparable technology on home GPUs and build services around it. SDXL followed, then a growing mix of open and closed models, including FLUX, released the summer before the talk. Access expanded from watching someone else generate images to operating the models and building products around them.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Cheaper creation changes advertising’s economics
The marginal cost of creation can fall without eliminating the value of creativity. Yurtseven deliberately separates the two. Storytelling and creative decisions still matter; once those decisions are in place, producing the next artifact becomes cheaper. He expects that shift to affect social media, advertising, marketing, fashion, film, gaming and e-commerce, eventually reaching content of every kind.
Advertising offers a precedent for software expanding a media market. Yurtseven claims that YouTube’s advertising revenue exceeds the revenue of other media companies except Disney. He points out that Disney’s total also includes parks, cruise ships and other non-media businesses. The comparison motivates his expectation that another technological change could increase advertising volume, although it is not presented with a consistent accounting basis across companies.
Yurtseven says advertising grew threefold since 2000, with all that growth coming from online or software advertising. The talk does not specify the geography, end year or inflation basis for that estimate. His forecast follows the same pattern: advertising will be among the first industries affected by generative media at scale, and AI will supply much of its future growth.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
From ad variants to audience participation
Lower production costs allow personalization at several levels:
- Demographic variants: Generate versions of the same ad for a hypothetical 10,000 demographic groups.
- Individual context: Generate an ad on arrival, using context such as the website a visitor came from.
- Interaction: Let the audience’s actions influence what gets generated, making the experience more than a fixed creative asset.
These approaches change both the number of ads a campaign can produce and when production happens: some content can be generated in response to the person viewing it.
Advertising can also absorb a volume of content that other media cannot. Adding a hypothetical 1,000 movies does not give a viewer more time to watch them. Ads occupy recurring opportunities on phones and television, where a different creative can appear on each encounter. Some consistency still matters, but repeated ad impressions create room for far more variation than a person’s finite movie-watching schedule.
Yurtseven describes a campaign fal.ai worked on the previous year for A24’s Civil War, a film about an imagined civil war in the United States. Its little green toy soldiers became a personalized audience experience:
- A participant submitted a selfie and a description through a live marketing website.
- The system generated a little green toy soldier incorporating the participant’s face.
- Participants’ personalized soldiers appeared on a display in Times Square.
The useful mechanism is the connection between an audience input, a generated representation and a public display. Instead of merely selecting which ad to show, the campaign let people become part of the creative itself.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Virtual try-on meets an existing shopping need
E-commerce provides another route into an existing market. Yurtseven describes it as gaining roughly “1%” of US retail each year, independently of AI. His wording concerns retail share but does not distinguish percentage-point gains from percentage growth. The relevant direction is the continuing movement of shopping online, where a highly visual experience gives generative media opportunities to add interaction.
Virtual try-on is his clearest example. By the time of the talk, he had seen it develop over several months to perhaps a year into one of generative media’s strongest early examples of product-market fit. Retailers were adopting it, and startups were building around it. Rather than asking consumers to find a use for a new model, try-on puts generated imagery inside an activity they already perform: evaluating something they might buy. That is why he sees every retailer and e-commerce website as a potential generative media user.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Video repeats the breakthrough—and the catch-up
When Sora appeared, Yurtseven recognized the pattern from DALL·E 2. The February 2024 research reveal looked, if anything, even more impressive. This time, however, he did not interpret OpenAI’s achievement as a durable barrier to everyone else. Having watched competitors catch up in images, he saw a demonstration that comparable video capabilities could emerge elsewhere too.
Yurtseven reports that video was barely represented on fal.ai in October, reached 18% in February and was around 30% when he checked the day before the talk. He introduces these figures as a revenue snapshot but also describes model usage, so the denominator and the October year remain unspecified. He treats fal.ai’s trajectory as a proxy for the wider market: video was growing quickly even while generation remained expensive and results imperfect. From that observation, he predicts video will eventually dominate generative media.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Why video could become a much larger market
Yurtseven’s market-sizing argument combines the cost of producing video with an assumption about its usefulness and engagement. He explicitly presents it as rough math:
| Factor | Estimate or assumption |
|---|---|
| Compute | Video models require roughly 20× more compute. |
| Engagement | Assume video is 5× more engaging. |
| Industry reach | Video serves more industries and use cases. |
| Market forecast | Generative video becomes 100–250× the size of image generation. |
The compute estimate has no specified models, hardware, resolution, clip duration or workload, and the engagement multiplier is hypothetical. Multiplying the two stated numerical assumptions gives the lower end of his forecast; the talk does not provide a numerical derivation for the upper end. These are expectations about market scale, not measured model performance or realized market size.
That forecast does not require image generation to stop growing. Yurtseven expects substantial image-market growth over the following years, while predicting that video will grow faster and ultimately become much larger. The distinction is between relative market sizes, not between a growing medium and an obsolete one.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Sound, then the possibility of real-time generation
The next source of growth is capability, not simply volume. Google DeepMind’s Veo 3 illustrates the progression: first video models improved consistency, then they added sound. Each addition can make a previously impractical use case possible. At the time of the talk, Yurtseven said Veo 3 was not yet on fal.ai; he was looking ahead to what people might build with it in advertising and e-commerce.
Beyond richer clips lies a different operating target: generate one second of video in one second. Yurtseven predicts that faster, cheaper generation will eventually reach that rate, allowing generated content to stream to a user. The change matters because the output could become an ongoing experience rather than a clip produced before playback.
If generated video can respond during an experience, the boundary between a movie and a game becomes less distinct. Yurtseven poses social apps and live events as open possibilities. Fortnite already hosts live events; more lifelike generated environments might extend that kind of participation to parents and other people who do not normally play video games. The question is not only how realistic the imagery becomes, but who can participate and how they interact with it.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Image editing opens another path into enterprises
Image models are still adding capabilities too. FLUX.1 Kontext and GPT-4o image generation bring new editing capabilities and better text rendering. Those improvements challenge the idea that image generation had reached a plateau: a model that can revise an image or render its text more effectively can address work that a visually impressive first draft cannot.
Yurtseven associates these capability shifts with adoption by more mature businesses. As editing and text rendering improve, he expects generative images to fit more enterprise use cases. He closes by inviting people to help build that infrastructure: fal.ai is recruiting machine learning, inference and product engineers, among other roles, and he offers to discuss generative media with attendees during the rest of the day.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
From the talk
The original announcement explains instruction-based image editing, character consistency, typography and known limitations.
OpenAI's historical launch post illustrates text rendering, conversational image refinement and transformations of uploaded images.
Introduces Veo 3's generated audio, including ambient sound and dialogue, and its initial access channels.
Further reading
The original Sora technical report describes its video representation, architecture, qualitative capabilities and limitations.
Historical Census estimates distinguish e-commerce sales growth from its share of retail sales. These estimates were subsequently revised.
Read the complete timestamped transcript
- 0:00
[on hold music] It's so [chuckles] nice to see a generative media track in the AI conference, AI engineer conference this year.
- 0:22
Um, we... My company, I work at this company called fal.ai. Uh, we call ourselves a generative media platform. Uh, and this is a term that's been around for a while, uh, but we kind of owned it, and we called it, uh, the name, name of our, of our company, Generative Media Platform.
- 0:42
Uh, and the way we define it at least is it, it's a generative video, audio, or image and, uh, our company is seeing all these kinds of models u-using our inference engine, and we are partnering up with some closed source model providers as well.
- 1:01
So I've been doing this for a couple of years, but it is a really, really new market and, uh, throughout the talk, I'm gonna walk you through how we got here and a little bit of the history and what's next.
- 1:18
I remember in 2022 when Sam Altman started tweeting, uh, about DALL·E 2, um, I was working at home. It was end of COVID. I know COVID took a little longer in San Francisco, but I, I remember sitting on the floor, could not believe my eyes.
- 1:38
People were tweeting at him, and he was tweeting back pretty high definition images of incredible things that people, people were tweeting. Like, looking back to it, obviously it all looks kind of bad quality, but I remember at the time I was ...
- 1:59
I thought this was the most incredible technology ever. And I was, I was in the industry. I, I knew what was going on, not as much as today, but I, I thought at that time, OpenAI was so far ahead of anything else, and it's gonna be so, so hard for, uh, normal people to catch up to this
- 2:17
technology. Uh, I was, I was... I remember this was one of the biggest WTF moments of my life.
- 2:26
But then you can tell me, "Hey, Gorkem, this was, this was all gonna happen." Uh, there was other AI waves before, [chuckles] before th-th-this last big wave. Uh, there was a GAN breakthrough that people did similar things using GANs.
- 2:43
Uh, Deep, DeepDream from Google, uh, went through a phase, and then there was even a viral, uh, consumer AI application of it called Prism. People uploaded their selfies, and they were able to change their avatars.
- 2:57
But, uh, the capabilities and the applications of the technology was not nearly as much as what generative media can be used today.
- 3:08
Not only the previous AI wave, generative media or being able to create art with computers has been around kind of since the computers has been around. Um, this guy, Harold Cohen, this is a recreation of his project, but basically, he created these massive computers to dr-draw on these huge canvas, uh, to, to
- 3:33
create art similar to how a human would draw. And then we have, uh, computer graphics and generative graphics, things like that. Uh, throughout the years, people tried to generate visuals and art using different computing technologies all along.
- 3:52
Uh, right after Sam Altman's tweet, uh, playing field evened out really, really quickly. So DALL·E 2 was April 6. Right after that, Midjourney released their f- initial model in beta as, as a Discord bot.
- 4:09
And then very quickly after that, uh, Stable Diffusion
- 4:14
open sourced their model, which was a huge, huge thing. People now were able to run a technology similar to DALL·E 2 in their homes, in their home GPUs. People started building services around it.
- 4:27
And then SDXL came out, and then now there's many different image models, uh, open and closed source, and most recently, FLUX was released, um, early, uh,
- 4:40
in, in the, in the, in the summer last year.
- 4:44
And with all the, this playing field evening out, the, the marginal cost of creation is approaching zero. And I'm, I'm very careful when I choose my words here. I'm not saying marginal cost of creativity.
- 4:59
It's marginal cost of creation. I think the storytelling is still really important. Creativity is still really important. But once, once you have that set up, creating that next new thing is, is approaching zero, and we believe this is gonna have huge impacts on different kinds of industries and, um, markets.
- 5:22
So anything from social media, advertising, marketing, fashion, obviously film and movies, gaming, and e-commerce is gonna be transformed by generative media, and this transformation is gonna continue until all content, one way or the other, is, is impacted by AI.
- 5:45
Um, so if you've been following, software has been eating media all along. Uh, YouTube, just from Basically ads, uh, is generating more revenue than any other me- media company except Disney.
- 6:00
This is, this is pretty remarkable. And with Disney revenue, there is clearly non-media revenue in there. They have parks, they have cruise ships, they have other things. So it's not too hard to say YouTube right now is one of the highest revenue-generating media companies in the world, and it is happening through, through ads.
- 6:18
And whenever [chuckles] a ad industry is, is impacted by technology, it usually grows in volume. So we believe the same thing is gonna happen with, with generative media and ads.
- 6:31
Uh, we believe ad industry is one of-- is gonna be the first industries to be impacted at a large scale by generative media. We believe the, the size of the industry is gonna increase.
- 6:43
So it's really funny. In-- Since 2000, every ad spend has been online. So a- ad industry grew three, three times since 2000, but all that growth happened, happened in software ads.
- 6:58
So we believe something similar is gonna happen with AI, uh, driven ads, and ad industry is gonna grow, and most of that growth is gonna come from AI. And there are s- s- several different ways how this can happen.
- 7:14
We believe ads themselves are gonna become hyper-personalized, so this might mean, um, you are generating many different versions of the same ad but maybe 10,000 different demographics really quickly, or it can, it can also mean it's targeted towards a certain individual.
- 7:32
If you are coming from a certain website, then the ad can be generated on the fly. Uh, and then it could also be, be interactive in, in ways, you know, that, that I just mentioned.
- 7:43
Things can be generated on the fly, and that, that might mean many different things, uh, in the, in the industry. One other thing why I think generative media fits the ad industry very well is, uh, the abundance of content.
- 7:56
For example, I, I probably won't watch a blockbuster movie every single day. So like even if we have 1,000 more movies this year, I, I probably have to sit down and watch a movie a day to go through all of them, but I probably won't be able to do that.
- 8:13
But ads, there can be kind of unlimited content. Every time I'm grabbing my phone, I'm seeing ads. On TV, there are ads all the time. And it doesn't matter if the ad is different.
- 8:25
Like maybe there, there needs to be some consistency, but ad industry can, can actually survive with, with a lot more content and things can get a lot creative.
- 8:38
So we, we were ahead of the time a little bit. Last year we did a, we did an ad, ad promo with A24 Civil War movie, and it was one of those interactive ideas that I was, I was talking about.
- 8:53
So if you've seen the, the, the movie, it's about, uh, a s- imaginary civil war in the US, and they had this campaign of these little green toy soldiers, and they created a live marketing site where you could put a selfie and then, uh, we, we created a little toy soldier o- with your selfie and your description,
- 9:16
uh, and they put this on Times Square. People were able to display their own faces on these little, uh, green toy soldiers. So AI is gonna help us create experiences like this that are interactive and personalized.
- 9:32
The, the other trend we are watching really closely is e-commerce. Um, if you've been paying attention to it, e-commerce is growing about 1%, uh, every year, getting a percentage of the US retail industry.
- 9:47
So this is a trend that's happening with or without AI, and we believe generative media is gonna play a big, big part on e-commerce's growth a- as well. Um, it's, it's already there are many companies trying to redefine how people shop online, and because online shopping is very visual, AI can add a lot of interactivity
- 10:12
to the experience. In fact, it's one of the earliest product market fits I've seen in generative media. This has been happening for a couple months, maybe a year, that virtual try-on is, is one of the clearest product market fits that I see in the, in the AI industry.
- 10:31
Many different retailers, e-commerce websites are adapting, adapting this technology. Many different startups are being built, uh, on it. So I believe this is gonna be everywhere. Every retailer, every e-commerce website is a potential, uh, generative media user.
- 10:49
And then there is video. Um, so when, when Sam Altman tweeted DALL·E 2, I thought OpenAI was so far ahead and no one was able to catch up. People caught up incredibly fast.
- 11:02
So this time he did the same trick with Sora when Sora was released a year and a half ago, basically. And this time around, maybe Sora was even more impressive than DALL·E 2 in terms of, uh, how far ahead things look like.
- 11:18
But, uh, this time around, I was incredibly excited that researchers at, at OpenAI was able to actually do things like this. And, uh, from, from the past experience and I know that if this is possible in, in one place, uh, eh, people are gonna go be able to do similar things in others.
- 11:37
So I was incredibly excited when Sam Altman started tweeting about Sora because I knew very soon, uh, a technology like this was gonna be everywhere. I- in fact, it started happening.
- 11:50
So this is a little snapshot of Our company's revenue, which I think is a good proxy of, of the entire market, uh, early this year in October, we barely had any video model, uh, usage in the platform.
- 12:06
And in February, this, this went all the way up to eighteen percent. Um, I didn't get time to update it, but I looked yesterday. It's around thirty percent today.
- 12:18
So it is growing really fast, even though it's expensive, even though it still doesn't work as well. Uh, video models are gonna completely take over the generative media market.
- 12:30
And I have some predictions about how much bigger the video market is gonna be compared to the image market. So rough math, but we believe video models are twenty x more compute intensive.
- 12:45
And let's say if it's five x more engaging and it's gonna impact more industries because it's gonna be more useful to the industry, we believe all said and done, the video market is gonna be gen- generative video market is gonna be one hundred x to two hundred fifty x, uh, bigger than the image generation market.
- 13:06
Uh, and we are just, just scratching the surface here. I believe the image generation market has a ton of growth that's gonna happen in the, in the next couple of years as well.
- 13:16
But video is growing much, much faster than that. And when all said and done, it's gonna be a much bigger, uh, market.
- 13:27
And yeah, video models are leveling up as well. Uh, you probably have all seen the, the newest model from DeepMind, from Google, Veo three. Uh, we keep adding new capabilities, uh, into the video models.
- 13:42
First it was consistency and then now with sound, uh, really the things people are generating with it is, is, is incredible. And every time a new capability is added, it unlocks a different use case in the, in the industry.
- 13:59
So, um, it's, it's not on our platform yet, but I'm, I'm very curious to see how people are gonna start creating using Veo three and what, what different use cases it's gonna unlock in the ad industry or, or, or the e-commerce in- industry.
- 14:15
So that is very interesting to see. So where is, where is the video market going? Uh, we believe there is so, so much to improve. Um, we are gonna have faster and cheaper video generation until video generation basically becomes real time.
- 14:36
Uh, so generating one second of video in one second. So you, you'll be able to stream generated content, uh, to the user, and this is gonna have very different implications on how people interact with this, this technology.
- 14:52
Everything, uh, potentially becomes interactive. The line between games and, and movies, uh, gets blurred. Um, so how is this gonna impact social apps? How it- it's gonna impact live events?
- 15:09
Uh, people, like if you play Fortnite or similar games, people are already having live events there. Is it gonna become, uh, more lifelike? Are more norm-- like the-- are parents like, you know, people who are not used to playing video games are gonna be part of this experience?
- 15:27
I'm, I'm really curious about the, the future of this technology.
- 15:32
And then image models are not that done yet as well. Um, there's, there's been a lot of different, uh, improvements in the past couple of months on the image models as well.
- 15:46
Uh, FLUX.1 Kontext and GPT-4o, uh, introduced new editing capabilities, better text rendering capabilities. At, at one point people thought, "Okay, maybe this is as good as image models are gonna get."
- 16:01
But with, with these new releases, uh, and new capabilities, it is opening up to more use cases in the industry. Whenever we see a, a, a technological shift like this happening, we see a lot of different, um, more mature players in the industry picking up, uh, th- these technologies.
- 16:22
So we believe something similar is gonna happen with FLUX.1 Kontext and GPT-4o, and it's gonna blend into more of the
- 16:33
enterprise use cases people are, uh, trying to do. Um, and then this is, this is pretty much it. Um, we, we are hiring, so please visit our website fal.ai/careers.
- 16:48
Uh, we, we are hiring machine learning engineers, inference engineers, product engineers, uh, all sorts of positions. And I'll be hanging around rest of the day today, so find me, talk to me.
- 17:00
Would love to discuss whatever, uh, related to generative media or about the industry in general. Uh, thank you so much. [upbeat music]