AI Engineer Europe 2026
Black Forest Labs: FLUX, Open Research, and the Future of Visual AI
About this talk
Black Forest Labs developer advocate Stephen Batifol traces the evolution from FLUX.1 to FLUX.1 Kontext, FLUX.2, and FLUX.2 [klein], highlighting contextual image editing, photorealistic multi-reference generation, and sub-second interactive workflows. He frames these models as steps toward broader visual intelligence, including robotics and action prediction, and presents a SelfFlow robot-manipulation demonstration before audience questions.
Chapters
- 0:00Introduction to Stephen Batifol, Black Forest Labs, and FLUX.1
- 2:11FLUX.1 Kontext and iterative image editing
- 3:58FLUX.2, multi-reference editing, and sub-second [klein]
- 17:54SelfFlow robotics and near-real-time visual editing
- 21:25Audience questions about action prediction and world state
Talk transcript
- 0:00
[upbeat music] Thank you for coming today.
- 0:16
I really appreciate it. Uh, thank you for coming to this talk, which is Black Forest Labs, FLUX open research and the future of visual AI. I'm gonna start quickly with a quick intro of myself.
- 0:27
Uh, I'm Stephen Bat-Batifol, sorry. I'm a developer relation engineers at BFL. And I wanna start with two questions. First, who here knows about BFL? Raise your hand.
- 0:39
Okay. Who here knows FLUX? Okay. About the same people, actually. Uh, but for the people that don't know BFL, I have a quick intro, so you're not lost. But BFL at G- at a glance, we are the team behind Stable Diffusion, Latent Diffusion and the FLUX models as well.
- 0:57
Our team has more than two hundred thousand academic citations. And we don't only build models, we actually al- also work with enterprises and customers with them. So some of our customers are Microsoft, Adobe, Canva, Mistral, and many more.
- 1:14
And the way we started is we started in August twenty twenty-four with FLUX.1. FLUX.1 was the first breakthrough, you know. That was the model that was the big competitor to Stable Diffusion back then, and that was really the breakthrough where people were like, "Oh, this is a really cool model."
- 1:30
We released it in open source in the first place. So really that was the one that was text-to-image only and you could run it on your laptop. That was a game changer as well.
- 1:41
And the anatomy was really good in comparison to the other models and especially in comparison to other models that were way bigger. So this is where we really had a breakthrough and Clem from Hugging Face actually gave us a shout-out back then.
- 1:54
This is fairly old, but FLUX was actually the model that was the most liked on Hugging Face back then. This is not true anymore, uh, but back then that was really the thing and that was really good and really big actually for a company that was, you know, coming out of nowhere and just released this model.
- 2:11
We then released FLUX Kontext, which was the first open source editing model in the world that was like the combination of text-to-image and image editing as well. This one, now what I'm showing you here, what I'm showing you here, you know, it's obvious because now we have editing models everywhere.
- 2:27
But back then there was a big breakthrough where you could do both text-to-image and image generation at the same time. If I have an example here, we have this input image and then you remove the snowflake from the face.
- 2:39
You know, you can see you have the character consistency. But you can also then move this person to be in Freiburg, which is where our headquarter is. You know, she's chilling, uh, taking a selfie in the streets of Freiburg and then you can, you know, do some local editing where you can change the background to be snowy
- 2:55
and then you also have snow on her face and everything. It was also the model that was really one of the fastest back then. You know, if you remember, this is the time where you had the first GPT image where it would take like forty, fifty seconds to generate or edit images.
- 3:09
Whereas Kontext, if I remember correctly, was like seven to eight seconds.
- 3:14
It was also really useful to tell stories. I've seen a lot of use cases from our partners, from our customers, where they would start with an input image and then create a storyboard like we see here.
- 3:25
We have the famous seagull here, you know, which has the VR headsets, drinking a beer in a bar. But then you can actually create other things. They can have a friend that is joining and that is then drinking with them.
- 3:36
Now, you know, I guess they got a bit tipsy and they're wearing hats in the bar. Then they're going outside and then, you know, this is a story you could create.
- 3:44
And that was really useful actually for video model or for animation models. You know, you would give those images as input frames or as end frames and then the video model could then create different content.
- 3:58
In November, we released FLUX.2, which is our steps towards what we call visual intelligence. FLUX.2 was, uh, as our base, still our base, our best model, sorry. Um, those are samples which I don't know if you can see clearly, but in my opinion they are like really amazing samples and it's impossible to basically tell they are AI
- 4:18
generated. If you look at the hands, if you look at the veins, if you look, you know, at the bracelets of the person on the left, there's personally no way that I would tell it's AI generated.
- 4:29
Same for the turtles you see on the right side. Same for the dog or cat actually that is in the bath. That would be very, very hard to get the sample like this, um, but those are AI generated.
- 4:42
And then you have more as well. So it's not only the people or animals. You know, you could do like some proper product photography. You can see it with a waffle here on the bottom right or you can make some very cute images like this person, you know, on the left that is riding, driving the moped with
- 4:56
some balloons. And this is what we released in November. But it's not only an image generation, it's also an image editing model at the same time. Um, you can see on the left we have six images that we give to the model and my prompt was literally like create an outfit with those images and then the model
- 5:16
is intelligent enough to actually, you know, make things that make sense. Like the jacket, you know, is worn properly. Same for the tie. And on the right side it's a bit more of a simple use case where you have the sofa and then you have to imagine, you know, maybe you are an e-commerce website or you're a
- 5:34
sofa maker and then you want to people to imagine, you know, what it looks like in your flat or what it really would look like if you were to buy it.
- 5:42
And those use cases are really, really important and those are like the main use cases we have currently for FLUX.2. But it also takes you up to ten images simultaneously, so you can really edit a lot of images at the same time and then you can create magic things.
- 5:57
It's very good at character, product and style consistency.
- 6:03
And what I wanna make clear is that BFL as a company and as a research lab, our first operating principle is to release state-of-the-art models. This is what we wanna focus on.
- 6:14
This is what we wanna do as a company. You know, we wanna raise the bar on quality with every release we do. So we did it in the past, you know, with FLUX.1 when it came out.
- 6:24
We did it with Kontext. Uh, FLUX.2 was our best image model to date. It's the first one we released that was actually multi-reference as well. It was state-of-the-art in the open source world.
- 6:34
It was state-of-the-art for text-to-image and image editing. In January, we released FLUX.2 Klein, which is a step towards, like, interactive editing and interactive generation. It generates and edits images in less than a second.
- 6:50
I'll talk a bit more about it later on during the talk. But I think the fastest it can do, if I remember correctly, it's five hundred milliseconds for editing and three hundred milliseconds for generation.
- 7:00
So basically real time. But this is not it. We also have more things that are coming, and this is where I want to talk about today.
- 7:10
So I mentioned it, we are a research company first. We publish things in the open. We publish paper. We really wanna make sure that the field is moving forward with us, you know.
- 7:20
This is our big focus as well, so it's state-of-the-art model, publishing things in the open, and that's what we wanna do. But I want to first take a step back and tell you a bit, you know, about, like, how do you train models, and especially gener- um, models that are generating content, generating images and everything.
- 7:40
You know, when they generate things, when you train them, they actually don't understand what they're generating. You know, they don't understand that my glass here should be actually on this table.
- 7:50
I shouldn't go through it. Because you train them, you have images, and then, you know, you add some random noise to those images, and then you just try to denoise them.
- 7:59
That's what you do. That's what those models are doing. And when you denoise images, you never learn, you know, that my glass shouldn't go through here. You, like, never learn that, you know, you're sitting on a chair, you shouldn't go through it.
- 8:12
So what do you do is that you use-- you do what is called, like, representation alignment. So you use an external model that actually knows about this and that is an encoder, that is like an image encoder, that is teaching our model, "Hey, he is currently sitting on a chair.
- 8:28
He shouldn't go through it." And those models are external, and they are really, like, trained to segment images, whereas our models are trained to generate images or generate videos or generate audio.
- 8:40
And you try to align them to be on the same objective so that our generative model actually learns, "Okay, you shouldn't go through the chair. My glass should stay on this table."
- 8:50
And this is great because it really improves the way generative models are working. We can see here on the right, you know, it is seventy times faster to actually converge and to reduce the loss when you use this external alignment.
- 9:04
So you're like, "Okay, this is great." But as usual, if something is working well, there are also counterparts to it. So the first one is that you have a scaling ceiling.
- 9:16
You imagine you're working with a model that is external, that has been trained. It's a checkpoint. You're not changing it anymore. What if you train a new model and you have a generative model that you wanna scale up?
- 9:27
You're still, like, limited by this encoder that you have on the side, you know? You're never actually scaling up fully with it. Also, those are specialized in modalities. You have an encoder, for example, DINOv2, and the other one that I can't remember, uh, it's specialized in images only.
- 9:45
What if you want your model to generate images, audio, video and more? You would have to have encoders for all of those. And you can imagine then you would have like a very Frankenstein setup, you know, nothing would really make sense.
- 9:59
And the objectives also misalign, so I've shared it before. We wanna generate content. We wanna generate images or audio. The other one is here to segment things. So, like, they have different objectives, and you're trying to make them work together, and it works great, but it's also not perfect.
- 10:17
Here you see on the right side we have DINOv2 and DINOv3. DINOv3 is a better model technically per se than DINOv2. But when you train your model, actually getting worse performances.
- 10:29
You know, DINOv3 is here in right in red and green, and so you're getting worse performances. So you're like, "Okay, like, this is supposed to be a better model, and yet when I do train a model to generate things, then it gets worse."
- 10:43
And there's also, like, not really any rules as to why, you know, certain encoders should work or otherwise shouldn't.
- 10:51
So how can we solve this? How can you teach, you know, a model representation directly without this external encoder? This is what we released about a month and a half ago now, which is a research paper.
- 11:04
You can read it. It's called SelfFlow. It's in the open. We released it to really make sure, you know, we're moving the field forward again, and it's not only us benefiting from it.
- 11:13
And it's basically a scalable approach to training multimodal generative models. So they use self-supervised learning, so you don't need any other models, you know, to train it. And I'm gonna try to go in tiny bit more details into it.
- 11:28
But we combine representation learning and generation in the same flow. And you see here on the left, you have videos, images, audio. You have different modalities.
- 11:40
What do you do when you usually train a model? You add some noise. You add some random noise. You try to denoise it, and then you align it with the encoder, you know.
- 11:48
How do we do it then? We actually add two different kind of noises that are both random, and they're both different. The first one we're adding is actually we're adding a lot of noise to the assets.
- 11:59
So this is the one you see at the top. And the other one we're adding, like, a low amount of noise. This is what you see at the bottom And the idea is that then we have two models that are actually working together.
- 12:11
We have the student one, which is always getting the images with the most noises and is trying to denoise them, and then the teacher one, which is basically a more stable version of the student's, is always getting the low, um, noises images.
- 12:26
And then the student one is actually trying to learn two things at the same time. It's trying to minimize the loss for the generation and the loss in representation.
- 12:36
And this is how then you actually work across different modalities. You know, this is then you only have one model, you don't have anything external. And if you actually scale up your model, then you're scaling up your student, you're scaling up your teacher, and you don't have to worry about the encoder that you have on the side
- 12:52
anymore. And this is where we're working on. This is something, you know, we are currently using for different models that we're training. Um, and this is, you know, where we believe the future is gonna be and to get rid of those encoders that we have.
- 13:08
We actually trained models. Uh, so those, disclaimer, those are research models. They're not meant to be released in production. Uh, but we released actually one model on all those modalities.
- 13:18
On the left, we're comparing flow matching, which is the usual way of training models, with ours. And you can see we are better in audio, so this is what you see in orange on the right side.
- 13:29
And then we're also better in images, so the dashed lines is the baseline, and then we are like the full line where we can see we are also better at images and also better at video.
- 13:40
So with this approach, without having the encoder and the external model that you may struggle with, you actually get better at every modality that you're training your model on.
- 13:50
It's also converging faster. You can see on the right, you know, the baseline is converging. It's actually hitting a plateau, whereas we are converging faster, and we're still, you know, decreasing the loss.
- 14:02
And I'm pretty sure that if we were to go towards two million steps, you know, the baseline would really plateau and then wouldn't really get any better, maybe actually get worse, whereas we would still go down in loss.
- 14:15
And this is the difference between the two. If you used FLUX in the past or if you used different models, you know, to generate, you may have noticed the text might not be perfect or, you know, things don't really make sense.
- 14:27
This is what you see at the top, where on the left it's like the future is FLUX. But you can see, you know, you have like some letters that are missing, or maybe you have two letters, like on worlds, for example, you have two Ls instead of one.
- 14:40
Whereas with this approach now, in the SelfFlow approach, you can see at the bottom everything makes sense. There's like-- They learn representation. You know, they learn that FLUX then for the letters should be like one next to the, to the other, and the same on the mirror, same on the tree.
- 14:55
And this is where we believe this is the future.
- 14:59
But on top of this, we can see some comparisons here. On the left is a baseline where, again, the letters are wrong, and on the right, you can see that the letters are correct.
- 15:10
Here is the same for the anatomy, where you see on the left, you know, you have like a face. It's looking a bit odd, let's put it that way.
- 15:17
And on the right, this is the one with SelfFlow. And again, this is not like a production model where, you know, you expect the face to be like perfect, but you can see that the anatomy is way better than what you have on the left.
- 15:30
But I wanna show you as well some different generation if it loads. Yes. Thank you. Uh, this is also possible...
- 15:39
This is also possible for video generation. So this is the same model that has been trained on images, now also can generate videos. On the left, you see the baseline.
- 15:49
It's a weird way to do a push-up, let's put it that way. Uh, whereas on the right, it's a perfect form. You know, the arms are correct, the hair as well is correct, and nothing is wrong in it.
- 15:59
And this is, you know, a way to actually fix all those artifacts that you may see usually in generations.
- 16:06
So same here for the birds. Oops, my bird. Yes. Thank you. Uh, it's the same here for the birds where...
- 16:15
Thank you. Where you see on the left side, you have the baseline. There's a lot of flickering. There's a lot of like, you know, weird things happening because the model was using this encoder or was trying to align things.
- 16:26
Whereas on the right side, with SelfFlow, it just does it perfectly and like the, the bird, you know, is walking on the floor, and there's no flickering or anything.
- 16:35
But it's not only about images or videos or audio. You train those jointly, so you can also actually generate things jointly. We have here an example of a video and audio sample where the idea is that we have someone that is saying hello from the Black Forest.
- 16:53
Uh, I will just play them, and you will hear the difference. Again, this is not a production-ready model, so it's not like perfect, but you can hear the difference, hopefully.
- 17:02
Hello from the Black Forest. Hello from the Black Forest. Hello from the Black Forest.
- 17:07
And so this one was the baseline where if you hear it correct-- if you try to pay attention to what he's saying, you hear like, "Hello from the Black Forest."
- 17:15
Ooh. There's a bit of like weird things at the end.
- 17:19
Hello from the Black Forest. Hello from the Black Forest.
- 17:22
Whereas on the right side, you can see, you know, the prompt is really just say, "Hello from the Black Forest," and then it ends here. And yeah, this is the same model that was trained on those images that we've seen before on video and on video and audio.
- 17:35
Hello from the Black Forest.
- 17:38
Thank you. But this is cool and this is great, but what if you could also teach robots on how to use this? This is also the same model. This one is trained on actions and not only on images, video or audio, so it can also predict actions.
- 17:54
And what I'm gonna show you now, it's a robot that is trying to pick up a can and make it closer to us.
- 18:01
On the left, this is a baseline. Again, you see some like flickering. You see like the, the arm is doing weird things. Whereas on the right, for the same amount of steps, you can see SelfFlow, the robot is picking up the arm directly and like bringing it closer And this is where we're going as well as a
- 18:16
company. This is where we're really interested is, like, not only image generation or video, but it's also doing actions and doing more things towards physical AI.
- 18:26
And there's more. It's also how do we make our models faster because this is really important for us. This is a demo of Klein, which is, you know, like near real-time editing.
- 18:37
You see it on the right side. This is, you know, generated with Klein on Krea, where you see the edits. And this is not a video model, this is-- those are images that are always editing in real-time.
- 18:49
Not only they are faster, they're also actually at least on par or better than other models. And I'm almost out of time, or they added five minutes, so I don't know if I'm...
- 18:59
Okay, cool. So I'm not out of time, so I can chill. Uh, so yes, here on the left you can see, you know, we have Klein, uh, that is four B and nine B that is compared to the other open source models.
- 19:10
So it's at least on par, while the latency, you know, it's like zero point five seconds, while Qwen is like around like fifteen seconds. You know? And if you are like on par and you're like way faster, then this is really, really good for us.
- 19:24
Same for image to image. You can see the editing for Klein nine B, we are at like a tiny bit more than zero point five seconds, whereas Qwen is still at around fifteen seconds.
- 19:34
And then same for multi-ref, you add it but we're still at less than a second, whereas Qwen is more towards the twenty seconds. And this is what is really, really important for us because you really want to actually generate things in real-time.
- 19:48
This is where we believe the road to world to visual intelligence, this is where we're going as a company in the future. And why does it matter? 'Cause like I mentioned it, real-time generation.
- 19:59
So you can imagine you render mock-ups as fast as you think. You know, you don't have to wait, you don't have to wait like ten seconds, twenty, a minute or two.
- 20:08
You do things in real-time, and you can guide them in real-time. This is also where we're going. You can think, you know, interactive visual engines for gaming or films where you really render a movie as you prompt it.
- 20:21
On top of this, there's also world models. The idea of world models and behind it and why it matters for us, it's you train your models to understand and simulate geometry, the relationship and like different interaction of the world.
- 20:36
And you may be like, "Okay, that's cool from a research perspective. Why do we care?" The reason is robots. Uh, that's why we care. That's why, you know, robotics and automation, this is where we're taking BFL, and that's why we also wanna go towards world models, is to train agents in those generative world to scale self-driving and
- 20:55
automate every manufacturing. And I think that is it. Thank you very much. [audience applauding]
- 21:05
Do we take... Yeah, I think we can take questions.
- 21:09
Can you share something where you base your training off for the world models?
- 21:14
Uh, you mean the data?
- 21:15
Yes.
- 21:15
Can't really. Uh, this is trade secrets. I mean, data is very sensitive, as you can imagine, so I can't really share this. We're partnering with a lot of people though for it.
- 21:25
Uh, how do you store the state of the world in those action prediction models?
- 21:29
Mm-hmm. Well, this is-- What the model is learning is basically like those representation. You know, the model is learning that in itself as like the state and it has like some kind of memory, and this is, uh, the way we do it.
- 21:41
What is that some kind of memory? Is it like in the, in the context window or is it external and you gather it or-
- 21:47
No, it's... Yeah, it's the context window that you have. You know, you train and then you have the tokens and then they'd be like, "Oh look, I've moved. Here is where I should be then next."
- 21:55
And can you then run it for long or does it do some compaction?
- 22:00
Uh, define long. What do you call with long?
- 22:03
Uh, indefinitely.
- 22:05
Indefinitely? Uh, that I'm not sure. I mean, there's always gonna be a limit. Uh, so you may have, you know, like a sliding window. Um, but this is the way we show it.
- 22:15
Thank you. [upbeat music]