AI Engineer Europe 2026
20 days of compute vs 7 hours: rethinking what state-of-the-art means — Bertrand Charpentier, Pruna AI
About this talk
Pruna AI cofounder and chief scientist Bertrand Charpentier argues that state-of-the-art image models cannot be identified from a single leaderboard or quality metric. Comparing Design Arena, Arena, and Artificial Analysis, he shows how rankings vary by benchmark and use case, demonstrates limitations of human inspection and CLIP scores, and highlights the compute and energy costs of large-scale evaluations. He advocates assessing specialized, compressed models on a Pareto front that balances task-specific quality against latency, cost, and efficiency, followed by audience questions about image-model and video-model compression.
Chapters
- 0:00Defining state-of-the-art beyond public leaderboards
- 2:34Why image-model leaderboards disagree across tasks
- 7:24Internal benchmarks, audience preferences, and CLIP scores
- 12:28Evaluation compute, energy, and faster image models
- 15:41Pareto-efficient FLUX optimization and text rendering
- 17:41Audience questions about image and video compression
Talk transcript
- 0:00
[upbeat music] Hello.
- 0:15
So, today we gonna try to ask, uh, to, to ask us the question: What AI model is state-of-the-art? And I guess this is a super important question for everyone because, of course, for our applications in research or when we deploy a product, we want al-always to have the m- the best performance out of our models.
- 0:32
But the problem is that state-of-the-art is a bit a confusing concept, and people maybe have different vision on this. So we'll just try to see a bit, like, first what, what...
- 0:42
how people approach this question and how they try to answer this question. And usually, there are two main methods that people try to, to use to know what m-model is state of the, uh, state-of-the-art.
- 0:52
The first one is simply, you know, they go on the, the Internet and check some public leaderboards, see, uh, what, what is the best model, uh, on, on the public leaderboard.
- 1:03
And another method is actually just perform some internal evaluation and again, see, based on their internal evaluation, like, what model is the best. The problem with these methods is, like, in most cases, if you apply them naively, you will always find, like, a kind of lazy solution, which is just to, to use a large foundation, uh, fo-foundation
- 1:24
model. So we're gonna just try to see a bit with these methods, like, what people try, uh, uh, tend to do, like, a bit, um, quick and can, uh, can, can be done better.
- 1:36
So the first method, again, is, like, just simply checking a public le-leaderboard. So for example, if you take the use case of, like, uh, image editing, let's try to find the best image editing model in this case.
- 1:47
So usually, the first step is just find some leaderboard. In this case, you can use, like, Design Arena, which is a very famous one. And then you just pick the top one, which is ChatGPT image, and then you feel that you are happy.
- 1:59
This is the best model for your use case. In general, it's, like, kind of a good solution. Like, you, you get, like, a first, um, reasonable, reasonable model at low effort.
- 2:09
But the problem is that, [lip smack] um, you don't know exactly a lot of things about how the users will interact with your model and so on, so you can actually make a much better choice.
- 2:20
So the first problem is that if you look at many leaderboards, but not a single one, you will see that each public leaderboard will have a different ranking. So here, maybe it's a bit small, but you can trust me, there are three leaderboards:
- 2:34
LM Arena, now called Arena, Design Arena, Artificial Analysis, and they try to rank image editing models. And if you try to draw a bit the difference between the, these leaderboards, you will see that it's not the same ranking.
- 2:47
Like, the top model is not the same. Also, like, relatively, if you compare model to be- between each other, they will be different. For example, there is one model, Hunyuan, um, that goes from rank ten on Artificial Analysis and is ranked five on, like, uh, uh, Arena.
- 3:05
So the idea is that each leaderboard add a different perspective, and sometimes there are also some models that have duplicate entries, so it's a bit noise, a bit what is the information you want to get out, uh, get out of it.
- 3:18
There are even some models that appear in the leaderboards and are not in some others. So it's hard to get the main information. And when you check actually in details like the Elo scores, so which is supposed, uh, to be the quality score you use to know what is the best model, we'll see that even these Elo
- 3:34
scores, they are very different. Meaning that for some, uh, area, uh, for some, um, leaderboards, it will be between one thousand one hundred to one thousand three hundreds, but for some, it will be a completely different range.
- 3:47
So relatively, models, we don't know how strong they are between each other.
- 3:52
And usually, the main solution for this is not to trust a single one, but you really to look at multiple one. And when you see that there is a lot of difference between, like, different rankings, it means that probably there are, like, some models which are approximately equivalent.
- 4:05
It's not because ChatGPT image is ranked to-top one on one, uh, on one leaderboard, one leaderboard, that it means that it's the best o-overall.
- 4:15
Another problem is, in most cases, you will have applica-- you, you will
- 4:20
release and use your model for specific application. What we've seen before is, like, some aggregated score over a lot of different task. I don't know, uh, for example, we'll, we can see, like, removing object, changing background, editing text.
- 4:36
But we can actually build some leaderboards for each of these specific use case. And this is also some leaderboards from, like, um, [lip smack] Design Arena, I guess. And
- 4:47
I think there is a problem with the... Okay, it, it comes back. Um, and you can see that actually, if you draw the difference for each specific use case, you will see that again, the ranking, they are completely different.
- 4:58
And ChatGPT image even i-is never to-top one i-i-in this, is th-- in this ranking. There are always some new model and some models which are super good at re- at removing objects, some models which are super good at doing some other things.
- 5:11
The idea is there is no model consistently o-outperforming the others, and, uh, there are very different models working well for d-- for different target use case. And this is normal because this is just due to, to the fact that, you know, s-some models, they have been, you know, for example, trained more on some specific task than others.
- 5:30
And the solution for this is when you check, like, public leaderboard, you should always try to target what your use case will, will do in the end. Like, if you focus on removing objects, look at this leaderboard and not the others.
- 5:44
Another problem is that usually leaderboards, they are not really statistically significant for your specific use case. So here, I try to show two d-- two different things. So, a fair thing is on how many samples these leader-leaderboards are built.
- 5:58
And if you check on the left, for example, Artificial A-Analysis, this i-- all these, um, these rankings, they are built on, you know, few s-- um, thousand samples for each of them.
- 6:09
So it's not much if you compare to the, the load of inference you have for many applications. It's probably super low. Uh, for some of, of our, of our models that we have, we have millions of inference per day, so probably we'll get more information by just, like, just training the model on our API rather than just,
- 6:26
uh, looking at this, um, uh, at this leaderboard. Another thing is Elo scores. Usually, you can also compute what is the win rate of each model. So when you build this c- the, the, these rankings, what you, uh, what you do is you actually make models battle against each other and ask people, "Okay, what is the best,
- 6:43
um, what is the best, uh, model between the, the two?" to, to a lot of users. And what you can see is actually the win rate, usually there are no models which are close to one hundred percent win rates.
- 6:54
It means that most of the models, they lose at least forty percent of their, of their battles. And if your use case is in this forty percent of the battles, it means that you will just-- if you take the best models, you will just, uh, take the wrong model.
- 7:10
So again, here it's important really to evaluate on more samples and always, like, have, like, uh, evaluation which is close to the final, uh, setup, um, the final, uh, use case conditions.
- 7:24
Now we can check also the second solution, uh, to, to, to try to know what is, uh, the state of the art for, uh, AI model. And the second solution is to do just internal benchmark.
- 7:36
One way to do it is what I see the most, uh, actually in image and video, uh, generation,
- 7:44
um, research and so on. People just do manual insp- inspection. They try a couple of prompts, a couple of models, and they, they get a feeling intuitively a bit what is the best model.
- 7:53
Another thing that sometimes people do is they just, like, run some benchmark, automated benchm- benchmark out there, and then try to see, okay, based on this benchmark, which is the one that w- that has the best performance.
- 8:05
So there, yes, basically then you, you just select the preferred model. So it can be, I don't know, for example, the third model or the one with the highest score.
- 8:14
The problem is that-- So there are a couple of problems with this, and we can start this with a little game. So here I'm just gonna show, like, three images.
- 8:23
And maybe one question for you is, how many people in the room prefer the first im- prefer the first image among these three?
- 8:32
What was the requirement?
- 8:34
So the, the-- This is a question, like, in general, you can ask a lot of questions. So does it, uh, adhere the prompt? Is it, is it what image do you prefer?
- 8:42
And so on. But-
- 8:44
What was the prompt?
- 8:44
The prompt was I, I think a little guy and a parrot or something like this. And yeah.
- 8:49
I like the first.
- 8:50
Okay. First. Who prefers the second image? Okay. A couple of people. Who prefers the third image?
- 8:58
The second image.
- 8:58
Okay. So what is great here is that we have seen that people have different preference. So it's important to see that if you do, like, manual inspection, you will be super biased to your own preference.
- 9:12
Right.
- 9:12
So it's very important to not trust only your preference, because then you have big surprise that actually it's not the, the, the models that are preferred by, by everyone.
- 9:21
Now we can do it again. Same question. Who prefers the first image in this case? I think the prompt was, like, probably a man eating some soup with past-- with, with pasta or something like this.
- 9:33
Okay. Who prefers the second image? Okay, great. And who prefers the third?
- 9:42
Okay. So that's also super interesting because I've seen some people changing their minds. So always on the left it was the, the, the Sidri model, middle Flux.1, and on the, on the right, like, uh, with some models we developed one image.
- 9:57
And the idea is, like, also you are super biased toward the few samples that you look at. So when you do manual inspections, you are two times biased by you and by also the number of samples, the specific samples you, you look at.
- 10:09
So, so in general, the idea is, like, you should never only trust the ma- the, the, the, um, manual inspection. It's good to get a feeling, but it's not enough.
- 10:19
You should always ask many people, uh, to do it. And human evaluation is usually great, but you have to scale it, um, properly.
- 10:28
Another problem is that when you do now not human evaluation, but more, like, proper, like, uh, automated evaluation with, uh, with metrics, sometimes you have, like, non-consistent results. So for example, this is a bit small, but you can trust me.
- 10:44
We ranked like eight models regarding some metrics, like a very standard metrics which is called CLIP score. And sometimes people, when they try to evaluate image models, they do-- they check, uh, this metric first.
- 10:56
And you can see actually that if you check, like, the rankings for, uh, the three, uh, metrics we look, like CLIP score on different datasets, it change all the time.
- 11:05
And these metrics are supposed to be between zero or, uh, zero and one, or zero and one hundred. And actually the variations between models, they are super small. So it means, it means that it's hard to know from this metrics which is the best model.
- 11:19
What you should do is actually first having some clear understanding of what the metric does, and also use multiple m- multiple of them.
- 11:29
So here, for example, this is another type of metric. When you know your, you know your use case, for example, you know to, you, you know you want to be the best at text rendering, you-- there are a lot of text rendering metrics that would be better to evaluate your models.
- 11:42
So here you can see again, like, the ranking is way more consistent. You have always Z-Image being the first and P-Image be- being the second model. And also the variations, they are way more signi- significant.
- 11:53
So the models are supposed to be zero, uh, the metrics are supposed to be between zero and one, and there are, like, clear difference, uh, between, like, every, uh, every model.
- 12:04
So yes, in general, very important. Understand your metrics. People usually tend to just use some metrics and say, "Okay, I did my benchmark," and then I stop here. But it's important to understand what you actually measure with this.
- 12:18
And now a l- a last problem, which is actually common to the, the, the, the first and second method to, that we're re- that we've seen before, is that usually quality is driven by compute.
- 12:28
So here, this is ChatGPT image. And for the evaluation, uh, Design Arena, uh, I think, or maybe it's, uh, LM Arena, um, they did like 27, uh, 26, uh, K, uh, battles.
- 12:43
So it means they generated 20, um, 26K, uh, images. And each of these image takes one minute to generate. So here I summarize all this information. S- sixty-two seconds, uh, 26K, uh, evaluations.
- 12:58
And in total to do these 26K evaluations it takes 20 days of compute. In terms of cost, it's 5K, uh, just 5K just to, to, to run this evaluation.
- 13:09
And in terms of energy it's approximately, you know, 556 kilowatt, uh, kilowatt hour. So I know that people might not have the order of magnitude of what it represents, this amount of energy, so just to give, like, some idea, I check my Strava and check how much energy I was consuming by running a marathon, and actually it
- 13:29
represent 400 marathon just to generate all these images. So it's a lot. Uh, I'm tired after one marathon, so I don't want to do 400 for sure.
- 13:38
Now there are some alternative. You can use some different models. So of course this is a model that we've, we've done that does, like, um, time gen- generation-- editing of, of images in less than one second.
- 13:49
And for the same amount of, uh, evaluation, it takes only seven hours. It takes also, like, way less edi- it uses also way, um, way less, uh, money, so $265.
- 14:01
And instead of running 500, uh, 500 marathons, I just need to run four marathons. So if three of you want to run a marathon with me, it will be enough to, to, to do this.
- 14:13
So again, the idea is, like, people tend to just look at quality, but it's important not to look only at quality, but also at efficiency, because sometimes the, the additional gain you get with quality is not worth the efficiency of the, the, the compute cost.
- 14:27
So to the question, what model is state-of-the-art? The answer is there are multiple state-of-the-art model. And the tool I prefer for this is usually the Pareto plots, where basically on the X-axis you have one efficiency metric.
- 14:42
For example, on the left it's, uh, ta- latency for the generation of an image. On the right it's the price for the generation of an image.
- 14:51
On the Y-axis, you have some quality score, let's say the Elo score. And here you can draw the Pareto front in red. And you can see that there is not one single state-of-the-art model, but there are actually multiple of them, and there are like three or four.
- 15:05
And you can see that even though the quality score is not-- there is no big variations, it's alway- always between one thousand uh, one hundred and one thousand two hundreds, there is a big difference in terms of efficiency.
- 15:18
So you can be, like, really like times, times, I don't know, twenty times faster just by using the different model.
- 15:26
Even better, if you know the specific task you want to do, you can do like the Pareto for- front not with, um, a quality metric which focuses on general, general capability, but really based on quality metrics which is for the target use case.
- 15:41
So this is some Pareto front focusing on text rendering. And here, for example, we optimized a lot like the Flux2 model, Flux2/Flux models. And we worked with VFL for, for this.
- 15:53
And you can see that you can get way faster, and you can still be on the Pareto front, uh, for the specific use case of text rendering.
- 16:02
So is benchmarking dead? The idea is it's not dead. We can do it properly and get a lot of useful information out of this. And if you use it, uh, in a better way, like, by taking all these, uh, you know, rules, um, when using the evaluation, you will usu- usually not find a large long, uh, a
- 16:20
large foundational model, but more like a lot of small, uh, preference models that will be very good for your use case. So I just listed a couple of takeaways, which are, like, evaluate on many samples.
- 16:32
You look at the user-- use case conditions. Use multiple benchmarks or efficiency. Uh, which are key things to keep in mind when evaluating models. And how, how to reach like state-of-the-art models in general.
- 16:44
Like this is what we are doing at Pruna. We are actually building a lot of what we call preference models with, with-- that are served behind hand point, hand points.
- 16:54
We have the fastest, for example, image, uh, models, video models that can run between one seconds to five seconds.
- 17:02
And but we also try to give a lot to the open source with a lot of open source contributions with a package to show you how to compress your models on your own.
- 17:11
Uh, also a lot of materials on all the best re- research papers for efficiency or even like some efficiency course.
- 17:18
So thanks for your attention. [audience applauding] I think we are out, out of time, but if there are any questions, happy to take them.
- 17:39
Okay, perfect. And, ah, you have a question?
- 17:41
Yeah. Yeah. So about the compression that you guys do, like, uh, could you like e-elaborate like on how you do compression on video models and, uh, image models?
- 17:52
Uh, sure. So actually, there are multiple, like, I mean, you know it as well. There are a lot of family of compression, uh, methods. So of course you, you can guess like quantization things we do a lot.
- 18:03
And we do it a different quant- quantization for every specific module in the, in the model, which is super important. Uh, we can do also some pruning where we just remove some components which are not important.
- 18:14
And, uh, for all these image and video models, something that works quite well is working on the step that the denoi- the denoiser, um, like when you generate a video or an, an image, you usually use like twenty to fifty steps to generate like, um, uh, the, the content.
- 18:31
And you can actually reduce it a lot either via distillation of or caching methods.
- 18:36
Okay.
- 18:36
So you can instead of doing like fifty times the computations using the same backbone, you can do it way less, I don't know, twenty times or even like four times, depending on how aggressive you want to be.
- 18:47
No, I'm asking because I, I'm actually on MLS video we are, we're like doing some caching.
- 18:52
Yeah.
- 18:52
And I wanted to understand if, if you guys know something different that I could like use-
- 18:56
Yeah
- 18:56
... to make it even faster.
- 18:58
So we have a, like, in our package, we have a lot of open source, uh, algorithms for good caching. Uh, but we have also some internal, you know, algorithms that we have for the models we serve behind the, the hand point.
- 19:11
But, uh, yes, there are really advanced caching methods and so on, but yeah.
- 19:14
Thank you.
- 19:15
Sure.
- 19:16
Awesome talk.
- 19:16
Thanks. Okay, then [outro music]