AI Engineer World's Fair 2026
The Desktop Frontier — Ahmad Osman, Osmantic
About this talk
Osmantic founder Ahmad Osman argues that rapid improvements in open-model capability density are moving frontier-class AI from data centers onto desktop workstations and consumer GPUs. He compares successive Llama, Qwen, GLM, and Nemotron models, connects declining hardware requirements to published research on the densing law, and advocates sovereign AI, enterprise ownership of local compute, and the increasing practical utility of existing GPU hardware.
Chapters
- 0:00The desktop frontier and a single-GPU prediction
- 2:23Model efficiency, shrinking GPU footprints, and the densing law
- 5:09Desktop frontier models, DGX Station, and efficient training
- 7:31Sovereign AI, enterprise ownership, and open-model history
- 13:08Dense-model comparisons and the future value of GPUs
Talk transcript
- 0:00
[upbeat music] Hey everyone. We are about to start this presentation.
- 0:17
Uh, it's called the Desktop Frontier, and, um,
- 0:23
it's basically about where we started and how far we've come with local and open source models. Um,
- 0:35
how many like... J- just a quick question. How many of you here follow me on X? [audience cheering]
- 0:42
I'm, I'm amazing. Love you all. Love you all. Uh, so you know, I sometimes every now and then I would say a prediction. Uh, here is a new one.
- 0:52
Within roughly 18 months, we are gonna have the equivalent of GLM 5.2 class intelligence running on a single RTX 5090 with 32 gigabytes of VRAM. Um, that's basically late 2027.
- 1:10
Um, this, this is conservative. We might actually get there faster.
- 1:17
So, um, you know, for a long time, uh, the story has been bigger models, bigger models, bigger models. How can we get to the next five trillion? How can we get to the 20 trillion?
- 1:29
And I'm not saying that there won't ever be like a gap between frontier intelligence, um, and, um, you know, open source models. There will always be a gap. But that gap, um, will shrink and the efficiency of the models will get exponentially better.
- 1:51
So the, the term that I like to think about is impact per parameter. Um, you know, what capability are we talking about? What could the model do? Um, what footprint, like hardware footprint did it have last year in comparison to now?
- 2:10
And, uh, what hardware does that use, and what hardware did it need to use a year ago? And, uh, you know, are we moving down for the same kind of quality on the, on that hardware?
- 2:23
Again, as I was saying earlier, I used to run Llama 2 on an RTX 3090. It's now running Qwen 3.5, 3.6, 27 billion parameter. That's better than Llama 3 405.
- 2:35
That's a 400 billion plu- 400 billion-plus parameters model that you beat with a 27 billion parameter model a year and a half after.
- 2:48
So, um, yeah, as I was saying, similar capabilities are moving into smaller hardware footprint. Um,
- 2:56
benchmark scores are one thing, but also, you know, a year ago, this time a year ago, we didn't have any local models that were able to successfully run within Claude Code, right?
- 3:10
It wasn't until GLM 4.5 that came out in late July, and GLM 4.5 Air required at least four RTX, uh, 3090s or an RTX Pro 6000. Now that footprint for hardware is not needed anymore.
- 3:27
All that you need is a single RTX 3090, 5090, and you have something much more capable, much more intelligent. So is this trend just random or is there more to it?
- 3:41
That's a question that everyone should ask. Um, um, is it just by, by random chance that we've gotten this far from models that were- weren't able to sustain more than four thousand, um, tokens in terms of context lengths, and now we have things that are million param- million, uh, tokens locally on your hardware that you own.
- 4:06
It's not by chance, you know, it's not just a coincidence that we got here. Um, there is research being done. There is efficiency gains to be made. There are architecture hacks that compound, and they will continue to compound.
- 4:21
And, uh, I think I like this line. It's not that small models are beating big models, it's that newer, more efficient models are beating older, less efficient ones.
- 4:33
Uh, so yeah, capability density, uh, is, you know, the literature I back this up with. Uh, uh, Nature Machine Intelligence, uh, calls this pattern dancing law. And, uh, basically, um, you know, every three and a half months, we are having 50% fewer parameters.
- 4:55
Whether that's in dense or activated, that's a different story. But we are getting way more intelligence out of the models that we're running.
- 5:09
So, you know, right now we're at, at, uh, GLM 5.2. That's, uh, our, you know, biggest player. And, um, it's 744 billion parameters total with only 40 billion parameter activated, and that supports up to one million context lengths.
- 5:26
You can run this in NVME four on a machine, uh, on a DGX station or on a server with eight RTX Pro 6000s. That's something that you-- like a DGX station is something that you can sit under your desk, and it's running this kind of frontier intelligence.
- 5:43
Whether, you know, it, it's a on one benchmark it actually beats GBT 5.5 extra high. Doesn't that mean that we're getting somewhere with local and open source models that we can compete with the frontier?
- 5:56
That we're not that far off from the best that you can get from the cloud?
- 6:02
We also have Nemotron 3 Ultra, which proved that NVFP4 training, more efficient training can be done on, uh, on, on hardware, right? That's, that's very important. That means that the footprint even for training these models, for fine-tuning them, for making small and specialized models as I was talking earlier, could be more efficient, could be done cheaper, and
- 6:22
could be, you know, could deliver you value in terms of economics way sooner or, you know, for much less money than you used to.
- 6:33
Yeah. So, you know, again, Llama 2, uh, that was a 70 billion parameter model. If you try to run that right now, you're, you're-- you'd laugh at it, right?
- 6:41
That used to take eight RTX 3090s to load up, and it-- those same eight RTX 3090s could run something like 15 parallel agents right now with Qwen 3.5-27B. That's, that's a massive jump in terms of performance gains.
- 6:58
Um, so the fencing wall basically means that we have similar or better, um, capabilities with significantly fewer parameters. That's the impact per parameter, as I was saying. I want everybody to leave here thinking about this term and, you know, thinking where are we gonna get a year from today?
- 7:19
As I was saying earlier, everyone here has a phone I'm assuming. Raise your hand if you have a phone.
- 7:25
If, if you didn't raise your hand, we know you lie about other things as well. So come on, guys.
- 7:31
So, you know, you can ri- you can now run GPT-4o quality on your iPhone. That, that's massive. That thing required data centers to serve. So why wouldn't you invest, you know, in Sovereign AI?
- 7:48
Why wouldn't you as a consumer, as an individual, as a small-sized business, middle-sized business enterprise, why wouldn't you want to be in control of the models that you're on?
- 7:59
Why wouldn't you want to make sure that nothing gets taken away from you? That every little thing can be optimized for you later on. That the performance gains can be made specially and specifically for your use cases, and that you can save more money that way in the long run.
- 8:19
And, you know, ODS for consumers, it's basically the way that we support individuals, but enterprises also. And I think that there is something that we like as a community we need to think about deeply.
- 8:33
We need enterprises for open-source AI to win. We need these people that are using the cloud right now, that are basically supporting data centers being built for cloud providers to come on this side, to own their own hardware, to own the stack fully end-to-end so that we can keep delivering open-source models.
- 8:54
So that there is an incentive for open-source providers to actually come up with models so that we can come up with new licenses that allows open source to thrive and...
- 9:08
So again, um, open weight and the frontier I think I-- Yeah, sorry, that was a misclick. Um, you know, so smaller models started bouncing above the weight after Llama 2 with Mistral 7B, one of my favorite models.
- 9:23
If you try to build that model right now in, um, Claude Code or OpenCode, it's not gonna work. But it used to take so much in terms of hardware, right?
- 9:34
That you would now get from a 9B milli-- model that I can run with Telegram, with-- or with Hermes, for example, and, uh, do a lot of stuff with.
- 9:44
So, so we've come a long way. Uh, we had that. We had Mistral 8x7B, which, you know, everybody knows is an MoE. Then the progression went, uh, from that to Llama 3.
- 9:56
You know, Llama 38B was one of my favorites. Still is. Um, it had unique identity in my opinion. Uh, then we had like the 70 billion, which was like the, the thing that I would run basically on my eight RTX 3090s at home.
- 10:10
Then there was like the 405, the 400 billion-plus parameter Llama 3 which again required a lot of hardware and if you bought it now against Qwen 3.5, the 27 billion parameter would lose against it.
- 10:24
That's in the span of what? Two years? Two years and some?
- 10:28
No, I think, I think less than two years. That's summer 2024 to March 2026. That's, uh, that's about 21 months. And the next big thing in my opinion, Gemma 2-27V and then we had the Qwen 2.5 and that, that was the moment that I was like, "Okay, we actually are making progress," and the gap was shrinking between
- 10:49
open-source models and, uh, the frontier. Uh, really Llama 3 saved like, you know, it really helped us a lot and then, um, Qwen 2.5 delivered a massive improvement and there was a lot of fine-tuning and experiments that could be done on that one.
- 11:05
There was amazing papers and, um, they helped the community im-immensely in my opinion.
- 11:12
Then the next big thing was DeepSeek-R1 in my opinion, and the reasoning becoming something that you can run at home. That was a massive MoE, almost 700 billion parameters.
- 11:22
Um, you know, you had to have like a very beefy server to actually get it up and running. Um, and then, you know, the improvements that came from just
- 11:33
more training on that one and DeepSeek-R1 that was released in May last year made massive jump again. So it showed that both training could deliver more improvements on the same, on the same checkpoints.
- 11:53
Then GPT open source like GPT-OS 120B. Anyone remembers that one from last summer? Yeah.
- 12:01
Nobody here used it Come on, guys. I, I need some help here. [laughs]
- 12:07
Um, it was, it was, it was one of the first open source models that were able to successfully do tool calling. Um, and, uh, it was a step forward.
- 12:17
It showed us that we can do more with, uh, with the hardware that we have at running at home. Uh, that was a footprint shift, right? From like, you know, that massive 700 billion parameters DeepSeek, uh, R1 that was, yeah, 671 billion parameters to something that was one-fifth, one-sixth of its size.
- 12:36
And GBT-oSS was comparable, maybe better, more agentic performance. Um,
- 12:48
then the moment of, uh, Qwen 3.5, that's 39.7... That's three-- That's 397 billion parameters. Uh, that's, uh, that's a beefy MoE. And, uh, you know, what's funny is that about, uh, it's about 15 times the size of, uh, that Qwen 3.6.
- 13:08
And I'm-- here I'm comparing 3.5 to 3.6 of the dense 27 billion parameter model, and that dense model beats it. And that dense model has 40% higher number of activated parameters, so it's not that far off.
- 13:24
That's massive amount of performance gains in a very small amount of time with massively different footprint in terms of hardware, uh, requirements. And that trend happened in like what?
- 13:36
Two, three months. So, you know, how far could we go from here? Um, how far before we get to, you know, uh, a recent model that there were some, uh, news about, you know, that is, uh, finally relaunched again?
- 13:51
How far before open source delivers something of that quality that you could run on your own hardware and you can control and will not be taken away from you and will not refuse a request from you?
- 14:08
So again, these are just the-- some benchmarks where you can see that an iteration on the 27 billion parameter model, a little bit more post-training improved it across all benchmarks and made it win against a model that is almost 15 size-- for 15 times its size.
- 14:32
And again, remember, this is 27 billion parameters activated versus 17 billion parameter activated. It's still massively the same amount. Like, you know, it's, it's only 40% less in terms of the amount of time it would take to process things, but it's 15 times smaller.
- 14:50
That's a lot. So again, um, how long until the prediction I made earlier becomes plausible when I said that we're gonna have the equivalent of GLM 5.2 running on an RTX, uh, 5090?
- 15:06
This is the math. Seventeen months. And this is a conservative math. Uh, earlier this year in December, I had a, um, a very viral post that I predicted that we're gonna have the quality of, uh, Opus 4.5 running locally at home on a single RTX, uh, Pro 6000.
- 15:24
That happened by March. [laughs] So a question, um, hardware purchases today, does it get more valuable as models become more efficient and smaller in size? [laughs]
- 15:43
That's a good question. So why are you funding other people to build data centers so that you can subscribe to them and pay subsidized tokens and then later on get- [clapping] Th- those subsidies are gonna go away, and you're not gonna be able to run those models, and they will have so many limi-limitations.
- 16:02
So might as well ask yourself, why not own the hardware yourself and be in control?
- 16:07
Um, so yeah, the forward-looking question is basically, what will a DGX station be able to run in three, six, 12, 18 months from now? That's something that there is a reason that I'm not selling any of my RTX 3090s if you follow me, and I have a lot of hardware, guys.
- 16:22
Um, but I'm interested in seeing what I could do with them in a year or two from now, more than in the amount of money I would get for them today.
- 16:32
This is not a financial advice, by the way. Like, let me make that very clear. [laughs]
- 16:36
Um, so yeah. Um, the disk size frontier potential, um, you know. An NVIDIA DGX station could run a lot of, uh... Today it could run GLM 5.2. What will it be able to run tomorrow, six months, 18 months, two years from today?
- 16:53
We know that, you know, RTX 3090s, Amber, uh, architecture from 2020 sells at a higher value than MSRP today and is still being utilized for a lot of use cases.
- 17:06
So what will a DGX station, the actively developed Blackwell architecture, will be able to run in a few months, a couple of years?
- 17:16
That's a good question. So the question you have to ask yourself, um, if an RTX 3050 90 with 32 gigabytes of VRAM runs in the equivalent of a GLM 5.2 in 18 months.
- 17:30
And this is the question that everybody should be asking themselves, and I want you all to be looking at the screen, taking this very seriously. Okay?
- 17:40
Should you buy a GPU? [clapping] Thank you. [outro jingle]