AI Engineer Europe 2026
Frontier AI at Home (literally)
About this talk
EXO Labs co-founder Alex Cheema presents a hands-on workshop about running frontier language models on local hardware. He introduces prefill and decode, argues for local ownership and privacy, discusses energy-efficient hardware-software co-design and Mac-based infrastructure, and explores open-model harnesses, test-time compute, quantization tradeoffs, cluster dashboards, and large-prompt demonstrations through extensive audience discussion.
Chapters
- 0:00Introduction: EXO, local models, prefill, and decode
- 2:23Local AI ownership and dependence on cloud providers
- 16:04Energy efficiency, consumer Macs, and full-stack co-design
- 44:47Open-model harnesses, test-time compute, and local privacy
- 1:24:40Cluster dashboard, quantization tradeoffs, and hardware optimization
- 1:41:24Large-prompt MacBook demonstration and closing
Talk transcript
- 0:00
[upbeat music] I guess let's start simple.
- 0:18
Uh, who, who's familiar with an LLM? [chuckles] Yeah, so everyone. Um, who's, who's familiar with the concept of prefill phase and a decode phase in a LLM inference?
- 0:37
Yeah. Um, who's used-- who's run a model locally? Okay. Wow. That's good. Uh, I, I always say, like, we're very early, but this makes me think maybe-- I mean, literally all of you raised your hand, so maybe we're further along the adoption curve than I think.
- 0:56
Um, and, uh, yeah, this is a, a scary one. Who's used EXO before?
- 1:05
Okay, a couple. Uh, cool. Um, so yeah, I guess, um, just a little bit of background on myself. So I'm the-- Uh, I, so I work on EXO. Um, we'm a, a lab focused on running frontier AI on local hardware.
- 1:26
So what we're doing is we're looking at, uh, full stack, what's involved in running inference locally, and, uh, working across the whole stack, so on the models software, the hardware side as well.
- 1:43
And, um, our mission is to drive down the cost of running frontier AI systems locally. Um, so current state of things is, like, most AI runs in the cloud.
- 1:58
Um, typically, if, you know, I guess this group, uh, a lot of you have used local AI before, but typically, if you're gonna use a model, uh, most people are using, you know, uh, models that are running in data centers.
- 2:13
Um, so why is that a problem and why, why should we even care? Um, well,
- 2:23
uh, I guess the name EXO actually comes from exocortex, so this is the idea that, you know, AI will go beyond just being this tool that you use for a chat interface, which it already is, and it's more of a kind of extension of yourself.
- 2:40
And if you think of it like that, it's almost like a part of you and a part of your brain. And then you ask the question, "Well,
- 2:49
do I actually want to rent my brain?" And I think Andrej Karpathy, uh, has like a nice one-liner that really summarizes this well, which is like, "Not your weights, not your brain."
- 3:01
Um, and, you know, there's this something, uh, quite deep about that statement, I think. And I think
- 3:12
we're now starting to realize, you know, as things like OpenCore, you know, and agentic systems are getting more popular and it's more than just a chat interface, that the,
- 3:27
um, there's a lot of concern about, you know, where is this thing running? Where is my data going? Um, you know, what if it gets cut off, right? Like, you know, I had, I had a friend that works in cybersecurity, and, uh, he was doing some, you know, penetration testing, which is, uh, uh,
- 3:51
pretty innocent, uh, you know, uh, doing it for securing a system, and basically got locked out of, like, three of the, um, API providers, uh, like Claude, Gemini and, uh, and ChatGPT.
- 4:08
Um, you know, there's-- this is becoming more than just, you know, this, this, uh, this chat, right? It's like, you know, you need it to basically be competitive in any field now.
- 4:24
And there's, uh, you know, with centralized systems you have, uh, you're relying on a few organizations basically for this. And, um, I think, you know, there's kind of this, there's two realities, right?
- 4:39
One is where we have this closed source world, and in that world, there's gonna be this massive power law where there's a few companies that basically have the most capable models.
- 4:50
And what they'll do is, I mean, th- they will rent-seek on, on that because, you know, that's how they make the most money. And, uh, you know, what we believe is there needs to be a kind of a competing force of, you know, keeping things open and actually being able to run these models without needing massive data
- 5:09
centers and massive amount of compute. Um, so
- 5:17
I just wanna talk a little bit about sort of, you know, the actual technical side of what we're doing. So, um, I guess the first, uh, the first thing to realize here is, like, a lot of the discussion is kind of around training or has been around training and, um, you know, that's, um, been
- 5:38
a lot of the focus of, for example, like open source efforts has been in like, okay, we need to have like, um, you know,
- 5:46
transparency around training. And, uh, what I'm talking about here is like specifically inference. So like if you already have the model, you know, how do you run that model?
- 5:54
And, uh, if the only option is to buy a million dollars of hardware to run like frontier model, then that's like a massive barrier and, um, that's kind of what we're focused on here.
- 6:05
And, um- Like, in terms of the hardware, because the focus has mainly been on training, um, what, uh, you tend to see is that
- 6:17
there's this idea of the hardware lottery, which is, you know, the research that's out there isn't actually, um,
- 6:25
you know, the best possible thing that you can do. Uh, there's a lot of research that's being done on the current hardware stack, which has primarily been built around training, and, uh, what that looks like is basically these NVIDIA GPUs, uh, that you stack up in a data center.
- 6:42
Um, but that's not necessarily the, the, you know, the, the best thing to do, especially then when you look at inference. So, um, you know, the, the, the current, uh, the current hardware is very much focused on, like, you know, FLOPS and FLOPS unit economics.
- 6:58
And, uh, you know, there's this... I really recommend, like, giving this a read. Uh, it's from Sarah Hooker, who was, um, she was at Google Brain and, and then Cohere.
- 7:09
Now I think she's doing her own thing. But, um, you know, basically the idea is, like, there's all these ideas out there that haven't really been explored because we just have a lot of inertia behind the current way of doing things with the current hardware.
- 7:21
So our thesis is, you know, there's all these things that haven't really been explored, uh, with different hardware and in the context of inference that, you know, can really, um, be a lot better than the current way of doing things.
- 7:38
And, uh, this is really an area that's super,
- 7:42
uh... There's a lot of, like, low-hanging fruit out there. So, like, you know, with a little bit of, um... Like, I'll give you an example. So, like, the other week, we were just looking at Qwen 3.5 and running that locally on Apple Silicon, and we found that, like, if you look at the theoretical,
- 8:03
uh, speed that you would get running this, it's, like, way off what you get in practice, and, uh, it was off by, like, 50%, so, like, 50% slower than what we thought it should be.
- 8:14
So we looked into, like, what it's doing, and we found, like, basically there's, like, a bunch of overheads that is introduced by inefficient kernels, and specifically, like, having a lot of unnecessary kernels that kind of get launched separately, which leads to a lot of, uh, overheads when you're running inference.
- 8:34
And, you know, each one of these, um, kernels adds, you know, quite a significant amount of overhead if you're thinking about running a model like, uh, Qwen 3.5 locally.
- 8:46
You know, theoretically, you might be able to get, like, 150 tokens per second, um, which is, you know, a token every less than 10 milliseconds. So if you've got, like, a few milliseconds delay here and there, it's like, it adds up to a lot.
- 8:59
And so what we did is we just did a little bit of work on sort of looking at, okay, what's going on? And we realized, you know, there's all these separate, um, kernels being launched that are unnecessary.
- 9:11
So we just did some pretty basic work to fuse that all together and increase the inference performance by 30%. So, um, just an example of, like, you know, there's a lot of stuff out there that maybe you would think is optimized and you would think is already, um, pretty close to, you know, getting the best utilization out
- 9:34
of the hardware, but it's just not the case. And, um, you know, this, this exists across the whole stack. So that's on the kernel level, but there's also a lot of stuff, you know, in the orchestration, um, in terms of, you know, how you connect different pieces of hardware, um, you know, communication overhead, uh, even stuff like
- 9:56
the harness, right? So, like, um, I think what people are starting to realize is actually there's a lot of value in the harness, and, um, you know, for example, if you use Claude code with a certain model, um, and then use OpenCode with exactly the same model, you get completely different performance.
- 10:13
And, um, you know, this is specifically interesting in the context of local because you're super resource-constrained. So
- 10:23
anything you can do on the harness layer that kind of is aware of the hardware, you'll be able to get a lot of gains there. So it's just across the whole stack.
- 10:33
Um, so broadly speaking, like I said, training is about FLOPS, right? So that's basically everything is compute-bound, and what matters then at scale is basically how cheaply can you get FLOPS and, you know, how much energy is it gonna use?
- 10:49
Um, but inference is mostly about memory. So, um,
- 10:56
you know, basically everything, all the operations you're doing, depending on the model, is like most of those operations are memory-bound. And specifically, like, when you're running stuff locally, if you're running so-something for yourself, you don't have the kind of, um, you, you don't have the ability to batch together, you know, multiple users' requests.
- 11:15
So everything you do is kind of, uh, low batch size, which is memory-bound. Why is that? Well, yeah, basically, uh, this is kind of how, uh, inference looks broadly speaking with most models today.
- 11:29
So there's kind of like a prefill stage and a decode stage, and my argument is prefill doesn't actually matter that much, um, especially when you're running stuff locally. And, uh, the reason is basically...
- 11:43
Okay, so just explain, like, prefill stage is, uh, loading your context with your prompt. So, you know, if, if you're loading in a PDF or something, it's, um, the part that actually, uh, generates all your KV caches, and then you have the decode stage, which is autoregressive, and that, uh, just generates token one by one.
- 12:04
That part's memory-bound. Prefill part is compute-bound. With the, with the prefill part, um, what you're seeing actually is, like- And again, this goes back to this idea of like, you know, the harness matters a lot.
- 12:16
Um, so a good harness, what it will do is it will get a lot of cache hits. So it will keep the prompt mainly the same. And if you look at-- if you type in a /context when you run Claude Code, you can actually just see this.
- 12:27
So you can see that, um, you know, it'll basically show you like, uh,
- 12:33
all the parts of the prompt. And you can see there's like a big part of it that is just the system prompt and system tools, and that stuff doesn't really change.
- 12:40
So maybe if Claude, you know, maybe if they push an update, this will change. Uh, but broadly speaking, you can, you know, keep most of the prompt the same, uh, when you're running, you know, actual workloads end to end.
- 12:53
You know, maybe in the benchmarks you'll see people that are doing stuff at like, you know, really long context, uh, really long prompt sizes. And, you know, my, my argument here is it doesn't matter actually as much as people think.
- 13:06
Um, so you know, really, um, it's about decode. And, uh, what matters for decode? Well, it's three things. So like I said, it's memory bound. Uh, but you know, the first thing is you have to actually fit it into memory.
- 13:21
So, um, if you wanna run a model at a good speed, then if it doesn't fit into memory, you're gonna be loading it from disk, which is super slow.
- 13:32
Uh, so that's like the first like hard requirement, basically. Uh, second is memory bandwidth. So this tells you how fast it's gonna run. So how fast can you actually load, you know, your model weights, load your KV- KV caches into a GPU?
- 13:45
Um, and the third thing, and this is something that, um, is particularly like important locally, is, uh, the energy. So like, you know, um, here I, I'm talking about energy per bytes.
- 13:59
I'm talking about like in these memory bound, in this memory bound decode phase, like how much energy does it cost to move one byte, one gigabyte? Um, and that tells you basically, if I'm gonna run an inference, like how much power is it gonna consume?
- 14:14
Um, yeah, this matters a lot. I mean, you-- I don't know, like, um, there's a lot of these kind of demos of doing stuff on phones, uh, that you see on Twitter.
- 14:24
And, um, we did a lot of stuff with phones, you know, a while ago, maybe 18 months ago. Um, but we found like
- 14:34
there's a big issue, which is the energy. Uh, and you know, the batteries on, on phones are quite limited. So, you know, maybe you're talking about ten to 15 watt hours, uh, on like an iPhone.
- 14:47
And, uh, inference, you know, it might be consuming something like ten to 15 watts, right? So then you're talking about like one hour of battery life, which is, uh, not really usable.
- 14:57
So of course, um, you know, there's a lot of work being done to improve that, and I think phones eventually will be able to run, you know, better models on phones.
- 15:07
But right now it's just like, you know, it consumes a lot of power, and it also gets really hot. So actually, when we were running benchmarks, um, like 18 months ago or so, like the phone was getting so hot, um, that I couldn't hold it.
- 15:21
It was actually like too hot to, uh, to hold. Um,
- 15:27
so yeah, these are the, the, the three-- this is what these like actually map to in practice. Um,
- 15:33
and, uh, on the-- So like there's this, uh, paper that's pretty recent, um, and, uh, it's from a group in Stanford, uh, Hazy Research maybe. Uh, some of you know them from their other work, like, uh, Thunderkittens and stuff.
- 15:49
Um, but like now they actually have a group that's focused on this concept of intelligence per watt. And the idea is this kind of encompasses these three things, and it's, uh, a way of like tracking, like how are these things improving over time.
- 16:04
And the metric is, um, you basically look at like how good is a model at a specific task. You divide that by the energy that it uses. And if you track this metric, uh, you see that, you know, this is actually improving, um, exponentially.
- 16:20
So over the past two years, uh, it's been about 5X. Um, the correct term for this actually shouldn't be intelligence per watt, it should be intelligence per joule really, because you're not really, um,
- 16:33
concerned so much about the time here. You're concerned about, okay, for a given task, you know, how much energy is that gonna consume? Um,
- 16:43
yeah. So yeah, this is, uh, this is now looking at, uh, sort of memory improvement. So this is actually like, I mean, the fact that you can buy a commodity piece of hardware like this that has-- I mean, Apple got rid of the 512 gigabytes option recently, so you can't buy 5...
- 17:05
These are 512s. Uh, but uh, you know, you can buy a 256 gigabyte and it's like, you know, you can buy off the shelf basically. Um,
- 17:15
so you know, th-this, this is like kind of a very new thing. Like you wouldn't-- you weren't able to buy consumer hardware that had this much memory in it.
- 17:22
Um, so this is like improving dramatically and, you know, uh, yeah. Give you an example as well, like the new MacBook M5, um, Max has 120-- like you can get it up to 128 gigabytes, and the memory bandwidth is pretty, pretty damn good, um, 614.
- 17:41
So yeah, this is kind of, uh, looking now at the intelligence per joule. Um,
- 17:52
so, uh, da, da, da. Yeah. So it's actually-- So the 5X from the hardware improvements, then you've got another 3X from, uh, model improvements, and obviously these things compound.
- 18:05
So what you're seeing is that, um, a lot of the, uh-
- 18:12
Improvements here are coming from, you know, hardware, right? Um, and also, uh, the model layer. So this again goes back to the idea that, you know, it's, it's kind of about looking at the whole stack, right?
- 18:24
There's a lot of gains across the whole stack and, uh, there's different things you can do. Uh,
- 18:33
can I just ask where we are on the demo? Oh, I'm, I'm, uh, like a screen to be honest. Screen. Yeah, I need a screen obviously. Okay, maybe we can ask.
- 18:43
Yeah, yeah. Yeah, 'cause I wanna show you this as well, right? Uh, I don't, I don't just wanna talk. Also, if anyone has any questions or anything... Yeah, you have a question?
- 18:50
Yeah. So just in terms of the current gen Mac side of things, like where do you see this going in real life? Like I, I'm trying to tell my friends like think of this like an appliance in your house, or if you're gonna go spend money on a fridge, buy one of these.
- 19:04
And they look at me, they're like- [laughs] ... "You're telling me to spend 5 grand on a box," right?
- 19:09
Yeah.
- 19:09
Um, where is this headed in terms of like the consumer appetite? And do you know, is it purely just upstream supply chain issues, or is Apple seeing this as a new profit center and maybe repositioning pricing strategically?
- 19:19
Like-
- 19:20
Is Apple seeing it as what, sorry?
- 19:21
Like a new, a new blue water segment that they can go sell into. 'Cause it's-
- 19:25
Yeah.
- 19:25
I'm just-- What I've heard about supply chain is that when these are built, they're already funded. So like both-- they'll essentially get all their commercial contracts in for RAM, and then they'll price the, the computer.
- 19:36
And I think what's interesting is the pricing change to bump the storage and have kind of reasonable RAM.
- 19:42
Mm.
- 19:42
Um, I more, more just wanna hear your theories. Like I see you as a, a buyer of, you know, a lot of this hardware.
- 19:47
Yeah.
- 19:47
Curious what your thoughts are on like the future of pricing and just like consumption as a consumer.
- 19:53
Yeah. Um, maybe I can just repeat the, the question for the-
- 19:59
Yeah
- 19:59
... for the mic. Uh, the-- but the question is, uh, where do you see, where do you see like, you know, the hardware going? And, you know, are people actually gonna be spending $5,000 on some kinda inference machine?
- 20:13
Um, so my thesis is there's, you know... Like I said, um,
- 20:23
you can look at certain metrics like intelligence per watt, right? Um,
- 20:29
and what you're seeing is that progress is exponential also for local, right? And, um, you know, right now today, you need something like this, right? If you're gonna run frontier models, then basically your only option is...
- 20:45
Okay, first of all, the bar is always moving because, you know, maybe now the frontier open model, you know, we have GLM 5.1, which I wanna show you as well, uh, that came out yesterday, and that's probably the frontier model now for, for open source.
- 21:01
And if you wanna run that, that's a trillion, it's a trillion parameters. Um, and
- 21:09
natively, it's, uh, it's FP16 as well. So like, um, you're talking about like 1.5 terabytes or something, right? So you need to like fit all of that into memory, right?
- 21:23
Going back to the first thing. Um, and for that, you're talking about, you know... Let's assume that you still have the 512 gigabyte Mac Studios like these, then talking about like $40,000 of hardware, right?
- 21:37
Um, and even then, it's not gonna run that fast, right? So like it's kind of, you know, maybe gonna run at something like 20 tokens per second if you're doing this kind of setup.
- 21:47
Um, that's not acceptable for a lot of people. Like, you know, I think people are used to a bit more than that now with, uh, the cloud, so maybe something like 50 tokens per second is kind of what people are accustomed to.
- 21:59
Um, so th-that's like the, the current state of things, right? But, um, you know, our, our thesis is this stuff is gonna, um... There's 100X in there. So like, you know, if you look at like all these parts of the stack, because they compound, um, if you're making like changes to the harness, if you're making changes to
- 22:19
the, the models, if you're making changes to like the kernels, all of this stuff together, um, there's still, you know, like, um,
- 22:29
100X in terms of like price to performance in there. So where I think things are going is like quite soon you will, like I think within, uh,
- 22:41
within, let's say 18 months, you'll be able to spend $5,000 or $5,000 and have close to frontier level performance running quite fast. Um,
- 22:54
but it's gonna take kind of this co-design, right, of like looking at things across the whole stack. You already see that happening in the data center. Uh, so you know, um,
- 23:05
the Groq acquisition by Nvidia, you know, you're seeing like more specialization, different chips for different things, and this idea of e-extreme co-design. So looking at everything together and seeing like, okay, how can we build this, these things in a way that, uh,
- 23:23
you know, gets the best performance for like the end use case, right? And I think things have consolidated and s-- like even though things are moving quite fast, it's like some things that don't seem like they will probably change.
- 23:35
Like, you know, having like, uh, agents seems to be, you know, the, the paradigm and it will be the paradigm for a while. Um, I don't really see that changing.
- 23:46
Uh, you know, you see like massive MOEs, right? So like that also seems like that's kind of consolidating, and that's why Nvidia can actually now go and say, "Okay, we're gonna build like specific hardware for this architecture," because, you know, they know kind of more about how these end use cases are gonna work.
- 24:07
And, uh, yeah. So I would say like maybe now I wouldn't actually recommend to like friends to say like, "Oh, go spend $5,000" unless they wanna experiment Um, but I would say within, you know, definitely within two years, you a- you will actually have some products on the market that you can just say, "Hey, go buy this
- 24:27
box." And instead of s- having all these subscriptions or, you know, now if you wanna use OpenCloude, you can't even use that with, you know, the, the Opus API, right?
- 24:37
Um, so you're basically gonna be spending probably-- I know people who are spending like $1,000 a day on, uh, tokens at the moment, right? So instead of that, um, you know, just say like, "Buy this hardware and, uh, you know, you never have to pay for a token again."
- 24:55
Uh, it's basically free, right? It's just the electricity cost. Uh, so yeah. I, I would say in two years. Yeah.
- 25:05
Actually, that was-
- 25:05
Yeah
- 25:05
... my question, is to say, like when should you buy hardware? So you just answered it, but I guess just as a sort of riff on-
- 25:10
Yeah, it's an estimate. I mean, [chuckles] but yeah.
- 25:13
Yeah, to riff on that idea. There's like so many things happening at one time. One is like the, like the, the, the... you know, these models are getting sort of bigger and better.
- 25:21
Also, like there's-- they're getting smarter and smaller at the same time. So it's like, there's like, you know, some people are trying to make small- smaller and smaller models smarter.
- 25:32
Some people are trying to make bigger models smarter, but like even-
- 25:35
Mm
- 25:35
... even smarter. And there's also like price of energy is kind of like, well, it's going up, but people are trying to decrease it. There's also the price of like hardware is kind of, you know, arguably coming down because it builds a lot more.
- 25:48
So where do you see like all these macro-- I'm trying to figure out like where-- which of these macro shifts is gonna win, as it were. Like which is gonna be the driving force?
- 25:56
I'd love to sort of hear your perspective on whether we're still gonna be like, uh, having models that still cost like thousands and thousands of dollars to run each month or whatever it is.
- 26:07
Like or whether or not it's all gonna, you know, shrink down the cost of hardwa- hardware is gonna sort of, you know, cheap, great hardware will make all these models much cheaper.
- 26:18
Or are we always gonna be training larger and larger models or always gonna push the limits of hardware-
- 26:23
Yeah
- 26:24
... that exists? How, how do you see those macro trends kind of like going out from here on?
- 26:28
Yeah. So the, the question is like more about the macro of like, you know, where the models are headed and, and, uh, you know, you, you have like a lot of progress in making these models bigger.
- 26:39
So like the rumor with Myth- Mythos is like 10 trillion parameters. I, I think that's the rumor, like 10 trillion parameters. So like, um, and then you also had rumors about Gemini being something like 4 trillion parameters.
- 26:52
So like, and then maybe the next model run is like 20 trillion or something. So like, you know, you're seeing that the models are getting bigger. At the same time, Gemma 4 just came out, which is like tiny, and it seems to be better than, let's say, the, the best model from two years ago, right?
- 27:09
So like, you know, where are things headed? Um, I think it's all of, all of the above. Like, I think there's, there's gonna be like, um-- Well, there already is, and there's going to increasingly be a massive demand for compute, and there's always gonna be like a lot of progress being made in the data center to just,
- 27:31
you know, pack more compute and, um, I think what you'll see
- 27:37
at some point is maybe things will bifurcate. So like there'll be like all these things that you can run locally and, uh, you can maybe do like 99% of things locally.
- 27:48
Because if you look at like different use cases, all these use cases kind of follow, um, an S curve, where like if you look at how intelligence is, is, is changing, then at a certain point, there's massive diminishing returns on h- running a more intelligent model for a certain use case.
- 28:05
For example, I don't know if any of... Has anyone used Whisper Flow? Yeah. So like as f- as far as I understand, like with Whisper Flow, there's like a little bit of intelligence there to do like good transcription and, um, to me, like
- 28:19
having like a 10 trillion parameter model go-- First of all, like having a 10 trillion parameter go reason for a few minutes wouldn't work in that use case 'cause you need it to be low latency.
- 28:28
But secondly, I don't think there's actually much of a return there on the utility that you're getting as a user. Um, there seems to be like a threshold at which, okay, if you have enough intelligence, the transcription is good enough.
- 28:41
And I think every use case will kind of follow this, right? And you can point to like loads of other things that are already kind of like this. Uh, so summarization, for example, um, or, you know, something like, uh, creating a to-do list or, you know, summarizing emails.
- 28:59
Like simple things like that. It's like, do I really need this massive model? Like no, I think there's like massive diminishing returns. So
- 29:07
that's gonna happen, and then you'll have a point where, okay, there's still use cases where you just do need a lot of compute. For example, if we're gonna like cure diseases or whatever, then obviously, you know, you're gonna want like a really intelligent model and spend a lot of compute on that.
- 29:21
But, um, most things that like, especially consumers use, they won't need it. Um, and uh, that's where I think most things will be able to run locally. So
- 29:33
that's where then you have these two cases, right? Like either you need to spend like a load of compute, and you need like billions of dollars of compute to do, you know, this insanely complex thing or, you know, you can just run things locally.
- 29:45
Uh, that's my thesis. Um, I think it's hard to like the, the-- I guess, uh, a-- One other thing that, uh, Karpathy said recently is like the fog of war is closer and closer, uh, getting closer and closer.
- 30:01
So like it's more-- There's more and more uncertainty. Like for me, I've been quite surprised a few times over the last two years. For example, I think the biggest like inflection point for me was, uh, you know, the adoption of code code and Opus 4.5 and, you know, a lot of people, uh-
- 30:21
Came, uh, you know, uh, back to work after, like, the Christmas and used these latest models, and they're like, "Holy shit," like, things have really improved. And that was the case for me at least.
- 30:36
It was around Christmas time when I was, you know-- That was the first time actually I used, uh, Claude Code. Uh, and then, you know, I was still kind of skeptical at that point that it would be able to do a lot of the work that we were doing, but then it seemed like, oh, the frontier has
- 30:53
moved quite a lot. There's a lot of stuff that just wasn't possible before with, with the older models, and now it's suddenly possible. And a lot of that was also the work on the harness and, and stuff like that.
- 31:03
Um, so no, [laughs] it is just a lot of uncertainty here as well. Like, uh, I mean, who, who knows, like, the next generation of the models, what they're gonna be like.
- 31:15
Um, but I think generally speaking, though, this will be true that, you know, there's not kind of just unlimited returns on intelligence, right? There's a certain point where you just don't need any more intelligence to do something.
- 31:29
And, uh, if that continues to be the case, then it doesn't really matter where things go. Um, you know, as long as the, the kind of progress continues, then you'll just be able to do more and more stuff locally.
- 31:44
Uh, yeah. [laughs]
- 31:47
Um, going back to the consumer, uh, five thousand dollar inference box, is that what you're saying is that we'll get to this point where local models, uh, with slightly newer hardware will be good enough, uh, smart enough for most consumer use cases, right?
- 32:08
So we're kinda at this point where the standard use cases now that we might be using context models for work fine on a Mac Mini style device, not a Mac Studio.
- 32:18
Uh, but we'll still look for more compute, uh, in the-- because it gives us, well, a better inference because it gives us competitive advantage or we're solving frontier problems, and that's what pushes the curve at that point where
- 32:34
consumer advantage might, like, flatten for, for inference. Um, I'm interested in-
- 32:41
Yeah
- 32:42
... like another factor at play, which is that, um, everything's very new at the moment, so we haven't seen a lot of, like, model on a chip come through, I think at a level that really meets people's
- 32:55
needs for inference. But I think as we, like, uh, as we have open-weight models that stabilize and are smart enough for interesting use cases, like, I think they'll emerge in time of potentially being a sufficient cheap alternative, uh,
- 33:13
in a, in a world where maybe the cycle of obsolescence, like, slows down a little bit, and we can then get better, better use out of that. Is that--
- 33:22
Is, is that a, an interesting hypothesis or is it misguided in some way?
- 33:28
So what is the hypothesis? So the hypothesis is that-
- 33:32
That, that perhaps like ev-everything's changing so quickly now that we don't have time to bake a good model on a chip, you know?
- 33:39
Oh.
- 33:39
But everybody's moved on to, say, like, the next hotness already.
- 33:43
You-- Are you talking about, like, specialized chips for specific models?
- 33:47
Yeah. So like, say, like, when free or to-
- 33:50
Yeah.
- 33:51
Yeah.
- 33:51
Oh, yeah. Um, okay. So the question is about, uh,
- 33:58
things are changing really fast right now, but maybe if things kind of slow down a little bit or stabilize and consolidate, then the best, uh, thing is actually to have specialized chips.
- 34:09
Well, like when it becomes-
- 34:09
Potentially. Yeah. That's the thesis
- 34:10
... an option or like a, a more viable option in the mix that we-- like hasn't-
- 34:14
Yeah
- 34:14
... really emerged yet for a lot of use cases.
- 34:16
Uh, yeah, it's a re-really good question. Um.
- 34:24
Like Talos.
- 34:26
Yeah, like Talos, for example. I mean, Talos is on the extreme end of like, you know, literally hardware built specifically for a certain model. Then you have-- I think you have a spectrum, right?
- 34:36
Like, you have GP GPUs, like general purpose GPUs like NVIDIA, uh, RTX or something, or H one hundred, right? Which is like-- It was meant to do, you know, be this platform that you can basically do everything.
- 34:47
And then on the other end you have Talos, and then in between you've maybe got like Cerebras, Groq. Um, you know.
- 34:55
I think for now at least it didn't really make sense to, um, build these specialized chips for LLMs because they are changing so quick, so the frontier is constantly moving.
- 35:03
But it is really interesting, I think. If you then start to think, okay, well, if you are gonna hit diminishing returns on a lot of these use cases and you have a model that's kind of good enough, then maybe it-- At that point, you, you kinda have to look at, like, what is the cost of building that
- 35:22
versus the savings that you'll get, um, using that hardware and deploying it at scale. Um, like the math didn't really make sense at the moment 'cause maybe the frontier moves every three months.
- 35:33
So like, okay, by the time you built the chip, it's useless, right? Um,
- 35:40
but, uh, I think that will change. Yeah. So I think, you know, we're not, we're not really like a-- We're not building our own hardware. Like, we are doing a lot of work to, like, figure out what the best hardware is to use.
- 35:52
But, um, at some point, you know, maybe it would make sense to start doing that or like work with someone who is building hardware because,
- 36:00
yeah, like especially if you're kind of co-designing the whole stack, you can be very opinionated about the models. So it might even be the case that, okay, today there's-- If you look at the
- 36:10
closed labs, right? They have to provide an API that's quite generic because there's all of these use cases people are using it for. People are relying on, on it for, like, all these different things, right?
- 36:21
And, uh, as a result, they have to have this big monolithic model that can kinda do everything. But you could imagine, if you're gonna go be more opinionated and maybe things do kind- ...
- 36:31
consolidate and there's certain use cases we know, okay, this is what we want, then you could be more, um, kinda strategic about specializing the models. So you could say, for example, "Okay, we're gonna have 20 different models," and these 20 models cover, you know, pretty much the same things that this big one, one big model could do,
- 36:52
but like each one is specialized for a different task. And, uh, maybe then what that looks like if you then look at the hardware is you would have a few chips that are specialized f- you know, can run those models really efficiently.
- 37:05
Um, yeah, I think that is, uh, yeah, that, that, that, that should, um,
- 37:13
that should be where things go, right? If you assume that, okay, there is gonna be this, um, uh,
- 37:20
things are all gonna flatten, right? And there's, you know, for a consumer, you can do 99% of things locally, then
- 37:29
I think that thesis will probably play out. Um,
- 37:36
yeah, I, I, I, uh, I think that's a really good point. Yeah.
- 37:43
Yeah. Cool. Yeah. I mean, I like this. I, I think it's, this is better than just talking.
- 37:51
Yeah, I think, uh-
- 37:52
Yeah
- 37:52
... got a question on the hardware side of things.
- 37:54
Yeah.
- 37:55
Because, um, you compared the NVIDIA GPUs versus the Apple's metal hardwares. Um-
- 38:03
Compare the-
- 38:03
Compared to the Apple's-
- 38:05
Yeah
- 38:05
... metal hardwares, uh, based on this architecture. So if you load the big model into, uh, obviously they have big memory, but if you load the model, they can fit into it, but it, when you actually run the inference, the hardware get, like, degraded so, so quickly.
- 38:20
Like, you can't get that much. Even if you have, like, 120 gig, GB Mac, compared with the 5090 RTX, you can get way better inference on that NVIDIA hardware compared to the Apple hardware.
- 38:33
Did you feel something like that?
- 38:36
Yeah. The question is, uh, you, you can fit stuff into memory on a Mac, but it's slow. So, like, when you run things on RTX, it's much faster. So did you see
- 38:48
this and, uh, what do you think about that? So,
- 38:57
yes, uh, I agree with you completely. Like, um, these,
- 39:06
these things are very different, right, in how they're designed. So the Mac is kind of unified memory. It's a big pool of memory, and it's maybe not as fast in terms of memory bandwidth.
- 39:15
It doesn't really have that much compute. Then you have the RTX, let's say RTX 5090. You're talking about, if you compare that to a MacBook, it has way less memory, right, 32 gigabytes of VRAM, but it's
- 39:29
much faster, right? So I think it's, like, GDDR7, which is-
- 39:32
Yeah
- 39:33
... close to two terabytes per second, uh, maybe, like, 1.5 terabytes per second, whereas a Mac, you know, you're talking about maybe, let's say th-this Mac Studio, right? It's 512 gigabytes, so it's, like, more than 10 times the memory, but it only has 800 gigabytes per second memory bandwidth, so about half.
- 39:52
And it has about 10 times less compute, right? So,
- 39:57
you know, I think, um, that's all true, and, uh,
- 40:04
my... Our thesis is basically you want both. Um, so,
- 40:13
um, I talked about these two separate phases of inference, right, the prefill and the decode. But, uh, you know, like, we tend to, like, look at these models as just a bunch of layers, right?
- 40:26
That, you know, we don't really go much deeper often, especially, like, if you're thinking about, you know, running a model. A lot of people are just like, "Oh, well, you know, I just need a certain amount of memory bandwidth to run it at this speed," or whatever.
- 40:38
But, like, um, there's a lot more going on under the hood, right? Like, you know, if you start to look at the architecture of these models, then there's a lot of things happening there and
- 40:52
my thesis is you wanna actually, I mean, this is already happening in the data center, but you wanna kind of run different parts of the model on different devices.
- 41:00
So the-
- 41:03
The reason I ask because I got this one, 120 gig, uh, Mac. I got GPT 120 billion model right here. If I run this inference for 15 minutes, this Mac is useless.
- 41:15
It gets so hot. Battery, even if I plug the in, battery get, like, drained within five minutes, 10 minutes. It's useless. It's like, basically, compared to if I run the same inference on, like, RTX 5090, I can get much more, uh, I can, I cannot-
- 41:30
Yeah
- 41:31
... do a big model, obviously. I can fit like 30 billion models, but inference I get is amazing. Um, so that's why I asked the question-
- 41:38
Yeah
- 41:39
... about your experience with this. So I don't-
- 41:41
Yeah. So somewhere here, um, can, can we show the Spark?
- 41:45
Yeah. There you go.
- 41:47
So, like, uh, actually, NVIDIA doesn't generally sell their own hardware, right? Like, they actually usually partner with OEMs that then sell the hardware to the end customer. But in the case of the Spark, the, the one you've seen is probably the NVIDIA one 'cause they made an exception there.
- 42:02
They have, like, founder's edition, which is, um...
- 42:05
It's not this one. So this is ASUS one, but it's the same thing, basically. Um, and, uh, this costs about $4,000. And, uh,
- 42:21
one thing we did, uh, recently is, um,
- 42:26
combine this with a Mac, right? So basically, this is just a very simple thing you can do, right? So you can basically run- This has a lot more compute, so you can run the prefill phase of inference on here, and then this has more memory bandwidth, so you can run decode there.
- 42:42
Now, uh, with the RTX, it's like different because it actually has more memory bandwidth, right? And more compute, but it's a much smaller pool of memory. So-- But that's actually, um, kind of fine.
- 42:56
Um, so without going into like too much technical detail, like you can basically get more granular and start splitting up the model in different ways and, you know, that's actually the most cost-effective thing to do.
- 43:09
It's kinda crazy that you have to do this at the moment, like that there isn't actually some hardware that just has it all. Uh, but that, that is the case today, right?
- 43:17
So we-- today, like the optimal thing to do if you wanna run models locally is actually to do both.
- 43:23
Yeah.
- 43:23
So you would have a Mac, Mac Studio or MacBook.
- 43:26
Prefilling part and then prefilling. So prefilling is Mac and then decoding is the-
- 43:31
Uh, the other way around. So-- But, but like, you know, there's more that you can do as well in terms of splitting up the model and, you know, w-we're gonna like release some stuff soon, which makes it really easy to do this.
- 43:45
So imagine, you know, you can just plug in an RTX directly into your Mac-
- 43:51
Yeah
- 43:52
... and get like a 3x speed up, right? On running large models. That's the kind of thing that we're working on and,
- 44:00
um, again, it's kind of analogous to what's happening in the data center, right? You've got like, with NVIDIA, you've got these Groq chips, which are running part of the inference, and they're like very, very high memory bandwidth, and then you have like tons of them, so you have this massive pool of memory, and you pair that with
- 44:16
NVIDIA GPUs.
- 44:17
Yeah.
- 44:18
And, uh, same thing's happening with-- Cerebras is doing something like this as well now with Trainium, AWS Trainium chips. Um, there's, uh,
- 44:29
um... Yeah, there's a lot of like interesting work being done in data center side, but like we actually-- we think the same thing will happen locally. Just like the software isn't quite there and the hardware's a bit awkward.
- 44:40
Like, the fact that you have to stack these Macs like this and connect them all with Thunderbolt cables and stuff, it's like-
- 44:46
Yeah. And that's-
- 44:46
... a bit awkward
- 44:47
... definitely, I think, this is a blocker because instead of harness engineering side of it, which basically you probably don't need a cloud code or Codex because you can get any local, um, open source model GLM and all these things, and just manage to fit it and build your own harness on top of it, and then whatever
- 45:02
tools are needed, all this thing, and then run all these things, um, or coding agents and all this stuff. But I think the biggest blocker is like how can you just prefill and these things because you need at least, at least 120 plus, uh, billion parameters to train three or something.
- 45:20
At the moment, yeah. But like if you see, for example, um, with your RTX, actually, you could try running Gemma 4, the dense version, right? Should f-should fit in...
- 45:29
You might have to quantize it a bit, but should fit into memory. But like the-- That's also interesting, right? Because, okay,
- 45:37
m- you know, MoE is maybe the current paradigm, but maybe it makes sense in some cases to run dense models locally as well, right? So, and in, in the case of dense model, if you run that on a Mac, that's gonna be slow because, um, now, okay, this, the ratio of memory to memory bandwidth on the Mac
- 45:53
is very high. So like, whereas on the GPU, it's, you know,
- 46:00
a lot lower. So, um, basically, uh, it doesn't actually benefit you too much to fit this whole model into memory n-anymore that's like very sparse. Like if you wanna run a dense model, you just want as fast-- you want it like a small bit of memory that can run it really fast.
- 46:18
So, uh, in that case, actually, you know, the best thing might be to just run it on the RTX, right? So like
- 46:25
this kinda depends where things go, but like I think at the end of the day, you'll want a-- you want both. You'll want both. Maybe even like other things, like maybe, for example, you know, you have a specialized chip or whatever, and maybe we'll have,
- 46:39
maybe we'll have SRAM locally. It's, uh, at the moment, it's like i-it's, it's... Well, it's super expensive, right? And it's also like the density is very low of the memory, so you can only have a little bit of it.
- 46:52
But maybe that's enough for a lot of use cases, right? Maybe you can run a smaller model really, really fast. Um, so like I just think it's not so much like Mac versus NVIDIA or whatever.
- 47:03
It's like, at the moment at least, it's like, okay, actually, you want l- bits of, bits and pieces of all these things. Um, and that's the, the way you get the most, um, price, the best price to performance.
- 47:19
So can you give us just like one minute on how do you achieve that today? Like let's say I just have this sort-- like I have infinite swarms of hardware around me around the house, but I'm not using anything hosted.
- 47:29
Yeah.
- 47:29
Am I running one gateway that's routing those requests as it scales-
- 47:32
Yeah
- 47:33
... like what's like-
- 47:34
So-
- 47:34
Like, you know, set and forget?
- 47:36
Yeah. So it's really awkward to do at the moment if you were just like using the existing tooling out there. Uh, but that's kind of one of the problems that Exo solves.
- 47:47
So Exo is just an app that you can install on every device, and it runs in the background, and it will automatically discover any other devices that are connected.
- 47:57
Um, and it works in a mesh network, so you can connect things however you want. And, um, basically, the software, like Exo software figures out the best way to distribute your model depending on like what hardware you have.
- 48:10
So that's our goal with, with Exo, is to make that really easy, right? And then, you know, for us, like having these heterogeneous setups and stuff is, you know, we can solve a real pain point there because it is, you know, e-even talking to people at NVIDIA that have tried doing this kind of stuff locally, it's like
- 48:29
very awkward. Um, and, uh, you know- It's-- You run into all sorts of networking issues and stuff like that. Like, we wanna make it, like, turnkey basically. So you install this app and that's it.
- 48:41
Connect things however you want. Uh, hopefully we can get a dem-demo and I can show you how that actually works.
- 48:47
I think these four are ready.
- 48:48
Those are ready? Okay. Yes. Good. So we have, we have, uh, four Mac Studios, um, and I'll, I'll show a demo of how that works exactly in a, in a second.
- 49:01
Um, but, uh, yeah, the, the problem is, like, I'm kind of, like, hesitant to say, like, "Oh, yeah, just go get all your hardware and do that," because like I said, today it's-- especially if you're just picking random pieces of hardware, it's like,
- 49:17
there's, uh, not much you can do, uh, especially, like, if the hardware is not-- doesn't have GPU, right? So if you're just taking, like, Raspberry Pis or something, it's like, you know, you, you-- Again, like, I always see these demos on Twitter of like, "Oh, look, I combined these Raspberry Pis and ran this big model."
- 49:35
And then, you know, you look into the details and probably they've, like, quantized it heavily, so it's not actually that useful and, you know, it's probably, like, really slow on the prefill, so, like, it's not, like, really usable.
- 49:49
Um, but, uh, if you happen to have GPUs laying around, then I think, you know, you could create a cluster and have something quite capable.
- 50:01
So, so what is your philosophy versus your address networking? I guess that's, that's the handoff I'd love to understand is just, it-- to me, it feels like the solution you're building is how do you get hardware useful in any context.
- 50:13
But in terms of exposing stuff over a network, right, we're talking about, like, using an iPhone or something similar, but it's-- you're doing an offload of inference to something if it's in your control.
- 50:22
Yeah, like, again, like, doing inference on a phone I don't think is-- that, that would-- is many years out, I would say. Um, so yeah, like, as far as how that would be used, yeah, you're gonna use, like-- So with Exo, it runs, um, it exposes a HTTP, like just a API endpoint on each device that you
- 50:41
run it on. And then, you know, you can use something like Tailscale. Like, we use Tailscale, um, to access that remotely. Um, they solve that problem really well, just, like, securely accessing your local device.
- 50:56
And then, you know, you can be anywhere, right? So I can be here and then have my cluster at home and chat to it with an app, right? Like, whatever app.
- 51:05
Is that bundled or are you bundling with all the native Tailscale primitives, or is it unbundled and you build that?
- 51:10
Uh, we're not, we're not, we're not bundling with Tailscale. Um,
- 51:17
but we might have, um, our own kind of solution for this soon. Like, um,
- 51:27
it's, uh, it depends, like, how much of a pain point it is, 'cause I, I, I don't think it's that hard to, like, set up your-- Like, especially if you're setting up your own cluster and stuff, like, then having your own Tailscale isn't, isn't that much.
- 51:39
But, like, something that we've, we've experimented with, you know, a while ago, and, uh, it might be something that we do. Um,
- 51:49
yeah. But I, I do think this, this is kind of how, how you, how you make something like this usable, right? Um, you know, you wanna be able to access it on the go.
- 51:58
You want, like... Like, what I imagine is you have your cluster at home, um, or maybe it's just one box, right? And, uh, it has access to all your data, right?
- 52:08
And you can securely access it from your phone anywhere, and it's running something like OpenClaw, so you can just, like, tell it things. I mean, you can do this today, right?
- 52:17
It's just expensive. Um, but you know, if, if you're talking about now that setup where you can just have that, you don't have to worry about, you know, these privacy concerns of, like, where your data is going, um,
- 52:31
and you can pay, like, five thousand dollars for that, you know, I think there's a pretty-- there's a lot of people that would buy that.
- 52:41
Cool. Uh, I think we can maybe try-
- 52:44
Yep
- 52:44
... demo.
- 52:46
Yep.
- 52:46
Uh, is there anything else? Yeah, I guess, like, one thing to keep in mind here is, like, I mentioned it earlier, but in the cloud you can batch. So, like, in the clouds, you've got, like, you know, you might have, like, a million users or millions of users using your model, right?
- 53:07
Which is the case with something like Claude or OpenAI. And so you kinda have the benefit of just, like, being able to take all these requests and efficiently schedule them, also known as batching.
- 53:19
Um, this is kinda how it works, right? So instead of doing inference, each inference sequentially, you can do them kind of together and you get these, like, really nice, um, economies of scale.
- 53:32
Uh, you can't really do that locally because y-you might be a single user. Um,
- 53:38
and uh, yeah, this is kind of how it works. You can basically climb up this... Like, as you increase the batch size, you can climb up this, um, this line which allows you to get better utilization out of the hardware.
- 53:50
Um, and, uh, especially with the data center GPUs, this works really well. Um, but I would argue actually it doesn't matter. Um, so, like,
- 54:04
uh, I think this is one interesting thing before we go to demo I just talk about, is just, like,
- 54:09
I have, like, three, um, reasons why I think, um,
- 54:15
this... So, like, obviously, if you can batch, your unit economics are gonna be way better in the cloud, right? Which is then, okay, like, if your unit economics are a hundred times better in the cloud than local, well, then local is always gonna be super expensive relative to what you can do in the cloud, right?
- 54:30
So that's kinda the argument. But I would say, um, actually, um- You know, there's gonna be some level of batching, even if you're a single user locally. So first thing is like multi-agent.
- 54:42
So, uh, I don't know why this was partly in Chinese, but, uh, this is like Grok 420. I guess it's not in beta anymore. Uh, so it's actually out.
- 54:51
But basically it uses four different agents. I think even more now. I saw-- I was using it the other day, and I saw like, you know, it shows you like what the agents are doing, and it seemed like there were more than four.
- 55:01
Um, but basically, you're actually running now as a single user. If you do one request, instead of it just being one pass through the model, it's like these agents that are kind of, you know, collaborating, and they're all running together.
- 55:13
So, uh, if this is the, the paradigm, then you actually... You know, you're, you're not gonna be running stuff at batch size one. Maybe you're running stuff at batch size eight locally.
- 55:22
And, you know, with the, the kind of characteristics of the hardware, especially with something like a Mac, then you're able to get really good utilization at that kind of batch size.
- 55:32
Um, second thing is, uh, test time scaling. So actually, the first thing is a form of test time scaling. Um, but I think there's a lot of interesting work being done on more general approaches that like are like search-based approaches.
- 55:48
So a lot of the current AI, I guess, paradigm is like about learning, right? Um, so like how do you scale these models, train them on loads of data?
- 55:59
But there's actually another thing that not many people are doing at the moment, which is like, you know, scaling, um, with search. And, uh, what I mean by search is just like simple thing you can do is like best of N.
- 56:12
So instead of running one, uh, pass through the model, you do like ten passes through the model and p- you know, have some way of picking the best one.
- 56:20
And you can train a model that kind of can basically verify these responses and score them. And then you, you know, you can do more sophisticated forms of search.
- 56:29
So like there was a Hugging Face... I don't know if I have it here. Is this it? Um, yeah, there was like some research that came out of Hugging Face that, um, basically showed that you can run a...
- 56:42
So I believe this is a 1B model. Um, and it's looking at the accuracy of the 1B model as you scale test time compute. So, uh, with various different methods, right?
- 56:54
And basically what you see is like there seems to be some scaling law here in terms of you can run a smaller model and do more test time compute, more search, and get the same performance as a bigger model.
- 57:05
So, you know, maybe the next paradigm will be something like this as well. And then again, you can start to batch, right? So instead of batch size one, it might be batch size eight or something.
- 57:16
Um, the first-- the third one, which I think is, uh, really interesting and I think it will probably hit an inflection point maybe this year, is continual learning. So
- 57:29
this is the idea that instead of just, you know, training your model upfront and then, you know, at inference time you're just doing forward pass through the model, you're actually training the model at inference time as well.
- 57:41
So, um, do I have something for this?
- 57:48
Yeah, so, so basically the idea here is, um,
- 57:51
what you might have is like everyone actually has their own version of their model weights based on how they use the model and, you know, their own data. And, uh, you know, this is kind of, um...
- 58:03
There's this whole area of like test time training, um, where, you know, there's been a few papers recently that basically have shown that you can get...
- 58:14
You've got like this long context problem at the moment, right? Or memory problem, where
- 58:20
the models forget and like, you know, you have to have these different sessions because you have limits to context and stuff like that. If you have, uh, test time training, you can solve a lot of these problems because now there's no such thing as context anymore 'cause you're literally like updating the model weights as you're using the
- 58:34
model. Now, this would completely break cloud unit economics because now you can't batch. So you wouldn't be able to do this anymore because each model is actually a different model.
- 58:45
So, you know, depending on how this lands, 'cause there's scenarios w- here where it might just be a small part of the model that changes with test time training.
- 58:57
But on the extreme end, if the whole model is changing, then you can't batch at all. So, you know, that would basically put, you know, local will, will get like 10X better in terms of-- relative to the cloud if this happens.
- 59:12
Uh, anyway, I'm gonna stop. I have more slides, but I'm gonna stop there and try and get a demo, if we're ready.
- 59:20
Uh-
- 59:20
We ready for a demo?
- 59:20
I don't know if the Wi-Fi is working. Just a second.
- 59:24
Okay. [laughs] Any more questions in the meantime? Yeah.
- 59:28
Just like what, what's your, like, general uptime from owning that setup,
- 59:33
this setup here? Like how much do you get like 100% of the time together?
- 59:37
Ah, this.
- 59:38
I saw some post on Hacker News that someone like wanted to rent out their spare capacity. So like how much do you personally-
- 59:44
Yeah
- 59:44
... using it 100% of the time full time?
- 59:46
No. Uh, so I, I still use, um... At the moment, I still use a lot of like Opus, uh, and, you know, um...
- 59:58
Hard lines.
- 1:00:00
Yeah, I, I use a lot of like Groq and stuff. Like, um, I do use local for some things, so,
- 1:00:08
um-- And it's growing as well, right? So like the set of use cases where, again, it's like good enough is kind of growing. But I would say utilization is pretty low
- 1:00:18
at the moment. So I guess, uh, what-- To, to repeat your question. So it's about like utiliz- uh, util- how much utilization do you get out of the cluster?
- 1:00:27
And also, what about renting out that spare compute?
- 1:00:31
Yeah.
- 1:00:31
So this is I think a really interesting idea as well. It's not something we're really focused on at the moment because it doesn't make sense if we don't have scale.
- 1:00:38
But if we have scale, let's say we have a million ExoClusters, then- And everyone's kinda using it for themselves, then the utilization might be quite low. Uh, in which case there's all this spare compute, all this idle compute that's just sitting out there.
- 1:00:52
So why can't we make use of that for something? Maybe it could even be, you know, for like a,
- 1:01:00
a, a volunteer like science kind of, um, problem, for example, right? Where it's like, okay, well, any spare capacity I have, it can go to like solving this scientific problem, um, which requires a lot of inference, right?
- 1:01:16
I think this is a very interesting idea, and it's something we'll revisit once we actually have, you know, a scale. Um,
- 1:01:24
yeah, I think, um, there's, there's a possibility here as well, where actually utilization goes up a lot because e- even for your local cluster without this. Because, uh, imagine like you can actually get frontier-level performance locally, right?
- 1:01:38
Well-- And you can just give access to all your data. Like, what I would personally want is, um, I would just want this... Like, because I'm not paying for tokens, I just want this thing to run all the time.
- 1:01:47
And like maybe it can like proactively, you know, tell me about things. It can be scanning the internet, um, for things that are relevant to me. It can constantly be like, you know, it could be thinking about, okay, future direction of Exo, maybe things to look out for.
- 1:02:03
Um, you know, you basically have this like twenty-four seven agent that can be looking out for you and as like a companion. I think that would mean utilization goes up a lot, right?
- 1:02:16
And, uh, you would just want this thing to be running all the time. So in that case, maybe there won't actually be that much, uh, idle compute out there.
- 1:02:25
Um, but, uh, I think it's interesting. Yeah, once we reach scale, then we start to think about these questions of, oh, actually we have the equivalent of the biggest data center, biggest data center in the world.
- 1:02:37
It's just distributed across the globe. You know, um, is there some way to make use of that? And, uh,
- 1:02:45
yeah, I think inference is quite easy actually to do in this setting because you don't really need much, uh, communication happening between different machines. Um, especially if, you know, if people are already running these frontier models on their hardware, then they already have the capability to run that, right?
- 1:03:03
So you don't need to do any kind of clustering over the internet or anything like that because they already have this capability. So then it's a-- just a matter of distributing work, uh, in like kinda data parallel way, and that's really easy to scale.
- 1:03:16
So yeah.
- 1:03:18
Uh, on that topic, do you think there is like some parallelity back when everybody was mining like crypto at home and was building their GPU racks? Because then-
- 1:03:30
Yeah
- 1:03:31
... more and more people started doing it, prices dropped. You could argue prices per token dropped, right? So it's not worth it for lots of people.
- 1:03:40
Yeah. That, that's-
- 1:03:41
My argument would be it's never gonna be worth it to rent out your hardware to other people if you're just-
- 1:03:47
Yeah
- 1:03:47
... want, if you just want financial gains. If you're doing it-
- 1:03:50
Yeah
- 1:03:50
... for nonprofit, then it's-
- 1:03:53
I, I, I think I mostly-- So the, the question was like, uh, do you think it would be like crypto mining, where originally it was economical to run your own GPUs and mine Ethereum, for example.
- 1:04:08
Uh, but then eventually, you know, you had these large scale operations that
- 1:04:14
made it, um, just not worth it to mine y- yourself. Um,
- 1:04:21
yeah, I, I think like, um, there's again a few things here. So-
- 1:04:35
I mean, some things are different, right? Because-
- 1:04:38
Yeah
- 1:04:38
... running a local model, you have all the benefits. Your data is not going to the cloud. Crypto miners would just run at one hundred percent all the time.
- 1:04:47
It's not like the same incentive, right? But-
- 1:04:49
Yeah
- 1:04:50
... there are some things that seem quite similar in the first like if you just look at it.
- 1:04:56
Um, yeah, I, I, I think like that's why I, I said like it only makes sense once you r- we reach scale because, um, I don't think any project that just is-- its goal is to, you know,
- 1:05:11
give you money for renting out your compute is gonna work. And I think there's already been quite a few failures, you know. Especially like, I think a lot of the stuff out there right now is kind of at the wrong level of abstraction, where you're renting out the hardware.
- 1:05:25
Whereas what you should be renting out is the use case, right? So like it should be, you can go higher and higher up, up in the stack and you could, for example, charge for tokens, or you could charge for like something even higher level than that, like a task, right?
- 1:05:41
And I think the higher level you go, the more you can get creative with how you make use of the hardware. And for example, maybe it turns out that, okay, this compute that's out there, um, people are willing to rent out for like very cheap, right?
- 1:05:59
Um, because, uh, it's literally just spare capacity that's sitting there, right? So they wouldn't be getting anything for it otherwise. In which case, you know, if you can,
- 1:06:11
you know, if you can basically, for example, build something that's higher level that, let's say, is something like, uh,
- 1:06:19
imagine an API where you can just submit a task and it will get done in the next twenty-four hours, right? So it's not latency sensitive, doesn't necessarily-- If it fails, it can just be retried.
- 1:06:29
Something like that, you could maybe have an API that's really, really cheap that runs on this kinda network. Um, but yeah, I, I think like the first thing is first you need-- first of all, you need like people to have a lot of capable hardware, and I think local AI will be a big catalyst for that, and
- 1:06:46
then you need scale. Right? There's no point doing this on 100 Macs. It's literally like negligible amount of compute. You know, it becomes interesting when you start to s- you know, look at like, okay, what if we had a gigawatt of compute?
- 1:07:01
It's like, well, you know, how else are you gonna get that compute? There's a few companies that have data centers that need all these-- You have to go through all this-- You have to raise so much money, you have to have these permits and stuff.
- 1:07:15
Like, this is like kind of an- another way to access, like large amount of compute. So it could be interesting. Um, but yeah, it's, it's not really like something that we're thinking too much about until we reach scale.
- 1:07:31
Yeah. I think we're almost-- Like we're quite close to time, so I wanna show the demo. Uh-
- 1:07:36
Yeah, you can do it.
- 1:07:36
Yeah. Is it ready?
- 1:07:37
I think you can SSH into James' Mac and then start it.
- 1:07:42
I have to SSH?
- 1:07:44
Uh, no, I know you can look at it on Tosca, if you're on Tosca.
- 1:07:47
Should be.
- 1:07:48
There you go. Go ahead. I'm just trying to set up Spark so then I can point there. Unless they're not on the-
- 1:07:56
Okay
- 1:07:56
... Mac as well.
- 1:07:57
No, it's working.
- 1:07:58
Okay, cool. That's good.
- 1:08:00
Cool.
- 1:08:01
Get my Mac.
- 1:08:02
So, uh, ba, ba, ba. Okay. So yeah, we've got four Macs here, um,
- 1:08:13
and they're all running Exo. So basically, how that works is just a macOS app that you run in the background on each machine, and they, um... That's all you do.
- 1:08:27
So like you just install this app, runs in the background, and what happens is they're connected by Thunderbolt. So,
- 1:08:37
uh, I don't know if it's easy to show. You can see it better if I turn it around, I guess.
- 1:08:47
So basically, um, each Mac is connected to each other Mac. So they're connected with Thunderbolt 5, which is basically like, it's kind of just like, uh, a wrapper around, uh, PCIe, so it's like pretty fast.
- 1:09:06
And, um, we actually did some work, uh, recently, um, to integrate low latency RDMA into Exo. So prior to that, the latency between the Macs, because this is just like consumer hardware, it's running this bloated macOS, uh, it was really slow.
- 1:09:29
So you would get like 300 microseconds, like 0.3 mi- uh, milliseconds of latency if you wanted to send data between the Macs. Now, problem with that is, um, it doesn't really allow you to split workloads in a efficient way where you can actually scale up.
- 1:09:47
So what you wanna do if you wanna actually get a speed up is you wanna kind of distribute each layer across, um, your machines. Um, if you have 60 layers in a model, which is the case with Kimi or DeepSeek, then that's, um, at least 60 times that you have to synchronize every time you run one, like,
- 1:10:10
um, generate one token. So, and it actually, in the case of tensor parallelism, which is, uh, a way of splitting up, um, you know, your, uh, tensor operations across these machines, you have to do two synchronizations per layer.
- 1:10:24
So that's 120 synchronizations for a model like Kimi. If that's gonna-- Let's say it takes, you know, 0.3 milliseconds, that's a lot of time. That's like 40 milliseconds that's spent just on, uh, the communication, right?
- 1:10:37
So that would be really limiting in terms of the speed up that you get. And in fact, prior, like back then, you wouldn't even get a speed up. So by clustering stuff, like the benefit is you have all this memory that you can split the model across, so you can fit bigger models, but it would be slower.
- 1:10:53
Um, anyway, long s- uh, long story short, with RDMA it's 100 times faster, so it's like single digit microseconds, which now, instead of it being like, you know, 30 mi- uh, 30 milliseconds of communication, it's like less than one millisecond, right?
- 1:11:10
Which is perfectly fine if you wanna kind of scale up these big models. So GLM 5.1 just came out yesterday.
- 1:11:18
Is it a trillion parameters GLM?
- 1:11:20
Uh, 1, 1.2.
- 1:11:22
One point-
- 1:11:23
Uh, wait.
- 1:11:24
Oh, I think it's trillion parameters. Um, and, uh,
- 1:11:29
yeah, basically, um, how the hell are you gonna run a, like trillion parameter model?
- 1:11:34
Yeah, trillion, trillion.
- 1:11:36
Yeah. It's like massive. Um, so well, the idea is, uh, you wouldn't be able to run it on a single device, it's just too big. So we're clustering it across these devices and in combination with this RDMA capability, can actually run it faster than it would run on a single device.
- 1:11:54
So, uh, it's already loaded into memory. You can see it here. So GLM 5.1 and, uh-
- 1:12:02
Thank you so much
- 1:12:03
... yeah. The way Exo works is you can spin up these instances of models, and an instance is just, um,
- 1:12:11
one configuration of a model that's running on your cluster. So, um,
- 1:12:18
it's already loaded into memory, and you can see like the utilization on each the-- memory utilization on each of the machines is like 112 gigabytes or something. So this is, this, because of like, it takes like really long time to download these models, we just downloaded the 4-bit one, um, which is still pretty big.
- 1:12:35
It's like, what is it? Almost 400 gigabytes. Um, but, uh, the full model here would be like 1.5 terabytes. Just, you know, uh,
- 1:12:48
it literally came out yesterday, so we just downloaded the 4-bit one. But, um-
- 1:12:53
Yeah, I can just, uh, chat to the model. So, um, does it know about
- 1:12:59
AI Engineer? Curious. Alex, quick question about this. Have you converted into MLX before you-
- 1:13:07
Sorry? Have you converted this model to MLX? Yeah, yeah, we did that yesterday. Okay. So basically, uh, when a model comes out, like there's always this rush to like, make it compatible with the hardware.
- 1:13:18
So we did-- Uh, Leo did that yesterday, um, downloaded the model, converted it. Fortunately, with this one, it was quite easy because,
- 1:13:28
um, basically, um, GLM 5 was already out, right? So this is just like another checkpoint of the same model, so the architecture's exactly the same. We already had support for it.
- 1:13:39
So with this, we just had to convert the weights. Didn't have to-- Normally, if a new model comes out, like Gemma 4- Yeah ... we had to do a lot of work to get that working.
- 1:13:47
Especially like Gemma 4 is quite different in its architecture in that, um, the way like the KV cache works is very different to like any other model. Working. So we-- That was a bit of a pain to get w-working and it-- The-- Yeah.
- 1:13:59
Anyway. Uh, so yeah, that-- Did it know about it? Refers to a prominent community.
- 1:14:05
It's founded by Swix. Yeah. [chuckles] I guess this is a recent checkpoint, so it knows about, uh, it knows all about this. I don't know how long AI Engine has been around, but, um, yeah.
- 1:14:14
So this is running across these four machines, and if you, uh...
- 1:14:23
So if you look, uh, s- I don't know how big... Can you see it well? Like, you can see basically the utilization on all of these machines is, um, is like, uh, it-it's, it's 100% on all of them 'cause all of the machines are u- being used in parallel.
- 1:14:40
So this is tensor parallelism with Audima. It's a little bit like you can see the response is a bit choppy, but that's because this is-- I think I'm connecting over Wi-Fi.
- 1:14:48
Yeah.
- 1:14:49
Right? So like it's actually just the connection between my Mac and the cluster. But here all I've done is like this is running Tailscale. So we have, um,
- 1:14:59
a, uh, machine that has a host name James and, uh, you know, just-- I'm just basically connecting to that, uh, to the dashboard over Tailscale. Um,
- 1:15:13
yeah, that's, uh... I don't know if we can maybe run a different one as well. Just depends what we have.
- 1:15:21
Uh, da, da, da. Let's try this one.
- 1:15:31
Qwen 3.5 should be fine, right? Can I try-
- 1:15:34
Uh, yeah.
- 1:15:34
Yeah? [chuckles]
- 1:15:35
Uh, yeah, that should be fine.
- 1:15:38
So like I said, you can, um, you can basically have like different instances,
- 1:15:44
uh, of the s- of, um, of different models. Uh, and you can run these all together and, um, this one is downloaded on a few of them.
- 1:15:57
So I can try this. So this will be a small-- much smaller model, so it should run a lot faster. Should be downloaded on these. So if I launch that, um, it will load the model into memory.
- 1:16:13
So you can see on the two machines that it loaded it on, the memory went up a little bit. So on Mike and S13, the top and the bottom ones, it's now like higher.
- 1:16:22
And, uh, once it's loaded into memory, you can just, uh, chat with it. Uh, and it should be, you know, much, much faster than the other one. So again, it's a bit choppy 'cause of the Wi-Fi, but, you know, you can see the tokens per second is like 77.
- 1:16:36
Um, and, um, yeah. Uh... So like if I get it to do something a bit longer, you can see basically the utilization is going up on these two machines 'cause they're w- again, working in parallel.
- 1:16:55
Um, so yeah, it's, it's pretty simple. Um, but like, uh, a lot of the complexity is sort of in, um, just like making it this one app that you can install and, you know, not having to do any of this network setup, and it just kind of figures out the best way to just use the model.
- 1:17:13
I also wanna show-
- 1:17:14
Yeah. Uh-
- 1:17:16
Uh-
- 1:17:16
On setting up the network.
- 1:17:18
Okay. So yeah, I also wanna show like why we bought a Spark. Like, so there is, um,
- 1:17:25
uh, also the ability to split models across heterogeneous hardware, right? So you can take like Spark and, uh, Mac and, uh, in this case, obviously, like I said before, this has a lot more compute.
- 1:17:37
So there's some interesting things we can do in terms of splitting up the model in, you know, more kind of, um, granular ways than just like the way that we're doing it here, which is, you know, tensor parallel.
- 1:17:50
Um, cool. While, while we wait for that, I can go back to the talk. I mean, we have 15 minutes, so...
- 1:18:00
Also, if anyone wants to like try-- We can try this. Like if anyone wants to try, if they have Exo or wanna install it, it's like exolabs.net, and we could try also then adding that to the cluster.
- 1:18:13
So because it's like, you know, super easy to, um, to add, uh, devices. I'm not sure if it'll work over the Wi-Fi 'cause it depends on the,
- 1:18:26
the, uh, firewall, but we can always connect it with Thunderbolt, I think if, if we have another cable, if someone wants to try it. I got it.
- 1:18:35
Does it copy the-
- 1:18:35
You got it.
- 1:18:36
Does it just only copy the models to the memory of the other machines?
- 1:18:39
Yeah, yeah. So, uh, okay. So is this running the app on the-
- 1:18:47
Oh, yeah.
- 1:18:47
Yeah, it is running the app, right? Yeah. I saw that. Like the latest one? Yeah, I think so. I think that's the latest. So let's give this a try.
- 1:19:06
See if it works over Wi-Fi first, so- I've got the same Wi-Fi as you probably. Yeah. So if I open the dashboard, which should...
- 1:19:18
Open in Chrome, I think, so it's not working. It's not working in Chrome?
- 1:19:24
Yes. So yeah, the idea is like you can just connect these in any way you want. So it's, it, it like works through a mesh network, so, you know, we should in theory [chuckles] be able to just connect this MacBook and have it join the cluster.
- 1:19:36
Um, yeah, looks like it isn't working over Wi-Fi. It must be something with this shared Wi-Fi, but we can try plug in. Should be able to plug in, right?
- 1:19:48
Yeah. With that one. [chuckles] Yes. Let's see. [chuckles] And the other one.
- 1:19:52
Do you want plugin?
- 1:19:54
Yeah, just disallow- Oh, you don't have any models as well, so that might be... But we will do a small one. We'll do a small one. Small one.
- 1:20:00
Uh.
- 1:20:01
Yeah, but Gemma 4 GPT 120 billion, I think so. Oh, you have it? Yeah. It's on the sheet. Gemma 4? Yeah, Gemma 4. I don't know if that will work 'cause on this version of the app-
- 1:20:12
That was-
- 1:20:12
Um, we'll, we'll try it.
- 1:20:15
Yeah, it is. Um, the weird thing is, uh-
- 1:20:20
So one weird thing with this hardware is you should never use this port next to the ethernet. Um, we ha- we don't really have a good explanation for why, but it just doesn't work.
- 1:20:31
Um, but apparently Apple is fixing it, so.
- 1:20:33
You're not gonna be able to connect via the hardwire.
- 1:20:35
Yeah. You can use any other port though. And all of them are... So like you need Thunderbolt 5 to do RDMA. What, what MacBook is this? Um- Do you know what chip it is?
- 1:20:44
M4. M4. Max. M4 Max. Yeah. So this has, this has Thunderbolt 5. Uh, do you have RDMA enabled? [chuckles] 'Cause you have to, you have to actually... So-
- 1:20:53
You, you can use the pipeline break.
- 1:20:55
Yeah, yeah. I, I just thought it'd be cool if you can do RDMA, but, um,
- 1:21:02
so if we wait a bit, they should discover each other automatically,
- 1:21:07
hopefully. Yeah, you have Exo Thunderbolt, so it should...
- 1:21:21
Yeah, okay, it's working. So you can see like, um, actually the first thing that...
- 1:21:28
So this is just like an architectural thing about Exo, but the first thing that will happen is when a new node connects, it basically catches up on the whole history of what happened in the cluster, 'cause we use, uh,
- 1:21:39
um- Event sourcing ... event sourcing, which is, uh, basically, it's like commonly used in da- distributed databases as well, like where each machine basically writes its own, uh, append-only log, and then they kinda get merged in some way.
- 1:21:54
And the reason we do this is because of, um, you know, consistency guarantees across the cluster. So if you're working with something that's quite dynamic like this, then devices can come in and out, you know, uh, basically whenever.
- 1:22:09
Like a device could power off, like this could go t- to sleep or whatever. If you're working with that, then how do you, like, guarantee that if a request is going on, it's not just gonna get lost, right?
- 1:22:18
So that's kind of, uh, from the very start is like how we architected this. Um,
- 1:22:26
so it's just replaying that. Um, but yeah, I guess... I was talking about RDMA. So this, uh, everything... So like M3 Ultra, which is these, M4 Pro, M4 Max, M5 Pro, M5 Max all have Thunderbolt 5.
- 1:22:41
Thunderbolt 5 is required for RDMA. Um, and yeah, look, [chuckles] you can see it's literally replaying all the stuff that we just did. I don't know if it's hard to see, but...
- 1:22:53
Um, so, uh, yeah. But the issue right now is like you basically have to boot into recovery mode in the Mac to enable it, 'cause it's more of like a developer focus feature, and Apple is a consumer company, so, uh, they don't want this like on by default right now.
- 1:23:12
Um, so this one doesn't have it enabled, but what we can do is like, you know, try basically creating, um, an instance, uh, just over the TCP/IP. So it'll be slower, but it should, uh, it should be possible.
- 1:23:30
Uh, yeah. It's got to the point where we were running the second model now, so it's almost there.
- 1:23:37
I think it might... There we go. Okay, so now you can see like this has popped into the cluster, and you can see it's only connected to this one, um, which is this Thunderbolt connection.
- 1:23:46
So Exo like maintains a live view of the physical topology, and with that it can figure out the best way to distribute the model, basically. So... Oh, okay. We did get a warning here that there's incompatible macOS versions.
- 1:24:00
So this is on different macOS, but we can still try. Um, so sometimes you have... It looks like this is on macOS 26.2- Yes ... whereas these are on 26.3, and sometimes that can be an issue, but we can still, um, we can still try it.
- 1:24:13
So you said you had a model, right? We can look at what you have. Uh, ba, ba, ba. Okay. Wait, can I put this on the... Can I put this- Yeah,
- 1:24:25
yeah. Will this work? Does it have a HDMI, this Mac? Uh, I don't think it does, right? No. It has an adapter. Do we have an adapter? Do you have an adapter?
- 1:24:40
Okay. I'll show the dashboard from this machine. I can still do that.
- 1:24:49
So yeah, you can see now the MacBook is there, but it says that there's incompatible macOS versions 'cause this is on 26.2, whereas these are on 26.3. Uh, but actually I can do it from here.
- 1:25:00
So I don't even need to do it from there. So that's a nice thing as well. I can access-
- 1:25:05
You know, any of these machines are running the same API endpoint and the same dashboard. I can just send requests to any of them. So you said you had a model.
- 1:25:14
I don't see it though. I don't see you ha- you don't have any models. [laughs]
- 1:25:19
Well, um-
- 1:25:20
I mean, it's fine. We can, we can download a model. We can just do a small one.
- 1:25:23
Yeah.
- 1:25:24
So-
- 1:25:24
There were so much, uh... so many models downloaded with MLX and, um, with LM Studio as well.
- 1:25:30
Oh, it's downloaded to the exo/models folder.
- 1:25:34
Ah, I see. Ah, okay. So you have it from using MLX. Yeah.
- 1:25:39
Yes.
- 1:25:39
So that puts it in a different directory.
- 1:25:41
And LM Studio as well.
- 1:25:42
Um, it's fine though. Like, what we can do, so you can filter here. So you can basically select any set of nodes, and then it will filter configurations by those nodes.
- 1:25:50
So if I filter by ones that contain the, uh... Wait.
- 1:25:58
Is that filter working? Or whatever. That, that looks correct. So if I launch it onto here,
- 1:26:06
see how it plays with the, with the different macOS version. Hopefully it works. Oh, didn't like that. [laughs]
- 1:26:15
I don't think it liked... Oh, no, it is just downloading. Okay. So when you launch a model, obviously first it needs to be downloaded, um,
- 1:26:24
which is actually quite a big pain point with this stuff. So if you're running stuff locally, obviously you need to have the entire model weights. That's means, you know, sometimes you need to download like a terabyte.
- 1:26:34
Um, which, yeah, you just need high-speed internet basically. Um,
- 1:26:42
and obviously in, in the case of, like, when you're splitting the model across machines as well, like, you basically need pretty much the whole model weights on each machine.
- 1:26:51
You can maybe get smarter with, like, putting parts of the model on different machines. Um, but it's kind of difficult because now if a node dies, then the part of the model that it's responsible for will be different.
- 1:27:04
So you basically need the whole model on every machine. Uh, so if this finishes downloading,
- 1:27:13
we can run an inference. So this is Qwen 0.6B, which is basically as small as it gets. It's only like, what is it? Zero point, 0.3 gigabytes.
- 1:27:27
Shi-
- 1:27:28
Can we see the-
- 1:27:28
Shitty Wi-Fi, I guess.
- 1:27:29
Can we see this, uh, existing model, uh, with Exo, Exo?
- 1:27:33
You can use it, but it's, it would need to... So, like, right now, you have to, like, manually migrate the models over to Exo. So if you ha- you're using a different app like Llama CPP, it won't be, uh, possible to easily, uh, transfer it.
- 1:27:46
But we're gonna add actually that so that you can just... If you have models already from a different app, then it will just... E- Exo will recognize it. But it's downloaded, so it says it's ready.
- 1:27:57
Uh, so this should be running on the Mac. Yeah, so you can see...
- 1:28:05
Let's do something longer. Okay, maybe the utilization here is broken, but, uh, yeah.
- 1:28:30
Basically, this is running on the, the MacBook here. So yeah, the idea is like... I mean, you get the idea, right? You can basically just, like, run this application on any of your machines and connect it in any way you want, and then Exo will figure out the rest.
- 1:28:45
Uh-
- 1:28:47
We have, uh, the Spark on design.
- 1:28:49
You have it?
- 1:28:50
Yeah.
- 1:28:50
Prefill decode? Okay. So next thing I wanna show you is like, okay, what if it's not just Mac? So what if you have different hardware? So I'll just quickly show you, like, you have a Spark here, and you have a MacBook.
- 1:29:03
Do you wanna connect that to the HDMI?
- 1:29:04
Uh, yeah, sure. Uh, which HDMI?
- 1:29:06
It's this. Do you have HDMI port?
- 1:29:08
Yes.
- 1:29:08
Okay.
- 1:29:11
I'm just trying to get Wi-Fi to work again. [laughs]
- 1:29:23
But i- is it running? Like, can we-
- 1:29:26
Um-
- 1:29:26
... try it or?
- 1:29:28
I think it... Sorry. Um-
- 1:29:31
Yeah. So the idea is run prefill on there, run decode on the, on the MacBook. This thing has twice the memory bandwidth. The MacBook has five hundred and forty-six gigabytes per second.
- 1:29:42
This has two hundred and seventy-three gigabytes per second of memory bandwidth.
- 1:29:45
What do you mean with-
- 1:29:45
Whereas this has four times more compute. Um, so, like, the ratio there, difference in the ratio is, like, 12x. Um, so the idea is basically,
- 1:29:59
uh, you run... Like, for really large prompts, uh, you ge- actually get a big speed up by, you know, not just running it on your MacBook, but also splitting it up across here.
- 1:30:08
But you don't wanna run the decode on here 'cause it's slower for memory bandwidth. So basically, uh, yeah, you, you split up these two phases, and then there's, like, some complexity to, like, streaming your KV cache 'cause now the KV cache only exists on here.
- 1:30:25
Um, so you need to somehow get it to the MacBook, so it can do its decode, right? 'Cause the KV cache, it reads it every time it does a pass through the model.
- 1:30:34
So the way we do that is, uh, they're connected by 10 gigabit Ethernet. I don't know if you can see, but, uh, it's 10 gigabit Ethernet. The MacBook doesn't actually have Ethernet port, so we use a 10 gigabit Ethernet adapter.
- 1:30:46
And, uh, so it's a bit of an annoying actually because ideally, like, you can just have some USB-C cable or something, connect them with that. We don't have that working yet, but...
- 1:30:56
Well, we kinda have it working, but it's not really, like, production ready 'cause you have to run a bunch of scripts to get that working. Um,
- 1:31:05
but, uh- Is the Wi-F... Is it going over Wi-Fi?
- 1:31:13
Yeah. There's a bit, there's a bit more feedback.
- 1:31:18
Okay. Uh, is it... Should we skip that, skip this?
- 1:31:23
Yeah, I think we should skip this.
- 1:31:25
Okay. You can try a little bit longer. We have maybe five minutes, and then-
- 1:31:30
All right.
- 1:31:31
Uh, yeah, the issue is, like, um, for some reason it's sending the KV cache over Wi-Fi at the moment, which then the bandwidth becomes a bottleneck. So what you ne- that's why we need 10 gigabit Ethernet, 'cause you're sending this KV cache over.
- 1:31:45
You don't want that to, like, uh, block, you know, you being able to compute the decode phase, right? So you need to basically overlap computation and communication and, like, stream the KV cache over to the MacBook, and it needs to be fast enough so that you can fully overlap.
- 1:32:03
If it's not fast enough, then it would be sequential. So you'd li- do your prefill, then it will still be sending, and then it would do the decode, which is, um,
- 1:32:11
which is gonna be slow. Uh, can I have the HDMI back?
- 1:32:14
Yep. Sure.
- 1:32:16
Yeah, I guess we're almost... Well, we're at time now. So, uh, I will
- 1:32:24
close off with one thing, which is, okay, like, I'm telling you all this stuff, but why is this not already, like, kind of more understood or more known? Well,
- 1:32:35
basically, like, the best source at the moment for this stuff is, like, Reddit or Twitter, and everyone says different things. So, like, you know, I to- I, I mentioned before, like, you know, there's maybe a lot of threads that you see or, like, people that run experiments.
- 1:32:50
Now you have all these citizen scientists as well, which I think is cool, uh, because, you know, you have really capable AI tools, so people can quickly experiment and try things.
- 1:32:58
But then the flip side of that is, like, there's a lot of noise. And, uh, if you don't truly understand what's going on, then you might actually think you have a result, um, and, like, the LLM is telling you, "Oh, you've made a breakthrough."
- 1:33:11
Um, but actually it's, you know, it's, it's not so interesting, and it's, it's not really usable. Good example of this is, like, people heavily quantizing models. So if you, like, quantize a model to one bit, then you're better off just using a smaller model and not quantizing it.
- 1:33:30
Um, so i- it's not... Like, these, these models are not very useful, right, at one bit. So, uh, you might see some threads, for example, about, "Oh, I ran Kimi on a MacBook or something," but, you know, it's, like, the one-bit version, and maybe they also pruned it or, like, you know, instead of it activating eight experts
- 1:33:49
of the model, it's activating two experts of the model or something, you know. And, like, what we wanna do is actually, like, bring some transparency around this. So we have loads of hardware.
- 1:33:59
Um, yeah, we, we, we, we have probably the most, um, like, hardware specifically for this purpose. Um, and, you know, our idea is we wanna publish benchmarks in the open that show you, "Okay, what performance will I get on certain hardware?"
- 1:34:18
And, um, this stuff is changing really fast as well, right? So, like I said, I think there's 100X, you know, in there. So this will also be a source to be able to track progress, right?
- 1:34:29
So you'll be able to see, like, okay, you know, how is... how are things getting better? Is the software improving? Is the hardware improving? You know, are the models improving?
- 1:34:37
So this would be very soon. Uh, you know, let's say within the next month we'll come out with
- 1:34:44
a website that basically has thousands of benchmarks on there. We're already continuously-- We're, we're getting the data now, so we're continuously running these benchmarks, different models, different quantizations, different ways of pruning the models.
- 1:34:55
We also-- We wanna pair this not just with, like, raw performance of, like, tokens per second and prefill time, but, like, also the quality of the model, right? So, for example, intelligence per joule, right, is one way of looking at, looking at that.
- 1:35:10
Um, so basically you'll be able to select a budget, let's say $10,000, and see, like, a Pareto frontier of all the local, uh, setups at that budget. And, um, it will show you, like, you know, if you want really good quality, you might have to, like, compromise a bit on the performance.
- 1:35:27
So if you want, like, to run GLM 5.1, then maybe you're gonna get 20 tokens per second. Um, but maybe, you know, for you it's fine to use a smaller model, in which case maybe use Gemma, and you get, like, 100 tokens per second.
- 1:35:41
And all of these kind of exist on different points on this Pareto frontier. Um,
- 1:35:46
so yeah, I, I guess I'll just, uh, close with that. Um, so unless we have this or-
- 1:35:53
I don't.
- 1:35:53
Okay. [chuckles] Uh, yeah. Um, yeah, I, I think... I don't know if... I think we're basically at time, so I'm not sure if we have any time for more questions, but yeah.
- 1:36:05
Thank you so much.
- 1:36:07
Thanks. Thank you. [audience applauding]
- 1:36:15
Uh, can I ask how we are for time if anyone does any, have, have any quick questions? Or should we close it?
- 1:36:21
If you want, if you want to get more-
- 1:36:22
Maybe one question if anyone has one. Or not. [chuckles] Yeah?
- 1:36:28
Um, so Simon Williamson posted something recently about, um, somebody who's used, uh, Andrej Karpathy's AutoResurge loop in order to optimize the model for a particular spec of hardware.
- 1:36:45
So, like, I think they were running an M3 Mac 48 gig, and they were using load, like load the memory, loading into memory at, at each point off the, off the disk.
- 1:36:54
Have you done anything to optimize for your hardware-
- 1:36:58
Oh
- 1:36:58
... setups to, to, to make the models-
- 1:37:02
Yeah
- 1:37:03
... match, match what-
- 1:37:04
Yeah
- 1:37:05
... what you've got?
- 1:37:07
Yeah. I, I guess I'm on the more cynical side with that work. Uh-
- 1:37:15
I have seen a lot of stuff out there that is kind of, again, like cherry-picked. And I think especially if you're using auto-research without really understanding what's going on, you might-- it might not even be that you're trying to like hype this up, but you actually believe that you've made some breakthrough.
- 1:37:36
Um, so like I'll give you an example. I've seen like stuff where it is just a one-bit model running, and they say like, "Oh, look, we have, you know, Kimi running from disk," um, at-- And it's not even that good performance, right?
- 1:37:50
It's like five tokens per second or something. But, you know, and then, and then also you have a-- there's many layers to it, right? So like auto-research might make a change that really impacts the perform-- like the quality of the model, and you don't realize it because you think-- it tells you like, "Oh, look, it's running."
- 1:38:06
So I think auto-research is a really powerful tool, but my,
- 1:38:12
my opinion on this is you still need to follow the scientific method. So you need to have like a well-reasoned hypothesis, hypothesis. You need to then e- run experiments to test it and then,
- 1:38:27
you know, uh, iterate, basically. So like if it's just like, "Oh, hey, auto-research, go off and figure out how to make this fast," it won't work. I think that's, uh, pretty much a data...
- 1:38:38
It's a slot machine at that point. It's a slot machine. It's like-- And it might be addicted-- addicting, and you might think like, "Oh, wow, I've-- I'm making real progress here."
- 1:38:48
But I think actually the set of things it can come up with, with slot machine approach is very small, and it will be very sparse as well. Maybe e-eventually.
- 1:38:58
Maybe it will come up with certain things, um, you know, I... if enough people are running it, enough compute you throw at it. But it's like brute force. It's basically like brute force approach.
- 1:39:08
Um, so like I really like auto-research and I, I u-using it as well. Like, I think it's a really powerful tool, but it needs to be
- 1:39:18
used in the context of scientific method where you actually understand what's going on.
- 1:39:24
So yeah. The-- But then specifically on the disk thing, I just don't think it's interesting. I think, um, obviously the idea of like having memory hierarchies is, is, is interesting.
- 1:39:34
That, that's just like a fundamental thing about, um,
- 1:39:39
compute, I guess. But, um, specifically disk, it's just too slow. It, it, it-- There are just like fundamental constraints there. You're better off focusing on the memory problem. Like that's gonna be much more cost-effective.
- 1:39:53
Uh, unless you do some like crazy setup where for, for some reason, you know, stacking a bunch of SSDs is like way cheaper than, um, than using s- normal memory, right?
- 1:40:08
Which might-- Maybe that w-
- 1:40:11
Market, like when it went up.
- 1:40:12
Yeah. Like with memory prices and stuff, maybe that would become interesting, but like it's gonna be-- It's not gonna be a MacBook, right? Because MacBook has one disk, and it's not that fast.
- 1:40:22
Um, but like I have seen some people looking at the problem of like, "Oh, what if we put like sixty-four SSDs in a box and, you know-
- 1:40:34
With RAID zero.
- 1:40:35
Yeah, with RAID or something. And, uh, and then start to use these techniques where you can be smart about loading in experts. You know, you can-- There's a lot of research on like predicting which expert is coming next, for example, so you can like optimistically load it and then, you know, maybe you're right ninety percent of the
- 1:40:52
time, wrong ten percent of the time, then you have to load it again. And it's like, you know, you, you can do a lot of clever tricks there. But like for me, like okay, like do those tricks, but do it with normal memory.
- 1:41:02
Don't-- Like, why are you, why are you looking at, you know, doing this with disk?
- 1:41:09
Yeah. Cool. Thank you.
- 1:41:11
Okay. Thank you.
- 1:41:13
And Alex, in the background, it worked.
- 1:41:15
Oh. [laughs] Okay. We-- [laughs] It did work. Great. Do you wanna sh-show it then?
- 1:41:20
Yeah.
- 1:41:20
Like, so, so basically, uh, we can show like before and after, right?
- 1:41:24
Yeah.
- 1:41:24
Can we show before and after?
- 1:41:24
Uh.
- 1:41:24
So, so one is like the MacBook running on its own. Um, and you can see like this is a pretty large prompt, so it's a paper thirty-eight kilobyte or like, well, two of them, so...
- 1:41:35
Uh, I'll just do the after first, but-
- 1:41:38
Okay.
- 1:41:38
So.
- 1:41:38
Doing the after first. So the after is using both together, prefill decode. So it will do the prefill on the Spark and the decode on the MacBook.
- 1:41:50
And, uh, like one thing to keep in mind here is this doesn't make sense if you're doing small prompts. If you're just saying hello, there's no speed up. But for like large prompts where the prefill time also goes like quadratically with the size of the prompt, right?
- 1:42:05
For most like architectures of model. So if you're doing really big prompts, it's gonna actually take a significant amount of time of the inference. So in this case...
- 1:42:16
Yeah. So there's... I don't know. This is like a hundred kilobyte prompt or something?
- 1:42:21
Uh, yeah. It's a lot of, um-
- 1:42:24
Is it working?
- 1:42:25
Yeah.
- 1:42:26
Yeah.
- 1:42:26
Yeah. Okay.
- 1:42:27
Okay. So basically, what it's doing is it's now-- You can see the utilization here. It's like a hundred percent only on the Spark. So it's running the prefill there, and it's streaming the KV cache over to the, uh, MacBook in a way that's like overlapping.
- 1:42:40
So-
- 1:42:40
I think I got a Wi-Fi error. [laughs] Oh, well.
- 1:42:45
Okay. Wi-Fi, I guess.
- 1:42:47
Okay.
- 1:42:48
Should we try one more?
- 1:42:48
I need to download the... Okay, let me try that.
- 1:42:53
Okay. Never mind. You get the idea anyway. Uh, so yeah, like in this, in this specific example, I think it's about two X faster, is it?
- 1:43:01
Uh, yes. It's around two X.
- 1:43:03
Yeah. So you get like on the end-to-end time of, you know, doing this whole thing, the prefill and the decode stage, it's about two X faster. Um,
- 1:43:12
the, you know, just running it on the MacBook versus running it on both. So yeah. And then again, you can imagine like the next thing here is like, okay, instead of a Spark, you know, what about just doing this with just a GPU, right?
- 1:43:25
So you could do this with a RTX fifty-ninety, which is cheaper, and it also has a lot more memory bandwidth and a lot more compute. So then you can start doing more interesting things, um, where you split parts of the model.
- 1:43:37
Is it working?
- 1:43:37
That's a single node.
- 1:43:39
Okay. That's, that's single node?
- 1:43:40
Yep.
- 1:43:41
Well, I don't know if we need to...
- 1:43:43
And then save that.
- 1:43:45
How long did that take?
- 1:43:46
Uh, like s- seven seconds for one paper.
- 1:43:50
Okay. Seven seconds for one paper [laughs] with the-- just the MacBook. And then...
- 1:43:57
Let's try it. Let's see if this then works. So, so Spark is on.
- 1:44:08
We're, we're quite over. Uh, just try it one more time. If it doesn't work, it's fine. That's it?
- 1:44:13
Uh, no.
- 1:44:14
Oh.
- 1:44:14
I usually take the disconnect.
- 1:44:23
Okay, so now the same paper. Now it should be running across both. Does it have to do one more? Okay. Yeah. So in that case, four point eight. So again, this gets better as you increase the size of the prompt.
- 1:44:38
Um, oh, cool. Okay. I know we're over, so I'll stop there.
- 1:44:45
Thank you. [clapping] [outro jingle]