AI Engineer Summit 2025
How to Build Your Own AI Data Center in 2025
About this talk
Arista Networks technical lead Paul Gilbert explains how enterprises design dedicated AI data-center infrastructure, contrasting training and inference GPU requirements, isolated backend fabrics, front-end storage networks, and east-west traffic. He discusses NVIDIA H100, DGX, and HGX systems, simplified BGP routing, RoCEv2 congestion management using PFC and ECN, Arista EOS, GPU-to-switch coordination, and emerging NIC-centric network designs.
Chapters
- 0:00Introduction and AI infrastructure requirements
- 2:00Training, inference, isolated GPU fabrics, and routing
- 4:14H100 systems and DGX versus HGX server architectures
- 11:26East-west traffic and RoCEv2 congestion control
- 15:12Arista EOS and GPU-to-switch coordination
- 20:35Large GPU clusters, NIC-centric design, and conclusion
Talk transcript
- 0:00
[on-hold music] My name is Paul Gilbert.
- 0:17
I'm a tech lead for Arista Networks. I have an accent, but I'm actually based here in New York City, and I build or design or help build and design, uh, enterprise networks.
- 0:29
Uh, but what we do is the plumbing, uh, so I'm not gonna talk about agents, but more kinda how you train, uh, models, what the infrastructure looks like, and how you do inferencing on, on the infrastructure.
- 0:42
Uh, I, I, I normally teach people, uh, the very basic stuff, so I, I-- you guys probably know this already, but these are new terms for us when, when we built computer networks.
- 0:53
People will come to us and say, job completion time, uh, barrier. I'm, I'm pretty sure you guys know that the, the inference. And the question I get all the time is, you know, we, we can build a network to train a model.
- 1:05
There's a, an algorithm maybe you can use to, to look at what you, what you need, uh, but then, you know, what's inference? And, you know, it's changed a lot now because of chain-of-thought and reasoning models, the inference is a lot different.
- 1:18
It used to be X, but now it's Y. Uh, I'm, I'm pretty sure you guys have seen this slide, but I use these just to talk to enterprises around kind of what they might be thinking in GPU size.
- 1:30
Uh, this, uh, Walid Sosa came up-- Dr. Walid Sosa came up with this. On the left there is the training, and on the right there is the inference, and it's kind of, you know, on one you have times eight- eighteen times, the other times two.
- 1:44
Again, I think that changes now with chain-of-thought and reasoning. Not too sure kind of which way it's gonna go. And at the bottom there is a really interesting one, again, which I show customers 'cause I-- most of the enterprises I talk to kind of don't understand models and how they work and training.
- 2:00
I, I know a little, but not a lot. But, you know, the, the, the, the, the model they trained down here was two thousand forty-eight GPUs for one to two months, and then when you go to inference after fine-tuning and alignment, it's four H100s for inference.
- 2:15
So we talk to people about building different types of networks, which I'll speak about. But, uh, you know, kind of I always start at the beginning and, you know, this is-- I got this slide and I think it's really interesting.
- 2:26
You know, LLMs were kind of just trying a little bit of inference, but now with the, the next-generation models, it's a lot. So this is what we build or I, I build.
- 2:37
Uh, and these are new terminologies for us from the networking world. Uh, backend network. So this is where you connect GPUs to. When we build these networks, uh, they're completely isolated 'cause GPUs are really, really expensive.
- 2:51
They take a lot of power, uh, and they're really hard to get hold of. So when people build AI networks in the enterprise, uh, we don't connect nothing else to these networks.
- 3:00
The, the bottom part of that, the backend network, there's eight GPUs per port on those servers, and they can be NVIDIA, they can be Supermicro, they can be, be whatever.
- 3:10
They will go into a high-speed switch at the bottom there. Uh, you have a leaf switch and a spine switch, and nothing else attaches to that network. And then on the front-end network is where you get stor-storage from to train the model.
- 3:23
Obviously, you know, it-- the GPUs synchronize, they do something, they calculate, they produce an algorithm, and they call for more data, and that's kind of the cycle. The front-end network is not as intense as the, the backend.
- 3:37
The backend network, depending on the model that you train, uh, they-- the GPUs will actually work at four hundred gigabytes. And for, for us in the enterprise, you know, and I've built some big, big data centers, but I've never seen anything like that.
- 3:52
So this is-- in the networking world, this is a, a, a completely new world to us. And we, we make the networks as simple as possible because, again, these are really expensive and people want to get their money's worth.
- 4:05
They want these running twenty-four by seven. Uh, so we do IB-iBGP or eBGP, just, uh, really simple protocols.
- 4:14
Uh, I'm sure most of you have seen this, but again, I, I kind of teach this. So, uh, the, the-- this was an infrastructure presentation, but, you know, that's kind of the back of a H100, probably the most popular.
- 4:25
It is actually the most popular, uh, uh, AI server out there right now. You can see in the middle there, there's four ports, but those four ports are broken out into two, so there's eight ports.
- 4:36
Those are the GPU ports there. And then kind of over to the, the left there, there's the Ethernet ports. So that's what we connect to. We've never seen anything like this before [chuckles], you know, when you first speak to people about this.
- 4:48
You know, I've seen servers with four hundred gig, you know, and I do a lot of the big financial networks, but never before have we seen servers that can put this type of traffic onto a network.
- 5:01
Uh, you know, we always-- they always ask me about this and, you know, I, I got this from an NVIDIA slide. It's down there, but, you know, there's this thing called scale up and scale out.
- 5:11
I-I'm not really sure. Scale up, you know, when, when you buy-- when act-- my customers buy these servers, they always have eight GPUs in it. You can't add anything to an NVIDIA server.
- 5:20
You get the DGX. The-- If you go with the outsource model, it's a HGX, so it's a third party. You don't really add things to it. But-- So I don't see scale up.
- 5:29
But scale out, you know, obviously, you can build the-- we, we build the network so that you can add more GPUs. We can start very small, and we can go up to the hundreds of thousands of GPUs.
- 5:40
Not in the enterprise, but the cloud scale guys do.
- 5:44
So, so what's different? Uh, you know, for us, again, it's-- there's s- it's hardware and software. The hardware are those GPUs. We're not used to them. Uh, the first time I tried to configure one, it took me hours and hours, but I'd never seen them before.
- 5:59
Whereas Other stuff I've seen pretty quickly. You know, and you have software, so CUDA and NCCL are probably, you know, two of the biggest protocols, and you, you guys know more about that than me.
- 6:09
But we had to kind of understand, not CUDA, but NCCL because it has a collective, so we had to understand kind of how the collective works because that will put, uh, traffic onto the network in a certain way.
- 6:20
Uh, the hardware was completely different. Again, you know, we had the eight, uh, 400 gig ports and the four 400 gig ports facing the, the, the front-end network. Totally new to us.
- 6:32
The other thing was data center applications, kind of web app database. They're, they're really easy. Uh, they go from one to the other and in different parts of the network, and if one fails, you, you have some kind of load balancing or fail, and it fails over.
- 6:46
Uh, this is-- AI networks are not like that. The GPUs all speak. They all talk to each other. They all get stuff. They all send stuff. And if one fails, the, the job might fail, it might recover.
- 6:58
But it's a different concept to us, so it's, it's hard to imagine. Uh, and traffic is bursty because all of these GPUs, if you have a thousand GPUs at 400 gig, they will all burst at the same time, and if you-- if they can, they will burst at 400 gig.
- 7:14
So [coughs] there's a lot of traffic on a network, and I'd never seen anything like that. So when we build these networks, we don't build them oversubscribed. We build them one-to-one.
- 7:23
Uh, might-- In the data center world, we used to do one-to-ten. It went down to probably one-to-three, but never one-to-one because it's just really expensive to, to, to build that kind of bandwidth.
- 7:35
But with AI networks, we need to, so we have no oversubscription in the network. And from, from our point of view, if you look at what one of these servers can put on the network, you know, just a, a H100 is eight 400 gig GPUs and four 400 gig is four point eight terabytes, which is...
- 7:53
And that's just one server. The storage size, the n- the, the, the front end probably nowhere near that, but the back end is always wire rate. And then 800 gig is probably, you know, the, the Bs are just around the corner.
- 8:06
I think in March they'll be released. I think there's some people that have them. And those are 800 gig. We support 800 gig today on the network, but each one of those servers in is a possible nine point six terabytes per server.
- 8:19
And you-- most people in my world, in the, in the enterprise world, come from servers that may be one, two, three, or four a hundred gig Ethernet, but nothing like, uh, nine point six terabytes per server.
- 8:35
So the other problem we have is the traffic patterns. When we load balance from kind of leaf to spine, we use a thing called n-tuple, which is the five tuple IP address pool and MAC address, and we do pretty good load balancing.
- 8:49
But with GPUs, it's just one IP address, and it can sometimes match to a single uplink and oversubscribe it, which would be really bad because you'll start dropping an awful lot of packets.
- 9:01
So we have to take a lot of care on how we load balance within the AI network or how we build the back end and the front end. So we have some pretty cool tools where we don't now look at the five tuples.
- 9:14
We actually load balance on the percent of bandwidth that's being used on the uplink. And we can get up to about ninety-three percent utilization on all the uplinks to the downlinks, which is pretty good.
- 9:29
Uh, you know, and if, again, one thing that's really new to us is a single GPU can, you know, or a set of GPUs, if they fail, uh, sometimes the model will stop.
- 9:38
There are no checkpoints, but, uh, a single f-- GPU failure is a problem for us. And if-- one of the big problems is-- we've always had is optics and transceivers and Doms, which is the, the rates and the, the loss between them and the cables, et cetera.
- 9:53
And when you start building these networks with thousands of GPUs, you will have a lot of cable problems, and you will have a lot of GPU problems. So it's, it's really hard for us because we-- again, this world is, is new to us for the last year or so.
- 10:09
Uh, power, I, [laughs] you know, power is-- You c- you know, y-you've read the newspapers. You know, everyone's trying to buy, buy nuclear power stations to power these things.
- 10:18
The, the average rack in a data center today is about seven kW to fifteen kW, and you can put like, you know, ten, one, one IU racks into those, and you'll be fine. [coughs]
- 10:30
And when customers come to me say, "Yeah, we finally got GPUs," whatever, and I say to them, you know, "What kind of racks have you got?" [laughs] And they say, "Well, we're gonna put them in..."
- 10:38
And then you can only put one of these servers in one of those racks because they, they actually draw with eight GPUs, ten point two kW, so you need new racks.
- 10:48
Uh, most enterprises now are waking up to this, and they're building racks between a hundred and two hundred kW, and they're water-cooled. There's no way you could air-cool them in a data center.
- 10:58
So that's a whole new concept to people as well, is water-cooled racks.
- 11:04
Uh, traffic is, is both ways, which again is new to us. So north-south, you know, in a regular data center, you have users coming in, database, app, web, whatever, and it comes in, it goes out.
- 11:15
But i-i- in the AI world, when the GPUs speak, that traffic is east-west because it's speaking amongst each other. Uh, and then when they ask for more data from the storage network, it's north-south.
- 11:26
So you have both traffic patterns. The east-west is really bad. That's kind of where they run wire rate. The front end to the storage is much m- more calmer because most storage vendors can't put that kind of traffic on, on the, on the network right now.
- 11:42
I'm pretty sure they will one day, but they're more around a hundred, two hundred gig.
- 11:47
Uh, and, you know, in a network, there's a certain amount of buffering on these switches, and buffering is bad because it means it can't send traffic somewhere because something else is, is not receiving the traffic.
- 12:01
So you need a congestion control and feedback. Uh, and right now we use something called RoCE v2, which is two parts of RoCE v2. There's a PFC and an ECN.
- 12:11
Uh, if you were building an AI network, your engineers, your network engineers will definitely know about this. ECN is an end-to-end flow control, where if traffic, if, if traffic is-- if there's congestion somewhere in network, packets are marked, they go to the receiver.
- 12:28
The receiver sends back to the sender, "You need to slow down because there's congestion," and it goes through an algorithm. It pauses for a while, it slows down, and if it doesn't get any more ECN packets, it speeds up again.
- 12:40
Uh, and PFC is basically stop. Uh, my buffers are full, I can't take any more, so it kind of is a dead stop. So you have kind of a, a slow, uh, feedback mechanism with ECN and a kind of emergency stop with PFC.
- 12:58
The networks we build are really simple. We don't, uh, have things like... In regular data centers, we have DMZs with firewalls, load balancers, et cetera. We have connections to the internet.
- 13:08
We have L4 through seven service, a whole bunch of stuff. When we build these networks, they're totally isolated. Uh, the GPU, the, the back end is completely isolated. The front end possibly could have connections to something, but even then it's so expensive to build you, you don't wanna take the chance.
- 13:28
Uh, on demand, the applications that we're used to, you know, if it fails or something fails, something will recover and, you know, you, you may get a, a little skip or a jump, but if you've done the right thing, it's not gonna be that bad.
- 13:41
In this world, if something fails, the model may fail, and you-- the call that you get into the operations center is different call than you get that if your, your at-- kind of restarted and everything's good again.
- 13:54
Uh, the other thing is collectives. You know, obviously Nico will go out there and work out kind of where the GPU, GPUs are and what to do with it, but there's kind of different designs.
- 14:04
So I tell my customers that, "Speak to your data scientists and your, your, uh, your programmers, developers, and find out kind of what they're doing and what kind of models they're building," because it can, can affect the network on kind of how you build it and how you design it.
- 14:22
So network's totally isolated. Things are moving fast. We're at eight hundred gig right now, uh, uh, which is, you know, we have been for probably a year. We will see one point six terabytes on the network probably end of this year, uh, early twenty twenty-seven, and it will just keep growing and growing and growing, and these models
- 14:42
will get bigger and bigger and consume more and more and more, I'm pretty sure.
- 14:47
Uh, visibility and telemetry. I-- You know, all my customers, the, the call that they get when a model fails because the network is the problem is a different call than they're used to.
- 14:58
So we put different, uh, telemetry and visibility in there to make sure that if things are going wrong on the network that, you know, they know about it, hopefully before they get that call.
- 15:12
So yeah, I work for Arista. Our operating system is called EOS, and we have a whole bunch of features there. So if you were building an AI network, I'm not sure that you guys speak to the engineers, but this is the type of things we talk about.
- 15:26
Uh, lossless Ethernet. Everyone's thinks that you, you know, when you train the model, you can't drop packets. I, I've seen it and you, you can. I think dropped packets are okay.
- 15:35
Consistent latency is okay, but if you drop so many packets, obviously it's a problem. So flow control lossless Ethernet is really key. ECN and PFC are part of that.
- 15:45
As I said before, they're flow control mechanisms. One is a slow down, please, and the other one is a stop. And as you know, because GP- GPUs are synchronized, if something slows down, you slow down one port, one GPU, everything slows down.
- 15:58
So you really gotta be on top of kind of the oversubscription, and if you are getting queue in, where is it? Uh, we, we have really good buffer. Uh, we can adjust buffers.
- 16:09
We have different kinds of switches for different places in the network. But we found that models send and receive a particular size packet, and what we do is we adjust those buffers to accept those types of packets.
- 16:22
Buffering is a really expensive commodity in switches in networking, and if you can find a way to allocate the buffers exactly tuned to the packet sizes, it's a win-win, uh, and we, we've worked out how to do that, which is good.
- 16:37
Uh, yeah, monitoring is really key for us. Uh, I tell my customers there's probably five things you wanna do. One of them is RDMA. You know, uh, these networks train using RDMA, you know, uh, which is memory to memory writes rather than going CPU to memory.
- 16:54
And RDMA is a, a complex protocol, and it has ten or ten or twelve, maybe more kind of error codes. So if the network starts seeing problems, uh, and starts dropping packets, rather than just drop the packet on the floor, we can actually ca- copy that packet to a buffer or send it somewhere or just, just the
- 17:14
headers and, and why we dropped that packet. And if you think about it, it's really cool. Like most networks will, in congestion, your buffer fills up is you're gonna drop the packets.
- 17:24
We'll drop the packet, but we'll actually take a, a, a snapshot of the packet and the headers and the RMA information in it, and we'll tell you why we dropped it.
- 17:35
Uh, another thing we have, which is really cool, we have AI agent. Uh, you know, from the networking point of view, we can look at what's going on, but we don't really have any visibility into the GPU.
- 17:45
So now we have an agent, which is an API, uh, and some code that we load on the, the GPUs in NVIDIA, and they will speak to the switch.
- 17:54
So that agent will say to the switch how are you configured. So PFC and ECN is th-those flow control mechanisms- Have to be configured correctly, 'cause if they're not, it will be a disaster.
- 18:06
So the, the GPU will speak to the switch and say, "This is how I'm configured." The switch will say, "Yeah, you're good. We understand each other." And the second thing it does, it gives you a whole bunch of statistics about packets received, packets sent, RDMA errors, RDMA issues in there.
- 18:21
So you can, can correlate now if the problem is the GPU or if it's the network, which is a huge step forward for, for us.
- 18:30
Uh, another really cool feature we have is, uh, smart system upgrade. You know, if you're used to routers and switches, you know, you have to upgrade the software sometimes.
- 18:40
Uh, sometimes to get new features, sometimes to fix PSIRTs, which are security vulnerabilities on that switch. Uh, we've worked out a way now that we can do that. You can upgrade code without actually taking the switch offline.
- 18:53
So, you know, if you have 1,024 GPUs with 64 switches in your network, you actually can upgrade those and the GPUs can keep working. So it's a real big step forward for us.
- 19:07
So, you know, for us, [laughs] again, I, I don't know if... But no oversubscription on the backend. You can't because the GPUs use everything you give them. Address-wise is really important for us.
- 19:18
It's a, it's a point-to-point connection, so it's /30s, /31s. You could use IPv6 if you have IPv4 problems, address space problems. All my customers I tell BGP because it's the best protocol out there.
- 19:32
It's really simple and it's really quick. Uh, EVPN VXLAN if you have multi-tenancy, if you have a lot of different business units, lines of business, uh, using the network.
- 19:43
You need things like advanced load balancing. We have a couple of different... We, we actually look at the collective that you're running, can load balance on that collective now, which we call cluster load balancing.
- 19:54
You could deploy it rocky. I tell all my customers do it because if you don't, [coughs] your network's gonna melt down, you're not gonna know why. These things will give you an early warning system that you need to do something with your network, so they're, they're really key to have.
- 20:08
And visibility and telemetry is, is really good at all times because in the network NOC or the operation center, you always want to be aware of the problem before you get the call from the developers and the, the people that have paid a lot of money for that network.
- 20:23
I'm running out of time here, but this is kind of 1,400 gig cluster, what it would look like spine and leaf. Uh, again, no oversubscription. 800 gig links between the leaf and spine.
- 20:35
400 gig down to the GPUs. This is a 4,000 cluster.
- 20:41
These ones, th- these are the bigger boxes. These are 16 slot. One of these boxes can take five hundred and seventy-six, 800 gig GPUs, so eleven hundred and fifty to 400 gig GPUs.
- 20:54
So if you're building clusters with thousands of GPUs, then this would be the box for you, the 7800 series.
- 21:02
And, and putting it all together, this is kind of what we would build. There's three networks here. There's a backend network where your GPUs live. There's a front-end network where the storage live.
- 21:11
And then there's these, the, the inference that you take the model, you put it somewhere else.
- 21:17
And I'm out, I'm out of time. Oh, they're not coming, so I'm guess I'm not... So, [laughs]
- 21:22
so, so the other thing is, you know, there's Ultra Ethernet Consortium. You know, I don't know if this interests you. Ethernet hasn't changed the way it's built probably thirty years.
- 21:31
Uh, there's some things it could do better around congestion control, around packet spraying, around the NICs talking to each other. So there's this thing called Ultra Ethernet Consortium. Version 1.0 will be ratified probably Q1 2025.
- 21:45
And it's a kind of different way of building networks, and you probably won't see them till Q3, Q4. But most of the cloud scale guys were kind of really keen on this because it, it puts a lot more into the NICs and takes a lot more out, out of the network.
- 21:59
So we just get... We can do what we're good at, which is forwarding packets.
- 22:05
So summary, you know, for us, we have the front end, which is storage, the back end, which is the really important part for us. Uh, th- that part i- is really bursty.
- 22:15
Uh, the GPUs are all synced, so they send and receive at the same time, and if you have a slow GPU, that's a barrier because it stops everyone else.
- 22:23
Job completion time is what matters to us. If, you know, we get the call that, you know, my job completion time was one hour yesterday, it's four days today, you know, it's probably our problem.
- 22:34
Uh, you know, models can checkpoint, but they're really expensive. You guys know that. [outro music]
- 22:39
And I'm done. [audience applauding] [outro music]