← All AI Engineer talks

AI Engineer World's Fair 2025

Why We Don’t Need More Data Centers

About this talk

Hyperbolic co-founder and CEO Jasper Zhang argues that building additional data centers alone cannot meet growing AI compute demand because construction is costly, energy-intensive, and constrained while existing GPUs remain underused across fragmented providers. He proposes a GPU marketplace and distributed orchestration layer that aggregates available capacity, supports flexible rentals, and lowers acquisition costs. He describes Hyperbolic’s Hyper-dOS software, Kubernetes and MicroK8s integration, H100 pricing examples, and a Monarch-server architecture that coordinates provider provisioning and SSH access. The recording concludes with a question from an unidentified additional participant.

Chapters

  1. 0:00Introduction: Jasper Zhang, Hyperbolic, and the data-center thesis
  2. 1:33Growing demand, energy constraints, and idle GPU capacity
  3. 4:28GPU marketplace architecture and Hyper-dOS orchestration
  4. 6:19H100 pricing, supplier selection, and startup productivity
  5. 12:52Audience question: Kubernetes, MicroK8s, Monarch, and SSH provisioning

Talk transcript

  1. 0:00

    [on hold music] Nice meeting you guys.

  2. 0:16

    Uh, great to be here. And, uh, I'm here to present Hyperbolic, which is a AI cloud for developers. And so my topic is, uh, why

  3. 0:26

    we don't need more data centers. It's like a very [chuckles] eye-catching title. Uh, but what I want to clarify is I still think building, building data centers is important, but just building data centers alone can't solve the problem.

  4. 0:41

    So, uh, wait. Before we get started, uh, let me introduce myself. I'm Jasper. I'm the CEO and co-founder of Hyperbolic. Um, I did my math PhD at UC Berkeley.

  5. 0:52

    Uh, finished my PhD in two years, which make me the fastest person in the history of Berk- Berkeley. And then I also won a few gold medals. So, uh, after that, I work at Citadel Securities, uh, trying to use AI and machine learning to predict the market and execute strategy.

  6. 1:06

    So I always have a passion about how to make things very efficient and how to help you to save money because everyone knows that, uh, compute actually one of the biggest costs for your companies or for your startups.

  7. 1:18

    Uh, usually you need to s- if you want to run like one thousand GPU will spend you millions of dollars, uh, per year. And we, we think that these problems should be solved by not just building more data centers, but actually, uh, building a GPU marketplace.

  8. 1:33

    So let's get started, uh, with the problem that we're facing. Uh, first, uh, I think-- So, so everyone knows that AI is gonna integrate with everything in the future, and every companies will be AI companies.

  9. 1:46

    So the demand for GPUs as well as data centers are exploding. So by McKinsey, by twenty-thirty, we'll need four X more data centers built in one quarter of the time that we build in this speed.

  10. 2:03

    Uh, but what if I tell you that you actually don't need that many data centers? Uh, you actually need, uh, another solution. So, uh, we can break down the demand first.

  11. 2:15

    Uh, right now, uh, the current capacity for data center is fifty-five gigawatts. Um, w- by the median, uh, scenario, we're going to see twenty-two percent annual growth rate for the demand.

  12. 2:29

    So in twenty-thirty, we're gonna need two hundred and ninety gigawatts.

  13. 2:38

    And however, uh, it's like there are a lot of challenges building data centers, right? So first, uh, we everyone knows Stargate, so it takes like, uh, for the first Stargate data center, it takes like more than a billion dollars to build.

  14. 2:52

    Uh, and then also it's very slow to colla- connect data center to the electrical grid. For example, right now the, the wait- wait list is like seven years. So you need to wait seven years to connect a one hundred megawatts facility to the, uh, to the electric- electrical grid in Nor- uh, Northern Virginia.

  15. 3:11

    And, uh, and then, uh, it also very, uh, consuming a lot of energy. So, uh, currently, we're spending four percent of the total electricity consumption in the US for just GPUs and data centers.

  16. 3:24

    Uh, and also is not very environmental sus- sustainable. Uh, if you can look at the number, that's crazy, uh, CO2 emissions annually.

  17. 3:34

    Uh, and even say if we're gonna deliver all the data centers, uh, on time, there's still a data center supply deficit of more than fifteen gigawatts in the US alone by twenty-thirty.

  18. 3:48

    And so it means that just building data center can't solve the problem. Uh, on the other hand, uh, we think the GPU utilization is actually pretty low. So according to, uh, Deloitte, GPUs sit idle eighty percent of the time for enterprises and companies.

  19. 4:11

    According to SemiAnalysis, there exists a hundred plus GPU clouds. So we can see like how fragmented this space is, right? A lot of you guys need GPUs, but you can't find them or like you are going to pay extremely high price.

  20. 4:28

    On the other hand, there are a lot of GPUs sit idle in data centers or in different clouds. And so naturally, uh, a solution that we think we, we should build is actually build a GPU marketplace or like aggregation layer that aggregate different data centers and GPU providers to solve the problem for, uh, GPU users.

  21. 4:50

    Uh, it doesn't necessarily need to be Hyperbolic, but I just use Hyperbolic as example, uh, to show here. [chuckles]

  22. 4:57

    Uh, so, uh, I can, I can just like, uh, share what we are s- we're trying to solve. So we're building this like global orchestration layer. Uh, we invented a software called HyperDOS, which is short for Hyperbolic Distributed, uh, Operating System.

  23. 5:13

    So basically, it's like a Kubernetes, uh, software. So any, any cluster, as long as it installed our software within five minutes, suddenly the data center become a cluster in our network.

  24. 5:27

    And on the other side, users can rent GPUs, uh, in different ways that they want. Like, they can just, uh, do the spot instance, they can like on-demand, they can long-term reserve, or they can also like host, uh, models on top.

  25. 5:43

    Um, and so like we see that-- we see there are like several benefits. Um, one, we, uh, we k- kind of like solve the effi-, uh, like the matching problem of compute.

  26. 5:55

    Uh, and then second, like GPU become commodities, so you like, you don't need to spend too much time to wait for data center. You just buy them on the marketplace.

  27. 6:04

    And then third, uh, you can have different options.

  28. 6:09

    And so, um, we do some math modeling. Uh, I, I mean, I don't have time to kind of put down the math in the slides, but this is our conclusion, right?

  29. 6:19

    Basically, uh, we can save the cost by fifty to seventy-five percent. Uh, even if you look at, uh, the current... We, we're running like some beta version of our marketplace right now, and our GPU cost for H100 is ninety-nine cents per hour.

  30. 6:36

    But if you look at Google, for example, they have on-demand GPU, it's like eleven dollars. They have like Lambda, they have like two or three dollars. But on average, by have ga- aggregating more supply, uh, and then like have a uniform distribution channel, you can drama- uh, drastically reduce the price.

  31. 6:56

    Um, it's... Like the, the theory behind that is like the Queuing theory. Basically like, uh, it's M/M/c theory. I probably... Next time if we're gonna watch my talk, I will share more math, uh, behind that.

  32. 7:08

    Uh, but yeah, and then like you can just save time to vetting your suppliers because y- if you like think about... I, I mean, how many people here are founders or like need to acquire GPUs?

  33. 7:21

    Yeah. So, uh, are you frustrated when you are trying to talk to... How many suppliers are you talking to? If you have talked to more than five, raise your hands.

  34. 7:32

    Are you frustrated when you're like trying to have like five sales calls and like try to like know which data's, uh, GPUs are, are frustrated? Yeah. Uh, are good?

  35. 7:42

    Yeah. That's good. Yeah. So basically, by having like this uniform platform, like founders or like startups or companies no longer need to vet different data center. They just like pick the one that they, uh, they have high rating or like have the best price.

  36. 7:59

    And we're also going to do like, uh, benchmarking on the performance of the GPUs.

  37. 8:06

    All right. So, uh... Oh, sorry. All right. So, uh, sorry, somehow the graph didn't, didn't show.

  38. 8:21

    Uh, let's... Give me one sec. Yeah. So, um, basically, we can think about a use case example. Um, so let's say if you, if you are a startup and you want like one thousand GPUs at the beginning, so usually you will just reserve these one thousand GPUs for a year, right?

  39. 8:43

    You think like, "I might need to use these GPUs, uh, for training, and later I want to do inference." And so you run some training jobs, and then after three months, then you realize that, okay, now I have a s- I have a ni- good-- a better idea by running those experiments, and now I need one thousand

  40. 9:02

    more GPUs just for a month, right? And then after, after six months-- at month six, then you finish your training job. And then you realize that now I only need five hundred GPUs for hosting my model, but w- I still have five hundred GPU left.

  41. 9:19

    So, uh, on the traditional, uh, on Hyperbolic case, uh, you basically can say, "Okay, I will rent one thousand GPUs for a year at the beginning." But then, uh, in month three, I can say, uh, "I just rent, uh, an, an extra one thousand GPUs for just, uh, a month."

  42. 9:40

    And then, uh, a month-- in month six, then I can say, "Okay, I can re-list my idle GPUs on Hyperbolic and try to sell the, uh, sell them to the, uh, to other people that need them," right?

  43. 9:52

    Uh, but if you just like use some traditional cloud, then you need to rent one thousand GPUs at the beginning, and then on month, in month three, you need to rent actually one thousand GPUs for a year usually.

  44. 10:03

    And, uh, if you calculate the cost, uh, compare, compare that and then also like think about the price difference you will have, uh, it will... You can reduce the cost from forty-three point eight million to six point nine million.

  45. 10:19

    So it's like six X saving. Uh, and you also help other people to get cheaper GPUs too because you can re-list those idle GPU to other people. And so, uh, so that's-- this is how we think that, uh, we're gonna, we're gonna like increase the productivity.

  46. 10:35

    Like, people only think about saving, but actually, uh, this is not true for GPU, right? Uh, by scaling law, we know that the more g- compute you spend, the better quality your machine will be-- uh, your model will be.

  47. 10:49

    So it's not just about saving your cost by six X, it's more about with the same budget, you will increase your b- productivity by six X. And imagine how many startups that they used only need to rely on OpenAI and Anthropic, those closed AI models, but now suddenly they- their money become more valuable, and they can

  48. 11:14

    rent as many GPUs as they want for the training.

  49. 11:19

    Um, and so the, the next step that we think, uh, usually the GPU marketplace will evolve into is that, uh, it will be a all-in-one platform for different AI workload.

  50. 11:30

    Because what people really want is not just GPUs. They want, um, to run their different AI jobs, right? They will... You will have AI inference, uh, online inference, uh, offline inference, and then you will also have a training job.

  51. 11:47

    And so, uh, yeah, so this is like, um, to, to like, uh, some takeaway. Like basically, we don't think we n- we need like just focus on building data centers.

  52. 11:59

    We also need to do like smarter allocation for the resources. And then second, uh, we can reduce your costs, um, for by building GPU DS-- uh, marketplace. And lastly, um, I think Uh, just focusing on building data center is not very sustainable.

  53. 12:15

    We're c- causing a lot of energy, uh, taking a lot of land. Uh, we should better reuse, recycle [chuckles] those idle compute by, uh, se- sending it to others. And so, uh, if you're interested in trying out, uh, you can, uh, come to our website.

  54. 12:33

    Uh, the, the left QR code is, uh, the current product that we have, which is a marketplace. But then we're also launching our business card and enterprise card that, uh, give you, like, production-ready GPUs with 99.5% reliability.

  55. 12:47

    All right. Thanks. [audience applauding]

  56. 12:52

    Awesome. So I actually got... I'm curious, can you tell us more about the, the kinda Hyperbolic OS? How exactly does that turn... 'Cause I know a lot of times you have a data center plus a set of GPUs.

  57. 13:05

    Yeah.

  58. 13:05

    How, how does it actually work to connect it to Hyperbolic itself?

  59. 13:10

    Yeah. So, um, basically this is a Hyper- HyperDOS is like a Kubernetes agent. So, um, you just install that in your cluster as long as you have Kubernetes. I mean, um, most data center have Kubernetes, but then even for your MacBook or for your, uh, PC, you can just install like MicroK8s to kind of, uh, become a

  60. 13:31

    Kubernetes-ready, uh, machine. And, uh, so basically now you kind of have... We, we have terminology in-house. We call, like, our Hyperbolic server, uh, Monarch, and then we have, uh, different barons. [chuckles]

  61. 13:48

    So it's like a feudalism model. So different barons, they own different compute. And then anytime... And every, every time when a user want to rent GPU, they will talk to our Monarch server, and the Monarch server will send a request to, uh, the, like, the baron, and then baron will just basically, uh, provis- provision the machines and

  62. 14:07

    set up a SSH instance for customers to access. Yeah. [outro music]