← All AI Engineer talks

AI Engineer World's Fair 2025

Why We Don’t Need More Data Centers

Read the talk

Why More Data Centers Aren’t Enough

Idle GPUs and fragmented suppliers make compute scarcity an allocation problem as well as a construction problem. A GPU marketplace connects those two sides.

From a talk by Dr. Jasper Zhang

Before you start: Familiarity with GPU training, inference, and cloud reservations is helpful; Kubernetes appears only in the final provisioning explanation.

Compute is expensive before it is scarce

If a startup needs a thousand GPUs, will building more data centers solve its compute problem? More capacity helps, but it does not automatically make existing GPUs affordable or available to the people who need them. Construction must be accompanied by better allocation. That is the qualification behind Jasper Zhang’s provocative title: he considers new data centers important, but insufficient on their own.

Zhang introduces himself as Hyperbolic’s CEO and co-founder, with a background in mathematical and financial efficiency. He says he completed his Berkeley mathematics PhD in two years—a university speed record in his account—and won several gold medals before working at Citadel Securities on AI and machine learning for market prediction and strategy execution. The practical concern carrying into Hyperbolic is cost: Zhang frames a 1,000-GPU operation as a multimillion-dollar annual expense. His proposed complement to construction is a GPU marketplace.

0:160:26
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:16 · section reference included

Demand grows faster than construction

As AI spreads into more products and companies, demand for GPUs also becomes demand for buildings, power, and supporting infrastructure. Zhang attributes a stark forecast to McKinsey: by 2030, four times more data centers will be needed, built in one quarter of the time.

Dark slide reading “By 2030 we'll need 4X more data centers built in 1/4 the time,” with the speaker inset below.
The opening demand claim: four times more data centers by 2030, built in one quarter of the time.

He then expresses demand in electrical capacity: 55 GW currently, 22% annual growth in a median scenario, and 290 GW needed by 2030. These are the figures presented in the talk, rather than a consistent restatement of McKinsey’s published scenarios. The relevant McKinsey analysis measures demand through power consumption, not building count, and separates its median and accelerated growth cases. The distinction matters: a multiple of capacity is not necessarily the same multiple of buildings.

1:331:46
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

1:33 · section reference included

The constraints extend beyond buying GPUs

A data center must pass several bottlenecks before its GPUs can serve a workload. Zhang’s examples move from capital to power delivery and then operating impact:

  • Construction cost: he puts the first Stargate data center at more than $1 billion.
  • Grid connection: he cites a seven-year wait to connect a 100 MW facility in Northern Virginia.
  • Electricity consumption: he attributes 4% of total US electricity consumption to GPUs and data centers.
  • Emissions: he also points to annual CO₂ emissions as a sustainability concern.

These are the talk’s infrastructure examples, not a site-specific construction or interconnection estimate. They explain why adding hardware and delivering usable compute are different tasks.

Slide listing $1 billion for a single StarGate data center, seven years to connect a 100 MW facility in Northern Virginia, 4% of U.S. electricity consumption, and 105 million tons of annual CO₂ emissions.
“The Infra Trap” lists construction cost, grid connection delays, electricity consumption, and annual CO₂ emissions.

Even timely delivery of announced projects would not necessarily close the gap. Zhang cites a conditional US data-center supply deficit exceeding 15 GW by 2030, assuming the planned facilities arrive on time. That leaves a reason to improve the use of existing capacity while new capacity is being built.

2:382:52
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

2:38 · section reference included

Connect idle capacity to buyers

The other side of scarcity is underutilization. Zhang attributes an estimate of 80% idle time for enterprise GPUs to Deloitte; the population and measurement window behind that estimate are not specified. He also cites more than 100 GPU clouds, a fragmented market described in SemiAnalysis’s GPU Cloud ClusterMAX analysis. A large supplier count does not mean every provider is ready to serve every workload. Buyers can struggle to locate affordable GPUs while capacity sits unused elsewhere.

A marketplace addresses that mismatch by aggregating data centers and GPU providers behind a shared access point. Zhang uses Hyperbolic as the example, but the allocation mechanism does not depend on that particular company: make dispersed supply discoverable and rentable by people outside the provider’s existing customer base.

Hyperbolic calls its orchestration software HyperDOS, short for Hyperbolic Distributed Operating System. Zhang initially describes it as Kubernetes-like software that brings a provider’s cluster into a global network. He says a cluster can join the network within five minutes of installing HyperDOS. The installation is the proposed bridge between a provider owning compute and that compute becoming accessible through the marketplace.

On the buyer side, Zhang describes several ways to consume that capacity:

Access optionBuyer’s choice
Spot instanceUse spot capacity
On-demand rentalRent as needed
Long-term reservationCommit capacity in advance
Model hostingRun a model on the platform

The intended benefits are better matching between supply and demand, commodity-style purchasing instead of waiting for a new data center, and a wider choice of suppliers and rental arrangements.

3:484:11
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

3:48 · section reference included

Pooling supply and simplifying procurement

Zhang reports modeled cost reductions of 50–75%, without presenting the derivation. He then gives the following price comparisons from the period of the talk:

Offering cited by ZhangQuoted GPU price per hour
Hyperbolic beta marketplace H100$0.99
Google on-demand GPUAbout $11
Lambda GPUAbout $2–3

The comparison does not specify equivalent configurations, regions, included resources, or contract terms. It illustrates the price differences Zhang is pointing to, rather than establishing a like-for-like procurement benchmark. His economic explanation is that aggregating more supply through a uniform distribution channel can lower prices.

He names M/M/c queueing theory as the theoretical basis for the pooling argument, but defers the mathematics. The useful intuition is that a shared route to multiple suppliers gives demand more opportunities to find available capacity. The talk does not provide the arrival rates, service rates, or other assumptions needed to derive its savings estimate from that model.

Pooling can also reduce procurement work. Zhang asks founders to recognize the experience of contacting more than five suppliers and sitting through sales calls to learn which facilities have suitable GPUs. In the marketplace he describes, users would select providers by ratings and price instead. GPU performance benchmarking is a planned addition, intended to make the comparison more informative than price alone.

“Your Savings” slide showing 2–4× productivity growth, 50–75% reduction in cost, and zero hours vetting suppliers.
Hyperbolic’s savings slide claims 2–4× productivity growth and 50–75% lower costs.
6:096:19
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

6:09 · section reference included

A reservation outlasts the training job

Consider Zhang’s startup example. The company initially reserves 1,000 GPUs for a year, expecting to use them for training and later inference. Three months of experiments produce a better idea, and the team needs another 1,000 GPUs—but only for one month. By month six, training is finished. Hosting the resulting model needs just 500 GPUs, leaving 500 GPUs from the original reservation idle. The workload has changed twice, while the original annual commitment remains.

The marketplace path keeps the original reservation but changes how the startup handles the mismatch:

  1. Reserve the baseline: rent 1,000 GPUs for the year.
  2. Rent the temporary increment: at month three, add 1,000 GPUs for one month.
  3. Offer unused capacity to others: at month six, relist the 500 idle GPUs from the original reservation.

Relisting proposes a new use for capacity already committed. It does not mean the reservation disappears, or that another customer has necessarily rented it.

Relist idle capacity without changing the original reservation

Constructed example: The record labels and before/after presentation are teaching details. GPU quantities, reservation duration, month-six hosting requirement and proposed relisting come from the talk.

Month-six request — unchanged
Relist the 500 idle GPUs from the startup's original 1,000-GPU annual reservation.

Operation: Offer the idle half of the reservation on the marketplace while retaining 500 GPUs for model hosting.

Original reservation

Before: Before relisting
1,000 GPUs reserved for one year
After: After offering capacity · Unchanged
1,000 GPUs reserved for one year

Model hosting

Before: Before relisting
500 GPUs used by the startup
After: After offering capacity · Unchanged
500 GPUs used by the startup

Remaining capacity

Before: Before relisting
500 idle GPUs, not offered to other buyers
After: After offering capacity · Changed
500 idle GPUs offered to other buyers; rental pending

Marketplace listing

Before: Before relisting
Not present
After: After offering capacity · Added
500 GPUs available from the existing reservation
Relisting exposes unused capacity to another buyer; it does not cancel the reservation or guarantee a sale.

Zhang contrasts this with a traditional-cloud scenario in which the extra 1,000 GPUs also require a year-long reservation. Combining rental flexibility, relisting, and different GPU prices, he quotes a reduction from $43.8 million to $6.9 million, which he describes as roughly sixfold savings. The talk does not supply enough contract, pricing, or resale assumptions to reproduce those totals. The mechanism is nevertheless clear: avoid a long commitment for a short experiment, then try to recover value from reserved capacity that the startup no longer needs. A successful relisting also makes that capacity available to another buyer.

8:218:43
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

8:21 · section reference included

Spend the same budget on more experiments

Lower GPU costs can change what a startup attempts, not just its bill. Zhang invokes scaling laws to connect additional compute with improved model quality. He then extrapolates roughly sixfold savings into sixfold productivity at a fixed budget. That is his interpretation of the purchasing-power change, not a measured training or model-quality result: more affordable compute does not itself establish a proportional improvement in useful output. His example beneficiaries are startups relying on closed models from OpenAI and Anthropic that could now afford more of their own training.

The eventual product is a place to run AI jobs, not merely a place to rent GPUs. Zhang expects the marketplace to evolve into an all-in-one workload platform covering online inference, offline inference, and training. That progression follows the buyer’s goal: a GPU reservation is an input to getting a job done.

10:3510:49
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

10:35 · section reference included

Reuse existing compute while expanding capacity

The resource argument returns to the physical costs of expansion. Data centers consume energy and occupy land; redistributing idle compute offers another way to serve demand. Zhang’s closing emphasis is smarter allocation alongside construction, with the marketplace reducing the friction of sending unused capacity to someone who can use it.

At the time of the talk, the left QR code on the closing slide points to the existing marketplace, while Business Cloud and Enterprise Cloud are presented as launching products. Zhang announces production-ready GPUs with 99.5% reliability, without defining a measurement period or service-level terms. These are the historical product names and promise in the presentation, rather than a statement of today’s offerings.

“Go Hyperbolic & Optimize Your GPU Usage” slide with app.hyperbolic.xyz and two QR codes labeled for 99-cent H100 rentals and early access to Business Cloud and Enterprise Cloud.
Closing slide promotes H100 rentals and early access to Business Cloud and Enterprise Cloud.
11:4711:59
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

11:47 · section reference included

How a rental reaches a provider’s cluster

The Q&A makes the orchestration proposal more concrete: how does a data center with GPUs actually connect to Hyperbolic? Zhang clarifies that HyperDOS is a Kubernetes agent installed in a provider’s cluster. For a machine without Kubernetes, he offers MicroK8s as a way to make a MacBook or PC Kubernetes-ready. Kubernetes readiness is the point of that example; it does not by itself establish that a particular laptop GPU is eligible to supply the marketplace.

Zhang describes the internal naming through a feudal analogy: the central Hyperbolic server is Monarch, and the compute-owning providers are barons. A rental follows this sequence:

  1. The user sends a GPU rental request to Monarch.
  2. Monarch sends a request to the appropriate baron.
  3. The baron provisions the machines.
  4. The baron sets up an SSH instance for customer access.

The shared marketplace handles the customer’s request centrally, while provisioning happens through the provider that owns the compute. SSH access is where the allocation proposal becomes a machine the customer can actually use.

12:5213:05
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

12:52 · section reference included

Resources

From the talk

Updates since the talk

Read the complete timestamped transcript
  1. 0:00

    [on hold music] Nice meeting you guys.

  2. 0:16

    Uh, great to be here. And, uh, I'm here to present Hyperbolic, which is a AI cloud for developers. And so my topic is, uh, why

  3. 0:26

    we don't need more data centers. It's like a very [chuckles] eye-catching title. Uh, but what I want to clarify is I still think building, building data centers is important, but just building data centers alone can't solve the problem.

  4. 0:41

    So, uh, wait. Before we get started, uh, let me introduce myself. I'm Jasper. I'm the CEO and co-founder of Hyperbolic. Um, I did my math PhD at UC Berkeley.

  5. 0:52

    Uh, finished my PhD in two years, which make me the fastest person in the history of Berk- Berkeley. And then I also won a few gold medals. So, uh, after that, I work at Citadel Securities, uh, trying to use AI and machine learning to predict the market and execute strategy.

  6. 1:06

    So I always have a passion about how to make things very efficient and how to help you to save money because everyone knows that, uh, compute actually one of the biggest costs for your companies or for your startups.

  7. 1:18

    Uh, usually you need to s- if you want to run like one thousand GPU will spend you millions of dollars, uh, per year. And we, we think that these problems should be solved by not just building more data centers, but actually, uh, building a GPU marketplace.

  8. 1:33

    So let's get started, uh, with the problem that we're facing. Uh, first, uh, I think-- So, so everyone knows that AI is gonna integrate with everything in the future, and every companies will be AI companies.

  9. 1:46

    So the demand for GPUs as well as data centers are exploding. So by McKinsey, by twenty-thirty, we'll need four X more data centers built in one quarter of the time that we build in this speed.

  10. 2:03

    Uh, but what if I tell you that you actually don't need that many data centers? Uh, you actually need, uh, another solution. So, uh, we can break down the demand first.

  11. 2:15

    Uh, right now, uh, the current capacity for data center is fifty-five gigawatts. Um, w- by the median, uh, scenario, we're going to see twenty-two percent annual growth rate for the demand.

  12. 2:29

    So in twenty-thirty, we're gonna need two hundred and ninety gigawatts.

  13. 2:38

    And however, uh, it's like there are a lot of challenges building data centers, right? So first, uh, we everyone knows Stargate, so it takes like, uh, for the first Stargate data center, it takes like more than a billion dollars to build.

  14. 2:52

    Uh, and then also it's very slow to colla- connect data center to the electrical grid. For example, right now the, the wait- wait list is like seven years. So you need to wait seven years to connect a one hundred megawatts facility to the, uh, to the electric- electrical grid in Nor- uh, Northern Virginia.

  15. 3:11

    And, uh, and then, uh, it also very, uh, consuming a lot of energy. So, uh, currently, we're spending four percent of the total electricity consumption in the US for just GPUs and data centers.

  16. 3:24

    Uh, and also is not very environmental sus- sustainable. Uh, if you can look at the number, that's crazy, uh, CO2 emissions annually.

  17. 3:34

    Uh, and even say if we're gonna deliver all the data centers, uh, on time, there's still a data center supply deficit of more than fifteen gigawatts in the US alone by twenty-thirty.

  18. 3:48

    And so it means that just building data center can't solve the problem. Uh, on the other hand, uh, we think the GPU utilization is actually pretty low. So according to, uh, Deloitte, GPUs sit idle eighty percent of the time for enterprises and companies.

  19. 4:11

    According to SemiAnalysis, there exists a hundred plus GPU clouds. So we can see like how fragmented this space is, right? A lot of you guys need GPUs, but you can't find them or like you are going to pay extremely high price.

  20. 4:28

    On the other hand, there are a lot of GPUs sit idle in data centers or in different clouds. And so naturally, uh, a solution that we think we, we should build is actually build a GPU marketplace or like aggregation layer that aggregate different data centers and GPU providers to solve the problem for, uh, GPU users.

  21. 4:50

    Uh, it doesn't necessarily need to be Hyperbolic, but I just use Hyperbolic as example, uh, to show here. [chuckles]

  22. 4:57

    Uh, so, uh, I can, I can just like, uh, share what we are s- we're trying to solve. So we're building this like global orchestration layer. Uh, we invented a software called HyperDOS, which is short for Hyperbolic Distributed, uh, Operating System.

  23. 5:13

    So basically, it's like a Kubernetes, uh, software. So any, any cluster, as long as it installed our software within five minutes, suddenly the data center become a cluster in our network.

  24. 5:27

    And on the other side, users can rent GPUs, uh, in different ways that they want. Like, they can just, uh, do the spot instance, they can like on-demand, they can long-term reserve, or they can also like host, uh, models on top.

  25. 5:43

    Um, and so like we see that-- we see there are like several benefits. Um, one, we, uh, we k- kind of like solve the effi-, uh, like the matching problem of compute.

  26. 5:55

    Uh, and then second, like GPU become commodities, so you like, you don't need to spend too much time to wait for data center. You just buy them on the marketplace.

  27. 6:04

    And then third, uh, you can have different options.

  28. 6:09

    And so, um, we do some math modeling. Uh, I, I mean, I don't have time to kind of put down the math in the slides, but this is our conclusion, right?

  29. 6:19

    Basically, uh, we can save the cost by fifty to seventy-five percent. Uh, even if you look at, uh, the current... We, we're running like some beta version of our marketplace right now, and our GPU cost for H100 is ninety-nine cents per hour.

  30. 6:36

    But if you look at Google, for example, they have on-demand GPU, it's like eleven dollars. They have like Lambda, they have like two or three dollars. But on average, by have ga- aggregating more supply, uh, and then like have a uniform distribution channel, you can drama- uh, drastically reduce the price.

  31. 6:56

    Um, it's... Like the, the theory behind that is like the Queuing theory. Basically like, uh, it's M/M/c theory. I probably... Next time if we're gonna watch my talk, I will share more math, uh, behind that.

  32. 7:08

    Uh, but yeah, and then like you can just save time to vetting your suppliers because y- if you like think about... I, I mean, how many people here are founders or like need to acquire GPUs?

  33. 7:21

    Yeah. So, uh, are you frustrated when you are trying to talk to... How many suppliers are you talking to? If you have talked to more than five, raise your hands.

  34. 7:32

    Are you frustrated when you're like trying to have like five sales calls and like try to like know which data's, uh, GPUs are, are frustrated? Yeah. Uh, are good?

  35. 7:42

    Yeah. That's good. Yeah. So basically, by having like this uniform platform, like founders or like startups or companies no longer need to vet different data center. They just like pick the one that they, uh, they have high rating or like have the best price.

  36. 7:59

    And we're also going to do like, uh, benchmarking on the performance of the GPUs.

  37. 8:06

    All right. So, uh... Oh, sorry. All right. So, uh, sorry, somehow the graph didn't, didn't show.

  38. 8:21

    Uh, let's... Give me one sec. Yeah. So, um, basically, we can think about a use case example. Um, so let's say if you, if you are a startup and you want like one thousand GPUs at the beginning, so usually you will just reserve these one thousand GPUs for a year, right?

  39. 8:43

    You think like, "I might need to use these GPUs, uh, for training, and later I want to do inference." And so you run some training jobs, and then after three months, then you realize that, okay, now I have a s- I have a ni- good-- a better idea by running those experiments, and now I need one thousand

  40. 9:02

    more GPUs just for a month, right? And then after, after six months-- at month six, then you finish your training job. And then you realize that now I only need five hundred GPUs for hosting my model, but w- I still have five hundred GPU left.

  41. 9:19

    So, uh, on the traditional, uh, on Hyperbolic case, uh, you basically can say, "Okay, I will rent one thousand GPUs for a year at the beginning." But then, uh, in month three, I can say, uh, "I just rent, uh, an, an extra one thousand GPUs for just, uh, a month."

  42. 9:40

    And then, uh, a month-- in month six, then I can say, "Okay, I can re-list my idle GPUs on Hyperbolic and try to sell the, uh, sell them to the, uh, to other people that need them," right?

  43. 9:52

    Uh, but if you just like use some traditional cloud, then you need to rent one thousand GPUs at the beginning, and then on month, in month three, you need to rent actually one thousand GPUs for a year usually.

  44. 10:03

    And, uh, if you calculate the cost, uh, compare, compare that and then also like think about the price difference you will have, uh, it will... You can reduce the cost from forty-three point eight million to six point nine million.

  45. 10:19

    So it's like six X saving. Uh, and you also help other people to get cheaper GPUs too because you can re-list those idle GPU to other people. And so, uh, so that's-- this is how we think that, uh, we're gonna, we're gonna like increase the productivity.

  46. 10:35

    Like, people only think about saving, but actually, uh, this is not true for GPU, right? Uh, by scaling law, we know that the more g- compute you spend, the better quality your machine will be-- uh, your model will be.

  47. 10:49

    So it's not just about saving your cost by six X, it's more about with the same budget, you will increase your b- productivity by six X. And imagine how many startups that they used only need to rely on OpenAI and Anthropic, those closed AI models, but now suddenly they- their money become more valuable, and they can

  48. 11:14

    rent as many GPUs as they want for the training.

  49. 11:19

    Um, and so the, the next step that we think, uh, usually the GPU marketplace will evolve into is that, uh, it will be a all-in-one platform for different AI workload.

  50. 11:30

    Because what people really want is not just GPUs. They want, um, to run their different AI jobs, right? They will... You will have AI inference, uh, online inference, uh, offline inference, and then you will also have a training job.

  51. 11:47

    And so, uh, yeah, so this is like, um, to, to like, uh, some takeaway. Like basically, we don't think we n- we need like just focus on building data centers.

  52. 11:59

    We also need to do like smarter allocation for the resources. And then second, uh, we can reduce your costs, um, for by building GPU DS-- uh, marketplace. And lastly, um, I think Uh, just focusing on building data center is not very sustainable.

  53. 12:15

    We're c- causing a lot of energy, uh, taking a lot of land. Uh, we should better reuse, recycle [chuckles] those idle compute by, uh, se- sending it to others. And so, uh, if you're interested in trying out, uh, you can, uh, come to our website.

  54. 12:33

    Uh, the, the left QR code is, uh, the current product that we have, which is a marketplace. But then we're also launching our business card and enterprise card that, uh, give you, like, production-ready GPUs with 99.5% reliability.

  55. 12:47

    All right. Thanks. [audience applauding]

  56. 12:52

    Awesome. So I actually got... I'm curious, can you tell us more about the, the kinda Hyperbolic OS? How exactly does that turn... 'Cause I know a lot of times you have a data center plus a set of GPUs.

  57. 13:05

    Yeah.

  58. 13:05

    How, how does it actually work to connect it to Hyperbolic itself?

  59. 13:10

    Yeah. So, um, basically this is a Hyper- HyperDOS is like a Kubernetes agent. So, um, you just install that in your cluster as long as you have Kubernetes. I mean, um, most data center have Kubernetes, but then even for your MacBook or for your, uh, PC, you can just install like MicroK8s to kind of, uh, become a

  60. 13:31

    Kubernetes-ready, uh, machine. And, uh, so basically now you kind of have... We, we have terminology in-house. We call, like, our Hyperbolic server, uh, Monarch, and then we have, uh, different barons. [chuckles]

  61. 13:48

    So it's like a feudalism model. So different barons, they own different compute. And then anytime... And every, every time when a user want to rent GPU, they will talk to our Monarch server, and the Monarch server will send a request to, uh, the, like, the baron, and then baron will just basically, uh, provis- provision the machines and

  62. 14:07

    set up a SSH instance for customers to access. Yeah. [outro music]