← All AI Engineer talks

AI Engineer World's Fair 2026

Building: Gergely Orosz × Simon Eskildsen

Read the talk

Building Turbopuffer: From PowerPoint Games to Object-Storage Search

Simon Eskildsen traces the engineering habits behind Turbopuffer: testing real failures, checking benchmarks against hardware limits, and making search economics work before scaling the company.

From a talk by Gergely Orosz

Before you start: Basic familiarity with databases, caches, and cloud storage is helpful; the article explains the specific failure-testing and vector-search mechanisms.

A clickable shape becomes a program

In PowerPoint, a shape can link to another slide. Connect enough shapes and slides, and a presentation becomes a game: each click moves the player into another state. That was Simon Eskildsen’s entry into programming. In conversation with Gergely Orosz, author of The Pragmatic Engineer, the Turbopuffer founder starts with those increasingly convoluted games, joking about how quickly the possibilities escalated.

Two seated participants face each other across a low table; the participant on the left smiles while the other gestures.
The onstage conversation at AI Engineer World’s Fair.

FrontPage came next. Its promise was that building a website would not require becoming a front-end developer. Then someone opened one of Simon’s sites in Firefox instead of Internet Explorer, and the layout fell apart. An accidental click into FrontPage’s HTML view exposed the machinery underneath. He began borrowing snippets to change the cursor, moved to Dreamweaver, and eventually wanted pages that could generate themselves dynamically. That led to PHP.

Around age eleven or twelve, he exhausted the programming advice he could find in Danish. Four years of World of Warcraft supplied an unexpected prerequisite for going further: English. Looking back, he wonders how different that path would have been with an LLM that could explain programming in Danish. Instead, learning English opened a much larger technical web, and he began taking programming jobs during high school.

An internet friend on Australia’s International Olympiad in Informatics team suggested that Denmark probably had a team too. Simon found its website and encountered problems very different from HTML and PHP. His illustrative example is to assign M packages, each with specified dimensions, among N trucks while optimizing the arrangement. The point of his NP-completeness shorthand is the difficulty of finding optimal solutions efficiently as problems grow; it does not mean that individual instances cannot be solved. Competitive programming gave him practice turning a loosely familiar situation into an algorithmic problem.

0:300:51
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:30 · section reference included

A broken phone and an unexpected apprenticeship

Shopify found Simon through an article, rather than through his competition results. After dropping and breaking his iPhone, he returned to a Nokia brick phone. He wrote about calling people again and recovering his sense of direction. In his account, the article briefly reached Hacker News, The New York Times featured it, and a Shopify recruiter connected the dots. The company apparently had not realized he was still in high school.

Invited to interview in Ottawa, he first had to establish what Ottawa was. Walking into Shopify felt right. He finished high school, moved to Canada, and joined in 2013, initially imagining a one-year gap before university.

Not having studied computer science made him insecure, but IOI had already taught him that he could sit with a paper long enough to understand it. At Shopify, he wrote down unfamiliar terms during the day and studied them that evening. If somebody mentioned TCP, he assumed they understood the three-way handshake, how TLS sat above it, and what the exchange looked like in Wireshark. He later realized that assumption was too generous, but it gave his self-education considerable depth. Soon he no longer wanted to leave the work he had found.

That impulse to peel back layers is now something he looks for when interviewing engineers. While still working on product, he sat beside infrastructure engineers at lunch and asked about their work. Even terminology became a reason to investigate: why is a reverse proxy reverse, or an inverted index inverted? The naming jokes recur later, but the underlying habit is serious—an abstraction invites a look underneath it.

4:034:23
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

4:03 · section reference included

Scaling means deciding what happens when dependencies fail

On Shopify’s infrastructure team around 2013–2014, Simon helped containerize systems as Docker emerged. Simon recalls Shopify growing roughly 120–140% year over year. Preparing for Black Friday meant forecasting demand and ordering physical hardware in advance, while also making the software handle the next peak. Application bottlenecks repeatedly led him to the boundary between Rails and the databases. Much of the work there involved orchestration rather than modifying database internals.

His boss Camilo’s shorthand was that you cannot cache writes: a read cache does not remove the underlying requirement to persist incoming changes. Eventually the write workload must move beyond one shard. Simon joined around Shopify’s sharding transition and recalls a cutover about a week before Black Friday. Subsequent work expanded into multiple data centers and untangled a large shared Redis instance that had become a catch-all key-value store. When it failed, the team discovered how little they understood about its dependencies.

A dependency outage should have an explicitly designed scope. If the session store fails, the entire storefront should not automatically disappear with it. The team split apart the shared Redis responsibilities and built a matrix: for each service and failed component, what behavior should remain available? Writing down that expected degradation exposed the next problem—how to test the failures through the real application stack.

9:129:21
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

9:12 · section reference included

Test the connection, including the driver

Simon’s first fault-injection experiment attached GDB to the application process and closed the file descriptor connected to the database. This forced failure through the connection-handling layers instead of replacing them with mocks. It uncovered substantial Rails issues, but the approach never shipped in CI. He then built Toxiproxy: a proxy placed between the application and its database connection.

Toxiproxy operates at the TCP layer; Simon corrects his initial layer-seven description to layer four. Its control API lets a test make a connection slow or unavailable while the application continues using its real database driver. Although he also mentions later layer-seven additions, the mechanism established here is transport-level fault injection. The driver’s recovery behavior becomes part of the test, rather than something a mock silently bypasses.

Orosz sketches a block that takes a database dependency down while fetching a page or exercising checkout. In Ruby, the documented indexed proxy selection and .down block express that pattern. With the application’s session connection routed through a proxy named sessions, an illustrative storefront availability check is:

ruby

require "toxiproxy"
require "net/http"

storefront = URI("http://localhost:3000/")

Toxiproxy[:sessions].down do
  response = Net::HTTP.get_response(storefront)
  raise "Storefront unavailable: #{response.code}" unless response.code == "200"
end

The assertion represents one cell in the failure matrix: losing sessions must not make this public page unavailable. Other cells need their own expected behavior.

Orosz describes tens of issues exposed across Rails and the MySQL driver. Production outages are poor opportunities to investigate those details: everyone is trying to restore the database, rather than studying how the application could have behaved better without it. Repeatable faults turn that emergency into a development exercise. Simon says the resulting tests worked well and believes Shopify still runs them in CI, though he does not confirm their current deployment.

11:2711:41
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

11:27 · section reference included

What should this operation cost?

The planned gap year became eight years at Shopify, from 2013 to 2021. Simon left because he wanted a different environment in which to learn. His work had ranged across caching, database scaling, multiple data centers, and traffic spikes from Kardashian product launches. Simon says the storefront rewrite he undertook with future co-founder Justine powered almost all storefront traffic eighteen months after they began. He left without a fixed next venture, but with a continuing project: Napkin Math.

Napkin Math collects hardware latency, bandwidth, and cost estimates: DRAM throughput, NVMe and EBS bandwidth, S3 round trips, memory prices, and differences between Spot pricing and longer commitments. Simon describes roughly fifty reference values, a Rust script to generate measurements, and flashcards to keep the numbers available in his head. His examples include memory at approximately $2 per GB and object storage at $0.02 per GB; the reference table labels these as monthly estimates, rather than current provider quotes.

The project grew out of infrastructure reviews. A team would benchmark database A, find it slow, and propose database B. Simon wanted to understand why A was slow before accepting the replacement. For a three-term search, he would estimate the matching document lists, the bytes that must be read, and the work required to intersect those lists. In his illustrative search calculation, an assumed multicore DRAM bandwidth of 100 GB/s suggests about 10 ms of work, against a reported 10-second benchmark. That gap is a question, not a verdict: either the model is missing something or the benchmark is measuring something unexpected.

For example, a supposedly simple query might fan out across a hundred nodes. Its P99 latency would then reflect distributed coordination and slow responders, not just local list intersection. The same reasoning applies to a B-tree: estimate pages visited and the cost of random reads, then compare that model with the observed query. A discrepancy points toward a query plan, a database bug, a disk problem, or an incorrect assumption. A benchmark tells you what happened; a cost model gives you something to investigate.

14:5615:06
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

14:56 · section reference included

When the estimate is wrong: MySQL and fsync

That method also catches errors in the mental model. Simon once reasoned that if each durable write needed its own fsync, database write throughput should equal the number of sync operations the disk could complete:

tfsync=1msRone write per sync=10.001s=1,000writes/s\begin{aligned} t_{\mathrm{fsync}} &= 1\,\mathrm{ms} \\ R_{\mathrm{one\ write\ per\ sync}} &= \frac{1}{0.001\,\mathrm{s}} \\ &= 1{,}000\,\mathrm{writes/s} \end{aligned}

Simon recalls observing about 10,000 writes per second on a small MySQL machine, despite an initial estimate of 1,000 writes per second based on one 1 ms fsync per write. The machine was doing more work than his model allowed.

The missing mechanism was batching. Multiple writes can share a synchronization operation, so writes per second need not equal fsync calls per second. Finding that explanation took roughly a day of pre-LLM investigation: BPF traces, unexpectedly large writes, source-code reading, and a detailed MySQL article. His joke that the internet runs on small towns in Bavaria pays tribute to the obscure expert writing that finally makes a system understandable.

19:0919:18
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

19:09 · section reference included

A useful feature that could not afford to ship

Three experiences converged into Turbopuffer. First came an unsatisfying search project at Shopify. Simon leaves the vendor unnamed, but describes a traditional search system that was difficult to operate, lacked a query planner in his account, and repeatedly failed to match his hardware estimates. Reading its source did not consistently explain the discrepancies. Second, Napkin Math had given him a strong sense of what a machine should be able to do.

Third came what he called Angel Engineering: working inside friends’ companies for vested equity rather than merely investing money. After ChatGPT appeared in 2022, one of those companies wanted to connect documents to AI. The early context windows were small enough that retrieval quickly became necessary; the speakers describe them in kilobytes, without identifying a model or establishing a token limit.

At Readwise, Simon built a recommendation prototype that was surprisingly revealing. While testing on a co-founder’s feed—with permission, he stresses—the recommendations led him to infer that the co-founder’s wife was pregnant. The feature worked well enough to justify calculating what it would cost to offer it to everyone.

Simon estimated that deploying the Readwise recommendation feature to all users would cost $30,000 per month, compared with roughly $5,000 per month for its other infrastructure at the time. For a bootstrapped company, the feature needed to earn enough to cover that cost and leave a gross margin. It did not meet that test, so they did not ship it. Simon returned to work such as tuning Postgres autovacuum, but could not stop wondering why storing the vectors had to be so expensive.

21:0521:21
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

21:05 · section reference included

Cursor becomes the first customer

Orosz first encountered Turbopuffer while investigating Cursor’s infrastructure. He had assumed Cursor was one of several early customers. Simon corrects him: Cursor was the first. The contact came after a Twitter launch intended to answer whether anyone cared enough for him to keep working on the project. At that point it ran on an eight-core GCP node; he intended to block production adoption until he had set it up properly across multiple machines.

The launch pitch was a million vectors for a dollar. That historical offer was monthly storage pricing, with queries charged separately, as clarified in the company’s technical retrospective. Simon remembers the cheapest working alternative costing perhaps a hundred dollars per million, but does not identify it or give comparable billing conditions. The small deployment did preserve a non-negotiable durability invariant: writes went directly to objects, so shutting down every VM would not lose the committed data.

For a coding product, the storage arrangement matched the workload. An inactive codebase could remain in object storage; opening it would bring the active data into memory, where subsequent queries could run quickly. The cache would follow usage rather than requiring every stored vector to occupy DRAM continuously.

DataPlacementPurpose
Committed vectorsObject storageDurable, economical retention
Active codebase dataMemory cacheFast repeated queries
Inactive codebase dataObject storage until neededAvoid permanent memory cost

Simon imagines Cursor’s founders arriving at this same design over dinner, but explicitly says he does not know whether that conversation happened or whether they considered building it themselves. What he does establish is that Turbopuffer matched a problem they wanted solved.

Simon flew from Canada to Cursor’s San Francisco office. On arrival, he found a Postgres problem under discussion and suggested setting up pganalyze. His diagnosis centered on insufficient autovacuum and unnecessary heap access. The precise PostgreSQL distinction is an index-only scan: ordinary index scans already fetch heap rows, while index-only scans can avoid those fetches when visibility information permits it. Helping investigate the existing database gave Cursor a practical reason to trust the person proposing a new one.

By then Justine had joined as co-founder. Her first change replaced the nginx proxy cache with a direct file-based cache. Cursor decided that evening to migrate and completed the move over roughly one or two weeks. Simon reports that Cursor’s first Turbopuffer bill was 95% lower than its last bill from the previous vendor. When Orosz identifies that vendor as Aurora, Simon corrects him: this saving was not against Postgres; the replaced vendor remains unnamed.

Orosz says reliability was a principal motivation for Cursor and recounts Sualeh’s warning against betting a business on a tiny infrastructure supplier—with an exception for Turbopuffer. The apparent risk became more understandable after Simon showed up, helped with a real problem, and brought years of operating Shopify to the relationship. Trust came from demonstrated competence and direct cooperation.

28:4028:58
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

28:40 · section reference included

The CPU advantage meets the capacity limit

Turbopuffer’s enthusiasm for CPUs produced an unusual encounter at NVIDIA. Presenting to Jensen Huang and company leadership about potential partnership opportunities, Simon joked that Turbopuffer could pivot into vapes if the database business failed. In Simon’s retelling, Jensen replied, “Judging by your slide, maybe you should.” Simon then asked whether Jensen vaped, prompting a teammate to relay the exchange to the company.

The team had warned him not to say the C word—CPUs. He nevertheless kept praising them: plentiful machines, SIMD, and AVX-512 were attractive ingredients for Turbopuffer’s workload. The anecdote ends with Jensen taking an interest, rather than with an announced GPU migration.

But CPU availability was no longer as easy as that pitch suggested. Simon connects the pressure to reinforcement learning environments that execute real tools. Teaching a model to search, run grep, or use Bash requires CPUs on which those operations actually happen. Deployed agents also perform general-purpose work. As applications expand into domains such as CAD, more environments are needed to teach and evaluate those capabilities, feeding demand back into training.

Turbopuffer needs NVMe SSDs as well as CPUs, and memory demand overlaps with GPU-server requirements. Simon expects CPU availability to worsen before improving, while declining to make confident predictions about the broader GPU market. He describes competing for allocations with companies that are also customers and writing software to acquire capacity quickly. Orosz adds a report from Reflection: despite long contracts, it could not obtain more of the CPU and GPU capacity it wanted.

35:1735:31
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

35:17 · section reference included

Run where the machines are

Capacity planning now includes conversations with clouds and large customers about which regions have CPUs and where electrical power will support incoming hardware. Turbopuffer has a useful degree of freedom: a cluster needs CPUs, NVMe SSDs, and object storage. That makes additional clusters relatively straightforward to deploy, although architectural work to accommodate scarcity still consumes engineering time that could go elsewhere.

Supporting many instance types reduces dependence on one allocation pool. Simon names GCP’s C4 machines as a favorite, says C4D performs well after targeted optimization, and also likes the Arm-based C4A family. A SKU is simply a particular machine offering; examples such as i8g illustrate the wider set of options. This flexibility echoes Shopify’s Black Friday/Cyber Monday planning, when teams forecast usage and made cloud commitments months ahead. The cloud feels infinite mainly while your requirements are small.

41:1941:29
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

41:19 · section reference included

Price the efficient implementation, then build it

The financing story begins with an engineering promise. Simon says he guaranteed Cursor roughly $4,000 per month based on the estimated cost of a more efficient implementation than Turbopuffer then had. The software was reliable but simple, and the founders had work to do before their own infrastructure cost matched the price they charged. Their long tenures at Shopify—and Orosz’s at Uber—had also made them attentive to how well simple software ages.

Simon was not yet sure whether this niche search engine could become a venture-scale company. Taking venture capital would create an obligation to deliver a large return on a timeline, extending through a chain of investors that might ultimately include pension funds. His immediate operating model was more concrete: compare Cursor’s bill with the GCP bill, then optimize until infrastructure costs were roughly covered. Additional workloads might eventually pay the founders’ salaries.

He had few financing relationships and describes himself as an outsider twice over: from Aarhus to Ottawa, then from Canada looking toward San Francisco. Before accepting an investment that would redefine the business’s success threshold, he wanted more evidence that he could meet that threshold.

A specific hire changed the calculation. Simon wanted to work with Boyan, whom he knew from the North Macedonian IOI team; his teammates had nicknamed him God. But the founders had gone about six months without salaries and had already spent tens of thousands of dollars on GCP. Simon called Laki, one of his Silicon Valley contacts, with a bounded request: approximately $700,000 to support two engineers for the rest of the year, plus a buffer, while the founders remained unpaid.

His proposal included a stopping condition: if product-market fit and a large opportunity were not evident by year-end, they would shut down and return the remaining money. Some investors heard low ambition. Simon saw candor about what he knew and what he still needed to learn. They raised funding, hired Boyan, became profitable later that year, and continued hiring as their conviction grew.

43:0643:16
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

43:06 · section reference included

Be explicit about why you are raising

Simon separates fundraising into six motives:

  • Research and development: Pay to learn faster. Turbopuffer’s January raise funded its first engineers, Boyan and Morgan, after the founders had financed development through unpaid work and their own cash.
  • Growth: Spend to bring an existing product to a larger market.
  • Founder ego: Seek a large number and the attention surrounding it. Simon considers this a dangerous motive because it dilutes existing employees and changes the entry price and potential upside for future employees.
  • Employee rewards: Provide liquidity during a long company-building journey. He says Turbopuffer took capital in December so employees could sell some equity without waiting for an IPO or another distant event.
  • Strategic partnerships: Bring in a relationship that can materially change the company’s trajectory.
  • Mergers and acquisitions: Finance an acquisition or related transaction.

The distinction matters because the same fundraising announcement can conceal very different purposes. In Simon’s account, Turbopuffer’s first raise served R&D; its second served employee liquidity.

49:1849:24
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

49:18 · section reference included

A distributed company can still gather around a campfire

The closing question turns from machine allocation to how people work together. Orosz notes that many AI companies favor an office for faster iteration. Remote work felt natural to Simon because Shopify’s infrastructure team had long been distributed: persuading every specialist to move to Ottawa was difficult. Outside established database-company hubs such as San Francisco and perhaps New York, he sees a distributed model as a way to reach the people needed to build the company.

Distributed does not mean never meeting. Turbopuffer brings everyone together twice a year, with Banff and Mexico City among the examples. It also names smaller, informal gatherings campfires. When several colleagues happen to be in one place, they invite others to join. The San Francisco campfire around this conference included customer meetings and dinners, turning an existing trip into time together.

Participation beyond the company offsites is optional. Some employees attend those two gatherings and otherwise stay home with their families; others travel frequently. Both fit the model. Simon recounts one employee seeing colleagues together in a New York meeting room, taking an Uber to the Ottawa airport, and flying down to join them. The point is to make that spontaneity possible without making it everyone’s obligation.

Turbo credits add a small incentive. Giving a conference talk or publishing a blog post can earn a credit redeemable for a business-class upgrade on the next flight. The idea has already prompted jokes about a central bank, interest rates, and a betting market for credits. Volunteering for two days on a conference expo floor can also earn one: it is taxing work, and the reward encourages future time with the team.

Orosz closes by noticing how little of this AI infrastructure conversation required talking about AI itself. Its substance was engineering judgment, curiosity, and the human connections that make difficult work possible. The same company that places durable data in object storage and active data in memory also makes deliberate room for people to meet, help customers, and build trust.

51:4552:05
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

51:45 · section reference included

Resources

From the talk

Read the complete timestamped transcript
  1. 0:12

    [upbeat music] [laughs] All right. It's great to be here, and today with me here...

  2. 0:30

    Uh, I'm, I'm Gergely, author of The Pragmatic Engineer, and I'm excited to have a chat with Simon Eriksson, uh, founder, CEO of [REDACTED:username]. A very technical CEO, and we're gonna have a pretty technical discussion.

  3. 0:43

    But before we jump into it, Simon, I wanted to ask: Where did you fall in love with computers?

  4. 0:51

    Um, through PowerPoint.

  5. 0:54

    PowerPoint?

  6. 0:56

    You- I don't know if any of you know this, but in Power- Well, you probably know this, but in PowerPoint, right, you can make the, the diagrams and stuff.

  7. 1:02

    When you click them, go to another slide.

  8. 1:07

    That becomes Turing complete real quick, right? You can sort of, you know, create very complicated- [laughs] ... convoluted games, and then at some point, you know, you make it through the Microsoft Office suite and you discover FrontPage.

  9. 1:18

    Do you remember FrontPage?

  10. 1:19

    Yeah, I remember FrontPage. It, it, it was supposed to eliminate the need for all, uh, any front-end developers.

  11. 1:25

    Exactly. And it w- it only worked in Internet Explorer. I remember a heartbreak I had one day when someone opened a website I created in Firefox, and it just [laughs] it was, it was all over the place.

  12. 1:36

    And then one day, I accidentally clicked the HTML thing in FrontPage, and it just showed all of this stuff that I couldn't make sense of, and I just started looking at it, and then going online and finding little snippets that you could add in to make the cursor change and all of these different, uh, different things.

  13. 1:53

    And then it just sort of escalated from there. Then you upgrade to Dreamweaver, and now you're coding. And then you're like, "Well, how do you make the pages dynamically?"

  14. 2:01

    You learn PHP. And then for me, I exhausted the internet on [REDACTED:origin] language programming advice.

  15. 2:08

    Mm.

  16. 2:09

    Um, and I was, I was around [REDACTED:age] or [REDACTED:age].

  17. 2:13

    And so I just, you know, went and got addicted to World of Warcraft for four years, but that gets you really, really good at English. [laughs]

  18. 2:21

    So you kinda start hacking, get into deeper. Now, the logical step would have been to just, you know, go to university and learn properly about this stuff, but that's not what you did, did you?

  19. 2:33

    I mean, I just, um... I, I started just... I, I mean, you know, then I learned video games, then I learned English, and then, you know, this, like, massive arsenal of the web.

  20. 2:45

    I- now it'd be very interesting 'cause the LLMs would just speak [REDACTED:origin] to me, and you could just... You wouldn't have hit the wall like I did.

  21. 2:51

    Yeah.

  22. 2:52

    Um, so that would've been very interesting. Maybe I would've been better at programming. That would've been nice. And then I... Yeah, then I just started picking up jobs and things like that throughout high school.

  23. 3:02

    And when I was in high school as well, I got exposed to this thing called the International Olympiad in Informatics.

  24. 3:07

    Mm-hmm.

  25. 3:07

    Heard of this thing?

  26. 3:08

    Yeah.

  27. 3:09

    Um, and I had a, I had an internet friend, and she lived in Australia, and she was on the [REDACTED:origin] team. And she told, "Oh, there's probably something for the [REDACTED:origin] team as well," but I had never, I'd never heard about it before.

  28. 3:21

    And so I found it on some, like, little mysterious website, and then applied, and then solved these programming problems that looked very different from the HTML and PHP things that I'd solved until then.

  29. 3:31

    What were these, like, the algorithmical-ish programs-

  30. 3:34

    Exactly

  31. 3:34

    ... problems?

  32. 3:34

    It's sort of like this is not actually the kinda problem you would see there, but I think it illustrates well the kind of problem that you might get, right, is you could imagine something like, okay, here's, like, N trucks.

  33. 3:45

    Here's M packages. The M packages have these dimensions.

  34. 3:51

    Give me which trucks which packages should be in, right?

  35. 3:53

    Yeah.

  36. 3:53

    And then do something optimal. Like, that's an NP-complete problem. You can't solve that, but you could compete with everyone else in the competition of doing the best thing.

  37. 4:01

    Yeah.

  38. 4:01

    So it's these kinds of problems, right?

  39. 4:02

    Yeah.

  40. 4:03

    Um, and so I started doing that in high school. Was working, um... I was working as well, um, for a, a startup. Um, and then I just... Shopify found me while I was still in high school.

  41. 4:17

    And, and the whole, like, Shopify found me, was it through your open source contributions? Was it, was it something else?

  42. 4:23

    It was because I had written an article where I had, I had... I'd dropped my iPhone, and it was... You know, the iPhones are a lot... Like, there used to be a time, right, where you dropped your iPhone, and you just knew it was over for the screen.

  43. 4:37

    Yeah.

  44. 4:37

    It doesn't really happen as much anymore. Like, the screens have gotten a lot better. But back then it was like, yeah, one drop and it was dead, and it just...

  45. 4:43

    I couldn't use it anymore. And so I went back to one of these old Nokia brick phones, and this is back in 2013, and people hadn't really realized all the pernicious effects of smartphones at the time.

  46. 4:53

    And so I wrote this article about how, oh my God, I'm, like, calling people [laughs] and I have my sense of direction back. Um, and I wrote an article about it, and this article, it went on Hacker News briefly, and it, um...

  47. 5:05

    New York Times decided to feature it.

  48. 5:07

    No way.

  49. 5:08

    Yeah. And so a lot of traffic was driven to it, and then some astute Shopify recruiter put it all together, and, um, and I had a call with them, and then...

  50. 5:20

    I don't think they realized that I was still in high school. But, um, but I had a great call with them. They invited me on site to Ottawa, Canada.

  51. 5:28

    Um, I had no idea what Ottawa, Canada is. I think the email says something like, "What's an Ottawa?" I had no idea. Um, and so I went there, and it was just, like, walked into the building, and it was j- just a...

  52. 5:41

    It just felt right. Um, and so I, I, I interviewed with them, and then said, "Well, I gotta finish high school first." And then, uh, and then I moved to Canada, uh, to, to, to work at Shopify, yeah, in 2013.

  53. 5:52

    Yeah. I think that's, that's a, like, legit excuse for, like, not even worrying about [laughs] college and, and university.

  54. 5:59

    But I, I did.

  55. 5:59

    Did it cross your mind?

  56. 6:01

    It did. I thought I was going... I thought I was doing a gap year. I thought I was like, "Okay-

  57. 6:05

    Really?

  58. 6:05

    ... I'm gonna go work at Shopify for a year, and then-

  59. 6:07

    Mm

  60. 6:07

    ... I'll probably go back and l- do... But I would just...

  61. 6:11

    I was very insecure at the time about the fact that I hadn't studied computer science, and my only exposure had been all the IOI competition, which is a pretty good crash course in a lot of computer science.

  62. 6:21

    And if nothing else, it had really taught me that you can just sit down and read a paper and just figure it out if you spend enough time on it.

  63. 6:28

    So I did, I did that repeatedly, and w- in my first year at Shopify, I just, every time I heard something that I didn't know what was, I noted it down on a piece of paper, and then I went home, and then that evening I would just read about it.

  64. 6:38

    Mm-hmm.

  65. 6:38

    Because I felt insecure that, like, well, if someone mentions, mentions, like, TCP, surely they know exactly what's in the three-way handshake and how TLS is, like, layered on top, and they've looked at Wireshark and all of that.

  66. 6:49

    I don't think that's true, but that's what I thought.

  67. 6:51

    Yeah. [laughs]

  68. 6:51

    So I go in and did that for everything that I encountered. Um, so that was a really good crash course, and then very quickly it became clear that, well, I just wanna continue doing this.

  69. 6:59

    I don't wanna go, go somewhere else and then come back to this, 'cause I felt like I'd already found what I wanted to do.

  70. 7:06

    So it sound- sounds like it was a pretty good combination of, like, you just having this, like, very natural insecurity. Like, you know, you're young, you know you don't have the education that everyone else has, and inside a company that's just doing pretty, like, cutting-edge stuff even at the time, and e- e- even t- today, right?

  71. 7:20

    Like, they're, they're leading. So you just kept self-teaching yourself, like, just catching up and go- And, like, do I understand that you just went deep in every concept that you under- So you didn't, like, just, like, try to just at a surface level, but, like, go as deep as you can, search on the internet, buy books, whatever

  72. 7:35

    that is?

  73. 7:36

    I think it was just that I just wanted to know- keep learning how computers work, and I think that this is something that I now look for when we interview engineers, is that you just, y- you can't help yourself but trying to peel back the layers.

  74. 7:49

    And for me, that ended up with the infrastructure layer. That was, you know, the people closest to the metal at, at Shopify. And I would just always sit next to them at lunch, 'cause I was working on the, on the product side.

  75. 8:01

    But I just, I couldn't help myself. I was so... I just wanted to learn what it was. When they were talking about a reverse proxy, I'm like, "Why is it reverse?"

  76. 8:09

    I t- I still can't answer that. [laughs] I, I, I mean, o- okay. [laughs] Do you know? [laughs]

  77. 8:19

    Well-

  78. 8:20

    What's in reverse? [laughs]

  79. 8:22

    'Cause it's a proxy, right?

  80. 8:26

    I don't, I don't know. I don't know. It's like an inverted index. Like, what's inverted? It's like, it's a terrible name. Anyway.

  81. 8:32

    Yeah. I, I, I mean, it's still better when, when you get to the NAT, NAT tables, the lookups, some of those things. Like, some of that... But yeah, I, I hear you.

  82. 8:40

    There, there's some, like, weird names with this.

  83. 8:43

    But at, at Shopify, what were some of the kind of, like, hard engineering challenges that you s- Engineering challenges, outages, like, like, learnings that kind of defined you, that were really also fun at the time or interesting to learn, but it, it would've been hard to get it elsewhere?

  84. 9:03

    Yeah, so I think it was... You know, in the 2010s there was, like, a bunch of SaaS companies that, that scaled really quickly, and I felt so fortunate to have a front row seat to that.

  85. 9:12

    And so I ended up on the infrastructure team, and this was back in, you know, '13, '14, and, uh, Docker was coming out, and so we were containerizing everything.

  86. 9:21

    And we were just, every single year we had to... You know, the growth rates of, of, of SaaS sometimes seems quaint in comparison to the growth rates of companies today, but it was a company that was growing at, you know, 120, 140% year over year.

  87. 9:35

    Um, and so every year we were just preparing for a Black Friday that was gonna be a lot worse than the last. And this is back in the day of we're buying physical hardware, right?

  88. 9:43

    We have to, like, place an order at a particular point in time and do some interpolation based on that. Um, and the software also had to scale. And when you're scaling most software, a lot of the application layer problems end up back at the database layer.

  89. 9:55

    Mm-hmm.

  90. 9:56

    And so I just naturally found myself at this layer between Rails and the databases. Shopify didn't, at the time at least, contribute many patches to the databases themselves, but mostly just spent time orchestrating.

  91. 10:08

    So we were doing sharding, because as, um, my, my dear boss Camilo used to say, you can't cache writes. So there's a fundamental point where you, you just, you have to move beyond a single shard.

  92. 10:22

    Um, so I wasn't... I joined around the time when they did the sharding, and they did it, I think they did the cut-over a week before Black Friday, which is mind-blowing, uh, and very...

  93. 10:32

    But it worked. And then the, the subsequent years we worked on things like going into multiple data centers. We also had this big, mysterious Redis server that was, like, you know, 128 gigabytes of RAM, which was a lot at the time.

  94. 10:44

    Today it's not that much. And no one really knew what was in it, and then it went down one day, and people were like [laughs], "Well, that's super terrifying."

  95. 10:52

    Um, because people had just been treating it as this KV store. Um, and so we started splitting it out. We did all this stuff around making sure that if you, if you, if you go visit a Shopify store and the thing that stores your sessions is down, the right behavior is not just for the entire, the, of

  96. 11:11

    everything to be down. But that's kind of the default failure mode, right?

  97. 11:13

    Yeah.

  98. 11:13

    You're not gonna rescue all of that, um, unless you're in a programming language that really forces that decision. So we did things like, um, build this matrix out of, okay, well, this service when this component is down should act this way.

  99. 11:27

    Um, and I found myself writing the test suite for a bunch of that. And then I was like, "Okay, well, we can't just mock all of this," and so, um, I came up with this idea at the time of like, oh, what we're gonna do is we're just gonna, um, shell out to GDB and then into the

  100. 11:41

    process, and then close the file descriptor to the database to simulate d- through the entire layer that the database fails. That was a little crazy, and we never shipped that on CI, but it did uncover a massive amount of issues in Rails to be upstream and things like that, of j- around just, like, handling failures at the

  101. 11:57

    connection layer. So then I moved on to create this proxy called Toxiproxy, and-

  102. 12:02

    Toxiproxy.

  103. 12:02

    Have you heard of this before?

  104. 12:03

    No, no.

  105. 12:04

    Yeah. Toxiproxy is, it's just, like, a layer 7 proxy that sits in between, um- You and, well, layer four, but in between you and the databases. So you basically have just, like, this proxy and then MySQL, whatever, doesn't speak the protocol-

  106. 12:18

    Yeah

  107. 12:18

    ... but then you can do an API call, say, take, take, uh, take the database down, make it slow. Um, and over time it also added layer seven things of, like, do a bunch of failures.

  108. 12:28

    This way you're not mocking the low-level drivers, but you're testing the drivers and their failure handling as well, so then this entire matrix could be implemented in CI.

  109. 12:36

    So the, basically the proxy was just, like, a really thin layer which, like, was passed through, but you build the functionality to, like, simulate problems of database or things like data corruption or whatever you want it to do, so you could just do it in there and then you can...

  110. 12:51

    Anything that built on top of it. But, but then, uh... Oh, yeah, and then everyone had to, like, call this proxy or it needs to be on the, uh, in a, in a layer.

  111. 12:58

    Exactly.

  112. 12:58

    Oh, you could do, like-

  113. 12:59

    So it's like a-

  114. 12:59

    Do, like, MySQL, you know, Toxiproxy.MySQL.down, and then pass it a lambda of what you want it to do. Like, get this page, do a checkout, whatever, with the sessions table down.

  115. 13:09

    And this just uncovered tens of issues, right? In the MySQL driver, in the rails... Like, it's just like no one in the ecosystem had been testing for this, and it was very difficult to see this in prod, right?

  116. 13:19

    Because when MySQL are down, you're focused on just getting it back up and not, like, what could the application actually have done. Yeah, it's, it's interesting. Of course, we're gonna talk a bit more about databases obviously.

  117. 13:29

    But just thinking about how a lot of the problems or some of the most gnarly pro- problems in large systems are always to do with state, and I never connected until now that, I mean, state is usually there's a database.

  118. 13:41

    If there's no database, if you have stateless services, you know, y- I mean, you still have problems. You have nodes going down. You have, I don't know, corruption, whatever.

  119. 13:48

    But it's usually, like, more isolated. But basically, like, if we have state, we typically have databases. If we have databases, then if you can simulate these problems, suddenly you can...

  120. 13:59

    I mean, you, you can, like, predict a lot of things. So the problem with state oftentimes is it's really hard to simulate problems happening ahead of time unless when they happen.

  121. 14:08

    So did you... It sounds like you had pretty good success with, with-

  122. 14:10

    Yeah. I think I've... To my knowledge, it's still, um, running in, like, the CI system of Shopify today. I don't know if anyone in the crowd is from Shopify, but I'm pretty sure that it still does.

  123. 14:21

    Um, and so we wrote all these tests against it to implement all of these different, uh, different failure conditions and it just... Yeah. It was, it was... It, it worked out great.

  124. 14:30

    So you spent eight years in total at, at Shopify. So, like, started from, like, all right, just a gap year. It just went on a year, a year, and another year.

  125. 14:36

    Um, at what point did you think about leaving and why? And what was your kind of decision framework? It sounds like you were, you were, like, on epic running out.

  126. 14:47

    Even today, Shopify is doing wonderful. It's probably doing even way better than your... Like, like, you know, that growth kind of kept on. So I'm sure there would've been an argument to stay and, you know, stay on the rocket ship.

  127. 14:56

    Yeah. So I, I spent, I spent eight years there from '13 to, to '21. Um, and I, I think there just came a point where

  128. 15:06

    I wanted to see something different. Again, I'd been inside of Shopify since I was [REDACTED:age], right? I'd been seen one other startup in high school. I was like, if I wanna learn more about computers and learn faster, it might be time to inject some novelty into this function.

  129. 15:20

    Um, and so I, I left in, in, in '21, and I'd worked on so many different parts of the infrastructure, like caching, um... Me and Justine, who's now my co-founder, rewrote the entire storefront, um, storefront for Shopify, um, which powered almost 100% of traffic 18 months after we embarked on it.

  130. 15:38

    Um, we've worked on running Shopify in multiple data centers. We've worked on so many s- database scaling projects like caching, all of these different things, right? Um, a lot of the, a lot of the scalability came from the Kardashians launching lots of products on, on Shopify, which would force a lot of traffic.

  131. 15:54

    Um, but that's, that's eventually how I left. And so when I left, I didn't really know what I wanted to do, and so I... One of the projects I had while I was at Shopify was this Napkin Math project.

  132. 16:06

    Have you seen this?

  133. 16:07

    Napkin Math? No.

  134. 16:08

    No. Um, so Napkin Math was essentially just this table that I maintain on GitHub of how much bandwidth can you drive to DRAM, what does a round trip to S3 cost and how long does it take,

  135. 16:24

    how much bandwidth can you drive to an NVMe SSD, how much bandwidth can you drive to an EBS volume? Just a collection of probably... There's probably, like, 50 of these numbers, and then a Rust script that generates them all.

  136. 16:36

    Um, what all these things cost. What do you... Like, what does a gigabyte of memory cost? $2. What does a gigabyte of S3 cost? Two cents. What does a gigabyte of, um, this cost?

  137. 16:45

    10 cents, right? What does it cost on Spot? What does it cost on a three-year commit? Like, I just have a massive table and then create flashcards for almost every single cell so I know all these numbers.

  138. 16:54

    And this was a project I started taking on at Shopify because I found myself in, um, this role a lot where I would go in and review a project, right?

  139. 17:03

    So some product team would be like, "Okay, we gotta do... We gotta build this thing, so we gotta build this infrastructure to support the feature." And a lot of the times they would say, "Okay.

  140. 17:10

    Well, we've gone and benchmarked it on database A,

  141. 17:14

    but the benchmarks are not very good, so we're gonna go with database B." [laughs]

  142. 17:20

    And I hate benchmarks so much because that's not a satisfying answer to me. To me, it's like

  143. 17:31

    this does not jive with my intuition. Database A that you're saying takes 10 seconds to do this

  144. 17:38

    should take 10 milliseconds if you do the Napkin Math, right? If it's a search query, right, it's like, okay, you're searching for three terms. They're... E- each term has this many documents that match it.

  145. 17:49

    That's this many megabytes. We inter- intersect these many, this many lists. You have DRAM bandwidth on multiple cores of 100 gigabytes per second. This should take 10 millisecond. You're telling me the benchmark takes 10 second.

  146. 18:02

    One of us is wrong. Either there's a gap in my understanding, which is very likely, or you have benchmarked the wrong thing. And in some ways, some reasons, right, it's like, okay, you've done a benchmark.

  147. 18:12

    You don't, didn't realize that your benchmark is doing a distributed query across 100 different nodes, and so if- Of course the P99 is gonna be really, really high, right?

  148. 18:21

    Unless you've cut that off or, or made some different set of trade-offs. So just found myself in these discussions repeatedly where people were making infrastructure decisions based on poor benchmarks.

  149. 18:31

    And so I needed some, like I needed some ammo to go in and just be like, "Okay, we can just do the calculation right here and then," um, because I was always doing these like little demos or like writing little prototype scripts to, to demonstrate this.

  150. 18:43

    But it was just... I just... The argument of, here's how a B-tree works, this is how many pages we have to visit, this is what a random SSD read takes, it takes one millisecond.

  151. 18:53

    You have to visit a thousand, blah, blah, blah, blah, blah. And then present it back and see if this is the difference to your query. Well, like is the query plan correct?

  152. 19:00

    Like is there a bug in MySQL? Do we have bad disks? Like, what's the discrepancy here? And I just got caught with that bug. And so after I left Shopify, I was just writing a lot of articles about this.

  153. 19:09

    I was just like, "Well, how long does, should this query take?" And then I... One hypothesis I had at some point is like, okay, well, how many writes per second can MySQL do?

  154. 19:18

    Well, shouldn't the amount of writes per second that MySQL do equal the amount of fsyncs that you can do per second? That sort of makes sense, right? Every time you do a write, you fsync to persist to disk.

  155. 19:28

    So how many fsyncs could you do per second? Well, an fsync takes one millisecond, so you do 1,000 writes per second. That, well, that doesn't really match up. Like, feel like a database can do more than 1,000 writes per second.

  156. 19:37

    Why can it do that? So that was one of those things where I tested and it's like, okay, well, MySQL on a little dinky box could do 10,000 writes per second.

  157. 19:44

    Well, how is that possible?

  158. 19:45

    Mm-hmm.

  159. 19:46

    And now you would just ask-

  160. 19:48

    How, how is it possible?

  161. 19:49

    Because you batch. So an fsync happens on usually a 4K-

  162. 19:53

    Yeah.

  163. 19:54

    Right? But it's like, that's not intuitive. Like it's actually... I, I, I got caught. It was like, just like, you know, probably some like 24-hour period where I just got obsessed with this question, where it's like you're writing like the BPF traces and all of that to do all of this.

  164. 20:05

    This is like pre-LLM, so it took forever. And you... And then I found out that, oh, every fsync was like much larger than I would've inferred.

  165. 20:14

    Yeah.

  166. 20:14

    Like, oh, it's batching. You go into the code and you read it, and then you found some obscure article by... It's always someone in like a central [REDACTED:origin] town that's like written some article about like how some intricacy of MySQL works and a patch that they did to...

  167. 20:28

    It's like it, the entire internet runs on small towns in Bavaria, I'm convinced. Yeah. [laughs]

  168. 20:36

    And then you decided to start [REDACTED:username].

  169. 20:40

    Yeah.

  170. 20:40

    Did... How, how did you decide? Did you know what you wanted to build or was it more like, "I wanna build something, something databases"? 'Cause you were clear very into databases.

  171. 20:48

    You, you'd done a, an awesome job benchmarking, like what is the theoretical like limits. You were very familiar with this. Probably became, you know, like world expert in, in this niche.

  172. 20:59

    And then how-

  173. 21:02

    I-

  174. 21:02

    Did, did you want to go into databases again?

  175. 21:05

    I think it was... There's three things that sorta came to a head. Um, the last project that I worked on at Shopify was Search, and I didn't have a good time.

  176. 21:17

    What, what, what did you use back there?

  177. 21:18

    Um, I don't... We don't need to name names-

  178. 21:21

    Okay

  179. 21:21

    ... of, of other database companies, but it was a, it was one of the, one of the like traditional search companies that a lot of different, um, companies run.

  180. 21:29

    And it was just very difficult to get it to do what I did, and I was just like, the projects that touched that database just... I couldn't get them to perform at the napkin math.

  181. 21:39

    And like there's, there's no query planner, and like I couldn't figure out why it wasn't there. And sometimes it tracked, and then sometimes it really didn't track at all.

  182. 21:46

    And so I tried to learn as much as I could to figure out and like start reading the source code of it, and I was just... I couldn't get it to track very often.

  183. 21:52

    It was very difficult to operate. And so I just... That was sort of like in the back of my head. I never thought I would touch that again. Then the second ingredient was the Napkin Math project,

  184. 22:01

    because it sort of just gave me a lot of facility with all of these Napkin Math numbers of what might be achievable with the machine if you utilized it perfectly.

  185. 22:10

    Properly, yeah.

  186. 22:11

    And then the third one was that doing this, you know, leaving Shopify in '21, having spent eight years there, and during that time I did this thing, I called it Angel Engineering.

  187. 22:20

    So I was like, joined my friends' companies and then I just vested equity instead of, um, instead of just investing or something like that. And 'cause I wanted to have my fingers in it.

  188. 22:29

    I wanted to like see what else was out there. That's why I left. And this problem kept coming up again, again, and then again and again, right? Like ChatGPT came out in 2022, and I was working with, with a company then, and they wanted to connect a bunch of documents to AI, and that's when the context windows

  189. 22:45

    were really small, so you had to read up-

  190. 22:46

    Yeah

  191. 22:46

    ... a search very quickly.

  192. 22:48

    So it's like a few kilobytes.

  193. 22:49

    It was eight kilobytes or four kilobytes-

  194. 22:51

    Yeah

  195. 22:51

    ... depending on the model. It was very, very small, so you had to reach for Search very quickly, right?

  196. 22:55

    Yeah.

  197. 22:55

    And so I, I, I worked with them and I was, I was... I created a little recommendation engine, and the recommendation engine was actually quite good. Um, like I saw, I found out that one of the co-founder's wife was pregnant through the recommendations that I was getting when I was running it on his feed.

  198. 23:11

    Um, [laughs] like it was, it was, it, it-

  199. 23:14

    Weird, but yeah. [laughs]

  200. 23:14

    It was recommending. Yeah, I mean, it was just like, you know, he was reading about like... I, I did get permission. I just like-

  201. 23:21

    Yeah

  202. 23:21

    ... I don't think anyone expected it to be good enough. And just like, okay, it's, this, this thing is working. And then I ran the back of the envelope math on what it would cost to do this for everyone, like all the users.

  203. 23:32

    This is a company called Readwise, so it's like articles that you save and then, and then search later. And it was gonna cost 30 grand a month, and this was a company, it's a bootstrap Canadian company.

  204. 23:42

    They spend about five... They, at the time they were spending about 5K a month on all the other infrastructure combined. So it just... It didn't d- You know, fundamentally in a company if you're doing an investment, you have to, have to earn some gross margin on top of whatever you're paying, right?

  205. 23:55

    And it just didn't line up. Um, and so we just didn't ship it, and I worked on... I, you know, tuned to autovacuum on Postgres or something like that, which is a good pastime.

  206. 24:05

    And then you... I just couldn't stop thinking about why it was so expensive to store all of these vectors that we were using for the recommendations. And I just sat and did the Napkin Math one day of like, can we just use it all to, in S3 and do some clustering, and then organize the files in just

  207. 24:20

    the way? And it's like, maybe you could build that and then-

  208. 24:25

    One day I just kinda said, "Fuck it," and did it. And, like, sat down and started to, like, to write it out. Um, and I spent the summer of, of '23 just hammering my head against the wall trying to find an approach where I could get the latency that I wanted.

  209. 24:41

    Um-

  210. 24:41

    'Cause the problem with S3 is it has really good durability, but latency we're talking hundreds of milliseconds, right?

  211. 24:46

    Yes. The P99 on a, uh, 256 or 512 kilobyte object on S3, um, is around 200 milliseconds. Um-

  212. 24:56

    And, and you're, you're saying P99 'cause, like, when you're talking large scale, you wanna care about the P99, right?

  213. 25:01

    Yeah, I think when you're des-

  214. 25:01

    That's why we're not talking about P50.

  215. 25:03

    When you're designing a system, you wanna optimize for the P99, and especially because when you're designing a system on, on S3, generally in every round trip you're not doing one request.

  216. 25:12

    You're often doing lots of requests, right?

  217. 25:14

    You're gonna hit the P99 real quick.

  218. 25:15

    Exactly. So it's like if you're navigating a tree on S3, right, it's like, okay, you get the upper layer of the tree, 200 milliseconds. You get, like, another layer of the tree, 200 milliseconds.

  219. 25:23

    You get a bunch of leaves of the tree, and it's 200 milliseconds. So in aggregate you have, like... You wanna look at the P99, probably even the P999, to design the system properly 'cause you will need to minimize the number of round trips that you have to make.

  220. 25:35

    So I just sat and sketched that out, um, and tried a bunch of different approaches, and then, and then finally in, in, in July of '23 I, I, I got something end to end that seemed to work, and then rewrote it probably twice and then released it in, in October of, of '23 based on, um, based on

  221. 25:53

    just that, that summer of, of, of working through it.

  222. 25:56

    And then you kinda, you built it on, on top of a S3, 'cause I guess durability and, and all of it, and just really good. You... How did you make it fast?

  223. 26:05

    We didn't in the beginning.

  224. 26:06

    Okay.

  225. 26:06

    Or I didn't in the beginning. Um, it was just me at the time.

  226. 26:08

    That came later, yeah.

  227. 26:09

    And it was really... Like, it was, it was, it was a project. It was not a company. It was not... It, it was, it w- it was to satisfy a curiosity.

  228. 26:18

    It was not... I did not set out to do this as like, "I'm gonna go, like, raise $10 million and do..." Like, I was, like... I barely knew what a VC was.

  229. 26:26

    Like, I was like, I just had to do this thing, and I was so focused on doing it, and it was so clear to me that if I wasn't gonna do it someone else was gonna do it, and I just became fully obsessed that summer with it.

  230. 26:38

    And so the first version was the simplest possible thing. I think... I'm a very pragmatic person. Like, I, I didn't get buried... I barely read any in- like, of the literature on LSM.

  231. 26:51

    I sort of like, you know, read a bunch of it, just, like, got the b-basic idea. Barely implemented that because that would've taken too much time. It was the simplest possible version of what it could be.

  232. 26:59

    Like, really what you have to imagine is that the simplest way you could do this is you run some clustering algorithm on the vectors.

  233. 27:06

    Mm-hmm.

  234. 27:07

    You get the clusters-

  235. 27:08

    Yep

  236. 27:08

    ... and then you put the clusters in files. The clu- the files are called cluster one, cluster two, [laughs] cluster three.

  237. 27:12

    Yep.

  238. 27:13

    And then you have another file called centroids of the clusters, and then you do the search by downloading centroids, looking at the centroids, and then downloading the n closest clusters.

  239. 27:23

    There was a few optimizations around merging some clusters that were adjacent in files and so on, just to, like, control some costs and some performance, but that was basically it, and then getting that to scale.

  240. 27:33

    That was the first version. And then how do we make it fast? Well, I didn't even implement a caching layer. I just put the reverse proxy in front of S3 with nginx and then had a cache behind it.

  241. 27:43

    Do you, do you know what a reverse proxy is? [laughs]

  242. 27:44

    I, like... I do know what it is. [laughs] I just still don't know what, what the reverse is about. [laughs] But anyway, um, the reverse, the reverse proxy, reverse things, um, the, the performance in this case, um, maybe that's what it's about, by caching, right, all of the, all of the S3 objects.

  243. 28:03

    It was... Again, it was the simplest.

  244. 28:04

    Yeah.

  245. 28:04

    Like, it's like I'm just gonna put that in front. I knew how to configure nginx. Like, I've written more nginx Lua than, um, than... A lot of [laughs] nginx Lua.

  246. 28:14

    Very good software. Um, just had that cache in front. And then the way that I would do things like deleting in the cache was just, like, shell out to Xargs and just remove, like, things in the, in the cache and reverse engineer the directory structure on nginx.

  247. 28:25

    And that's what we shipped, and it was just running on a single server in a Tmux instance, and I was like, "Okay, let's see if anyone gives a shit."

  248. 28:31

    Yeah, so so far, I mean, this is kind of like cool engineering and, like, a cool side project and, like, a bunch of novel ideas, and I, you know, like, I think just some hardcore engineering.

  249. 28:40

    How did Cursor come into play? 'Cause, like, when I learned about TurboBuffer, I was talking with Cursor about, like, how they s- s- built their, their backend, their database, how they scaled, and they were telling me all these migrations, and they were telling me like, "Oh, yeah, so we, we're on Postgres," but it didn't r- No, they

  250. 28:58

    did something else in Postgres. It didn't really work that well. They went to AWS Aurora, which is AWS's managed service of Postgres, and it didn't work well, which was very surprising.

  251. 29:05

    And they were like, "Oh, yeah, and then we went to this thing called TurboBuffer," and they worked well. And I was like, "What's TurboBuffer?" And they're like, "Oh, yeah, TurboBuffer."

  252. 29:13

    I think, I think they said, like, "We were one of their first customers." And this never computed to me. Cursor was already massive at that point.

  253. 29:19

    Yeah.

  254. 29:19

    How did you meet the folks, and how did they become... Was, were they the first customer? One of the first?

  255. 29:26

    They were the first customer.

  256. 29:27

    The first.

  257. 29:27

    The first.

  258. 29:28

    No.

  259. 29:29

    Um, they, they, they reached out, um, after... I just launched on, on Twitter. I was like, "Hey, I built this thing." And frankly, it was like,

  260. 29:39

    in, it... I exa- Like, you... It was like, "Hey, launch this thing," and to me it was like, "I am so sick of working on this." [laughs] Like, I was like, "I've been working on this all summer.

  261. 29:48

    I don't know if anyone cares. I only wanna work on this if anyone cares. Let's put it on Twitter." Again, single Tmux instance on a se- eight-core node somewhere in GCP.

  262. 29:56

    I was like, "If someone goes to prod, I'll, I'll set it up properly on multiple, and, like, I'll just block on that. But let, let's see if anyone cares."

  263. 30:02

    It was like the MVP of MVP. Anyone who's actually worked in the internal on databases would never have had... Like, would have had too much pride to ship anything like that.

  264. 30:14

    Um, and I've just... You know, I've worked on... I was just releasing it like a SaaS project. Why can't you work on a database like it's SaaS? I don't...

  265. 30:20

    Do you know?

  266. 30:20

    Yeah.

  267. 30:20

    It's like, if anyone uses it, we'll do it properly. I know how to run software with a lot of nines. Um, but it was not a proper LSM. Like, it was very, very...

  268. 30:28

    It was the simplest version of what it could be. And then- I released it on Twitter, and I was like, "Yeah, you can do a million vectors for a dollar," and before that, I think the, the cheapest was maybe $100 per million-

  269. 30:38

    Yeah

  270. 30:38

    ... for something that actually worked.

  271. 30:39

    Yeah.

  272. 30:39

    Um, and I knew it was reliable, right? I knew... Like, I had these invariants, like, if you shut down all the VMs, like, no data is lost, like, all the writes are committed directly to objects.

  273. 30:48

    Like, it has all the same invariants it had today. Um, and Cursor reached out, and knowing them now, I'm sure... At the time, the Cursor was maybe eight people, and knowing the founders now, I am sure that they had sat at the dinner table one day and were like, "The unit economics of what we have right now,

  274. 31:04

    where all the vectors are in DRAM, are not working. Why hasn't anyone built it where we can put it in S3, and the actual code bases that are actively being used we can put in memory, and everything else just sit in object storage, and then we just hot load it in and out of the cache?"

  275. 31:17

    Yeah.

  276. 31:18

    Makes so much sense, right? You open the code base, a few seconds, and it's in RAM, and then the queries are as fast as on-

  277. 31:22

    Yep

  278. 31:22

    ... anything else. It made so much sense. So, I mean, at the time, they were... If you look at some of Aman, one of the co-founders, early tweets, he talks about, uh, using S3 for KV caching and things-

  279. 31:31

    Yep

  280. 31:31

    ... like that, which barely anyone is still doing even though, um, the economics-

  281. 31:36

    Very, like, economical, yeah, price-wise.

  282. 31:37

    It's... Yeah, it's, and it's, it's very, it's very uncommon, and I think it will happen, right? But they were ahead of their time.

  283. 31:42

    Mm.

  284. 31:42

    And they, I think they were... I don't know if they were thi- thinking of building it themselves. I think that's quite likely. Um, and they found [REDACTED:username], and it just perfectly pattern-matched into that.

  285. 31:52

    Again, I don't know if this dinner conversation happened or if this was just inside-

  286. 31:56

    We'll have to ask them now

  287. 31:56

    ... Arvind's head. Um, but it pattern-matched something, and so we exchanged a bunch of emails, and then something compelled... I didn't know anything about B2B sales. Now I love B2B sales. [laughs]

  288. 32:07

    Um, I didn't know anything. I was just like, "I just wanna help them," 'cause they, they were, they had some unit economics that didn't line up. So I just went to San Francisco, right?

  289. 32:15

    I live in Canada. I went to San Francisco, and I showed up at the office, and when I showed up at the office, they, um, they were having some Postgres problem that they were discussing.

  290. 32:23

    Yeah, the AWS Aurora problems, yes. [laughs]

  291. 32:25

    Yeah, early on, and I was like, "Oh, do you guys have pganalyze?" And they said, "Oh, no, we don't." I said, "Okay, let's, let's, let's get that going, right?

  292. 32:33

    Let's look at it." And it was the same thing as it always is with Postgres, which is auto vacuum hadn't run enough, and so they had all of these, like, going to heap when they should be doing index scans and blah, blah, blah.

  293. 32:43

    So we were talking about all of that. And so I was just helping them, right? It was like my, you know, my database genes just, like, kicked in. And I think this built enough trust with them that, okay, well, maybe if he knows how to help us with the database, maybe he also would know how to build

  294. 32:58

    one. And, um, at this time, I'd also approached who I thought was the best engineer who ever worked at Shopify, my co-founder Justine, um, and she'd come on, and the first thing that she did was, um, remove the reverse proxy nginx cache with a, a file-based cache. [laughs]

  295. 33:17

    Like, just a direct cache, which again, great. Like, the S3 thing worked. Um, and so she was online. She was starting to work on it, and, um, and Cursor c- Cursor, Cursor then that night was like, "Okay, well, we're gonna migrate."

  296. 33:28

    And so they migrated everything over the course of, like, a week or two after that. Um, but Cursor was a small company back then, right?

  297. 33:35

    Yeah, and they, they were, they were just in the beginning of their massive rapid growth.

  298. 33:40

    Exactly. And I, I told them that I was gonna reduce their bill by 95%, and I did. Like, we did. Justine and I did. We... Like, they came on, and their last bill with their previous vendor and the first bill with us, it was 95% lower.

  299. 33:56

    Yeah, and you're, you're, you're nice for not saying vendors, but I, I can say vendors 'cause I've talked to them, and, and it's in the deep dive about Cursor.

  300. 34:01

    It was, it was A- AWS Aurora specifically. Uh, so and-

  301. 34:04

    This was, this was not, this was not Postgres, no.

  302. 34:07

    Uh-

  303. 34:07

    This was, uh, this was, uh-

  304. 34:07

    Oh, it, it wa- it was a different one, but it's probably still in the writeup where we don't, we don't want to name names. But, uh, yeah, but they were...

  305. 34:13

    The reason they went there is reliability was their main, main, main pain point. I'm sure the unit economics w- would've been there. But yeah, this was... And then what Swalit told me is, he said, like, "Look," like, "there's a few things that we did that you should never, ever do."

  306. 34:24

    And he said, "One of them, you should never, ever bet your business on a tiny startup where you are their only or biggest customer except for [REDACTED:username]." And he said, "I love, love those guys."

  307. 34:33

    So I guess it just comes to show that even in your case, like, to me, what the story is, shows is, is you can do things, and wh- when you build high-quality things and you're pushing for things, good things can happen.

  308. 34:44

    And on the other side of Cursor, when you're a startup, it's okay to take sometimes irrational risks when you have conviction. And, uh, it sounds to me that you gave them conviction by showing up in person, by helping them, by showing that, you know, you know your stuff.

  309. 34:58

    Like, you suddenly brought in your, your 10-ish or eight years of Shopify experience and your curiosity, and they probably took a risk because of that, not because you were some, you know, random vendor.

  310. 35:07

    They probably would have never done that. So fa- fast-forward to today. Uh, [REDACTED:username] is now a lot bigger. You're, you're working on some s- some cool things, but

  311. 35:17

    you have this very interesting business where, for you, CPUs are important, right? You run on mostly CPUs. And you told me a story over dinner yesterday that, uh, you met Jensen, uh-

  312. 35:31

    Yeah

  313. 35:31

    ... and Jensen, he really wanted to s- sell you on GPUs. Can you tell me how that meeting went?

  314. 35:36

    Um, yeah. [laughs]

  315. 35:37

    Jensen Huang, right?

  316. 35:38

    Yeah. I just... I, I'd never met, uh, I'd never met, uh, Jensen before. We were, we were at an event at, uh, at, at NVIDIA, and we were just doing, um, presentations via-

  317. 35:48

    Was it in a big, big HQ, super impressive?

  318. 35:50

    Yeah, exactly. They'd invited a couple companies to go and, and, um, and, and, and talk about, um, [clicks tongue] uh, talk about our businesses and how we can partner with NVIDIA and so on.

  319. 35:59

    And I, I don't j- I, I don't know. I was like... I think I was in a goofy mood that day, and so I went up on, on stage, and [laughs] I said, um, "Hey, I'm Simon from, from, from [REDACTED:username], and, uh, and yeah, if you're wondering about the name, it's, like, if everything goes south, we can always

  320. 36:17

    pivot into vapes." [laughs] [laughs] I was kinda nervous. And s- this is what I, this is what I said. And then, and then he said back to me-

  321. 36:27

    And what, wait, who was in the room? Was it Jensen? Was it a direct report?

  322. 36:29

    It was, it was Jensen, and then, I don't... I... He has, like, f- I don't know if it's just 50 d- direct reports or it was, like, you know, it was, there was-

  323. 36:36

    It was no pressure

  324. 36:37

    ... it was Jensen and then a bunch of the, um, like, NVIDIA, NVIDIA leadership, right? Um, 'cause you go there, and then you talk about that and you find opportunities to partner and work together, right?

  325. 36:45

    And- S- so I'd said, "Yeah," I said, "You know, so plan B could be that we could pivot into vapes." [laughs] And then he said... I was already nervous. He said,

  326. 36:57

    "Judging by your slide, maybe you should." [laughs] [laughs]

  327. 37:02

    No, he did not. [laughs]

  328. 37:07

    And, and then I didn't know what to say back to that, so I said,

  329. 37:14

    "Well, Jensen, do you vape?" [laughs] [laughs] And he didn't, he didn't answer the question. [laughs]

  330. 37:26

    And then someone, um, someone on the, um-

  331. 37:29

    Oh, God

  332. 37:30

    ... someone on the, on the team, um- [sighs]

  333. 37:34

    ... wrote to the whole company, [REDACTED:username] company, "Simon just asked Jensen if he vapes." [laughs]

  334. 37:39

    Um, and then, you know, this is so- this is a great start, right? And, um, and then the team had s- team, team had sort of talked to me beforehand and was like, "Simon, we gotta make sure we don't say the C word.

  335. 37:50

    We can't say CPUs." [laughs] And so I just couldn't stop talking about CPUs. [laughs] I was like, "AVX-512 is so sick, like we love SIMD, and, um, s- like we, we l- we l- like there are so many CPUs, they're so easy to get, like, um, it's just a riot in CPU land."

  336. 38:10

    Like, you know, I've, I don't... I think I stopped short of saying, "I'm so glad I don't need GPUs." [laughs] But, [laughs]

  337. 38:17

    but, but it was just, it, I just couldn't stop talking about CPUs. Yeah. And so, you know, Jensen took an interest in that. Yeah. [laughs]

  338. 38:26

    So he- who, who knows? Like I'm, I'm, I'm sure you made, you made a memorable impression. Maybe he made it his mission now to like at some point get you guys onto GPUs.

  339. 38:33

    But speaking of CPUs, can you tell me a bit what you're seeing inside of the hyperscales of cloud providers? A- you're, you're now on AWS, you're, you're in GCP, you're on Azure.

  340. 38:42

    What I would think naively is there's a GPU shortage, and when I talk with inference companies, they are, and, and, and, o- AI labs, they're just getting whatever they can do.

  341. 38:50

    I would think getting CPUs is, should be easy. Is it?

  342. 38:54

    No. [laughs] It's not anymore.

  343. 38:57

    Why? What, what's happening? Can, can you tell us about the dynamics on, on, on, on the why and what you've learned?

  344. 39:01

    Yeah. So I think that GPUs will probably continue to be scarce. Like I don't know, maybe there's gonna be some surplus. I, I refuse to speculate too much about the macro.

  345. 39:11

    But I think as, as RL is becoming a very, very large amount of the workloads, that needs a lot of CPUs.

  346. 39:20

    Mm.

  347. 39:20

    So the labs are sucking up a lot of CPUs 'cause you need CPUs to be like, "Okay, we need to like teach this model how to, how to search.

  348. 39:26

    We need to teach it how to use Grep."

  349. 39:28

    Mm.

  350. 39:28

    "We need to teach it how to boot up Bash. We need the..." It needs to run real things and learn from that.

  351. 39:34

    Mm-hmm.

  352. 39:34

    It takes a lot of CPU.

  353. 39:35

    Mm-hmm.

  354. 39:36

    Um, and so I think as we, RL is consuming a lot of CPU, and then also just all of the agents are running on CPUs, right? They need to do all kinds of very general purpose things on a CPU.

  355. 39:45

    And so as, as, as the demand curve is sort of shifting to the right and it's becoming more and more applied, and that feeds back into RL, by the way, right?

  356. 39:53

    Because as things become more applied, it's like, oh, the models are not that good at CAD or shipbuilding. I don't know. And then, you know, you have to spin up even more RL environments to do that.

  357. 40:02

    So I think that's what we're seeing, and so we're on the other end of that, needing these CPUs. We need a lot of NVMe SSDs as well, um, and a lot of this right now is tied up in DRAM, right?

  358. 40:12

    Of, of where-

  359. 40:13

    Yeah

  360. 40:13

    ... like you need a lot of that also for the GPU servers. Um, but I would assume that it gets a lot worse before it gets a lot better on the, on the CPU side.

  361. 40:21

    Um, and I think even the big companies are fighting amongst each other, right, to get the allocations. And even we, you know, we're selling to companies that we also fight for CPU with and against, right?

  362. 40:32

    It's, uh, it's, it's, it's really difficult, and so you write things to try to make sure you get these CPUs as fa- fast as possible.

  363. 40:38

    Yeah, and yesterday I was at a dinner that you hosted with your team where you actually have a bunch of [REDACTED:username] customers. A bunch of them are AI, AI labs or, or AI startups, but a lot of them, one of them, uh, Reflection had, have huge, massive amount of footprint, and they were telling me that they're in

  364. 40:53

    a situation where they cannot buy more. Like they, uh, when it comes to GPUs or CPUs, they've maxed out. They have the longest contracts that are possible. And, uh, I didn't realize how competitive it is in the cloud when you're, you go beyond a small fish to like a medium size or e- even a-

  365. 41:09

    Yeah

  366. 41:09

    ... a large fish, that now, like, uh, it's, it's interesting. So, so now you have this e- and even you're having this, this, uh, kind of fight behind the scenes that is maybe not as visible.

  367. 41:19

    Exactly. And I mean, you, you work with the clouds, right? You work with them to talk about which regions have, um, have CPU, which regions are getting... It, it comes down to power, right?

  368. 41:29

    Of like, okay, well, where is the power, which is generally where they're gonna ship the new CPUs. Um, and so we have to work with some of our biggest customers on that.

  369. 41:36

    So these are real constraints, right, that are, that are making our way to us. We're just very fortunate that it's very easy for us to run lots of [REDACTED:username] clusters 'cause all we need are a, like a few CPUs and NVMe SSDs and an S3, and then we're in a good place.

  370. 41:51

    But there's lots of changes that we can make even to the architecture, um, to try to protect from, from a lot of this. Now, I'd rather spend that engineering effort on other things.

  371. 41:59

    Yeah.

  372. 41:59

    But we are very, very good at using a lot of very different SKUs, right? So we don't need everything to be a particular CPU or instance type. We can run with many even, many different types of machine types, um, on, on different-

  373. 42:11

    And, and SKU meaning that's the, it's, it's a fancy name for like the different machine types.

  374. 42:14

    Yes, exactly. Right? Like, you know, C4d or i8g or whatever they're called on the-

  375. 42:19

    What's your favorite one?

  376. 42:21

    Um, we really like right now the, um, C4s on, on GCP.

  377. 42:28

    GCP, yeah.

  378. 42:28

    Um, the Z4ds are also performing really well, um, now that we've done some, done a bunch of, of, um, of optimizations to them. Um, those are really, really great machine types.

  379. 42:39

    Uh, we really like those. Um, and then the Arm C4as as well, um, on, on GCP, um, th- we like those. But I think that in general, like when you're, yeah, when you're small, it's very easy to suck up a bunch of...

  380. 42:51

    But, uh, at, at Shopify, I was also part of, you know, deciding- Be- ahead of BFCM, right, a few months out, you have to tell the cloud providers how much you're intending to use.

  381. 43:01

    Do commits on all of that, right? The c- the clouds are not infinite as they seem when you're small.

  382. 43:06

    And one way, of course, to, like, get, like, infrastructure and, and, uh, also just, like, credibility is venture capital. If you raise $100 million, a billion dollars, some of your customers just raised $2 billion.

  383. 43:16

    Actually, I talked with them yesterday. You know, it gives you credibility, it gives you cash, you can pay for this thing. Your specific, [REDACTED:username]'s relationship to venture capital seems very interesting.

  384. 43:26

    I never heard you announce a raise until may- maybe just very recently. Can you tell me how you th- and, and you told me that when you started this thing you didn't think too much outside of just building some cool stuff.

  385. 43:37

    How did you think about venture capital, and how do you think about raising? 'Cause again, I feel you have a very fresh and different perspective than which, which is typical inside of Silicon Valley.

  386. 43:48

    Yeah, so I think to, to understand my, how I think about capital, you have to go back to the, the beginning of [REDACTED:username], right? Where I'd promised Cursor that Justine and I could get their bill to 4K a month.

  387. 44:03

    And this was based on some very rough napkin math on, okay,

  388. 44:07

    if, if [REDACTED:username] was a better implementation than it currently is, then it should cost this much [laughs]. And that's the pricing we ship with, and that's what we guaranteed, um, guaranteed Cursor.

  389. 44:18

    Um, but the software was not that good. [laughs] Like, it was very reliable, but it was very simple, right? And that's, like, a core engineering principle of me, is simplicity above everything.

  390. 44:29

    Um, you and I have talked before about how software that ages well and some of the advantages of seeing, be- having long tenures inside of companies. You had a long tenure at Uber, I had a long tenure at Shopify.

  391. 44:39

    So, you see simplicity just almost always wins. Um, and at the time, I was not convinced whether this was a venture scale opportunity.

  392. 44:49

    Mm.

  393. 44:49

    Because I understood that if you take venture capital, no matter how many smiles there are in the room, everyone's sort of expecting that you have to earn a big return on that on some timeline that makes sense to everyone involved.

  394. 45:01

    And everyone involved are, you know, pension funds in Canada, like, that, it's like, it, there's like a whole stack, right-

  395. 45:08

    Yeah

  396. 45:08

    ... of, of, of people that, that need to... So, at the time I was like, "I don't, you know, I don't know if this could be a billion-dollar company."

  397. 45:13

    I didn't know that in the very, very beginning. Um, it wasn't completely clear to me. It felt like a very niche kind of product, right, to build this particular search engine.

  398. 45:22

    Um, and that was completely fine with me. So, I, you know, it's, it's, it was fine. And

  399. 45:29

    so then I just, I just looked at the Cursor bill, and I looked at my GCP bill, which is what we started on, and, you know, as like a, you know, dumb [REDACTED:origin] person who's just like, "Okay, like, this number should just be lower than the other number."

  400. 45:42

    Yeah.

  401. 45:43

    That's sort of like, you know, and it's just, I don't think I'd spend enough time in San Francisco, 'cause I think the money over here, it works a little bit differently. [laughs]

  402. 45:51

    Um, that's just, that's all I knew.

  403. 45:54

    You, you were doing Business 101. As, as long as you're making a profit, you're good, right?

  404. 45:59

    Yeah. That's [laughs] it's like I'm, I'm, I'm not kidding in this exaggeration, that it was just like, that just made sense to me. That Justine and I were just gonna go optimize this until these r- numbers were roughly equal.

  405. 46:11

    A- and maybe if, if, if we could get some other workloads, we could start paying ourselves, but that, that was, like, very much the philosophy at the time.

  406. 46:18

    Mm.

  407. 46:18

    Um, because I didn't know if I could go raise a bunch of, of, of money. I didn't know anyone who had the money. I, I didn't have any relationships.

  408. 46:25

    Um-

  409. 46:26

    You were an absolute outsider to the, the world.

  410. 46:27

    I was, I was an outsider. I was like an outsider squared, right? I grew up in Aarhus, Denmark, and I, um, I then moved to Ottawa, Canada. So it's like I'm an outsider to Canada, and in Canada I'm an outsider to San Francisco.

  411. 46:41

    So I was just thinking about this from first principles. Like, oh, you're a venture capital, you need this return, you need it on this timeline.

  412. 46:49

    I don't know if I can deliver that yet. I would need more data to decide that, because I wanna, like, I kinda wanna keep working on this, and now I have to get to this point for it to not be a failure.

  413. 47:00

    Um, in, in, in January then I, uh, there was a person that I was at IOI with in, in, uh, in 2012 and 2013, and his name is Boyan, and he was on the Northern [REDACTED:origin] team, um, at IOI.

  414. 47:15

    Um, and he was, he's, he was really good. He was so good that the [REDACTED:origin] [REDACTED:origin] team called him God. Um,

  415. 47:22

    I don't know why, but that was what he went by. And he was, yeah, he was very good, grew up... And, and I really wanted to work with Boyan, but I couldn't afford to work with Boyan. [laughs]

  416. 47:32

    Um, and he was very much like, "This is what I can live off." Like, you know-

  417. 47:36

    Yeah

  418. 47:36

    ... I just, like, I wanna build this data- like, that would be, like, this is what it can be. But at this point, Justine and I hadn't taken a salary for, like, six months, and we'd already, we'd already spent, like, tens of thousands of dollars on, like, on GCP bills and all of that.

  419. 47:51

    And so I was like, "I don't think we can, I don't think we can, we c- we could do it." And so I had met one, one individual in, in, in Silicon Valley, uh, his name is Laki, and it just, I ended up just calling him and saying, "Hey,

  420. 48:03

    I kinda wanna learn a little bit faster here. Can I,

  421. 48:08

    can we raise, like, s- 700K?" That's, like, what I wanted to raise. So it's just like, "I wanna have, like, two engineers for the rest of the year. Justine and I still don't need to be paid, and then a little bit of buffer room."

  422. 48:20

    It's just like, "This is what I need. And if this doesn't have PMF and is a big opportunity by the end of the year, I don't think we're gonna bother, and we'll just shut the whole thing down, and we won't have taken a dime.

  423. 48:28

    We'll return everything to you." Um, I think that was the first time you heard anyone say it like that. Um, and, um, I told some other VCs that at the time, and that, that was terrifying to them.

  424. 48:39

    I think to someone on the West Coast, this sounds like you have low ambition or something like that.

  425. 48:45

    Mm.

  426. 48:45

    Um, and to me it was just like, I, I don't know, it just came from a... When I don't know how to play a game, I just play with open cards.

  427. 48:51

    Yeah.

  428. 48:51

    Like, this is how I see it. And so I, we were, it was very clear to us that we wanted to do this. And, but also it became clear to us that we didn't wanna just, like, keep working on this unless it could become big, and we were starting to develop conviction, conviction that this actually become really,

  429. 49:04

    really big And so we, we, we did that, and hired Boyan, and then became profitable later that year. Um, and then just continued to hire. And then it's like, to raise more money, you need sort of...

  430. 49:18

    There needs- there's six reasons to raise capital. The first reason to raise capital is to fund R&D.

  431. 49:23

    Mm-hmm.

  432. 49:24

    That was the reason that we raised capital in January, because we'd funded R&D with a lot of our own, you know, opportunity costs and not taking a salary, and then paying the bills ourselves.

  433. 49:33

    Um, but we wanted to learn a little bit faster, and so we hired Boyan and Morgan as the first engineers. And then the second reason to raise capital is to fund growth.

  434. 49:43

    You've, you, you've built something, and you want to tell the world about it, and you wanna spend more capital to do that. Um, the third reason to, to, to raise capital is for the founder's ego. [laughs]

  435. 49:57

    Um, it's a very popular-

  436. 49:58

    I appreciate the honesty.

  437. 49:58

    It's very popular. Very, very popular.

  438. 50:01

    Yep.

  439. 50:01

    Right? Big numbers, lots of press, like, um... And I think this is a very, very dangerous reason to raise money, and I wish that it was more talked about because you're diluting all of your employees when you do it.

  440. 50:14

    You are, um, setting a certain price for future employees and their upside. It's,

  441. 50:20

    it's, it- for some people, it can become a status game, and that's not what it's about. We're here to build a big business together, and

  442. 50:28

    this is not a reason to raise money. Um, but I, I do think that it happens. Um, the fourth reason to, to, to raise capital is to reward your employees, right?

  443. 50:37

    It's a, it-- You're on a very long journey, and you wanna work with the best people in the world, and by definition, there's not that many best people in the world.

  444. 50:45

    Yeah.

  445. 50:45

    So you wanna reward them. Um, that was the reason that we took more capital in December, um, was to allow the employees to liquidate, uh, some of their equity, um, instead of waiting for some, like, event, like an IPO or something, like, further out.

  446. 50:58

    Um, the fifth reason to raise is for a strategic partnership. There are strategic partnerships that have been made in this, in this city that have made companies. Um, and um, the sixth reason to raise would be do- s- doing M&A or, or something like that.

  447. 51:13

    But it's like you have to be very honest about what reason you are raising in those six. First reason we raised was one, and second reason we raised was four.

  448. 51:22

    Um-

  449. 51:22

    So which, which one's the first reason to raise was?

  450. 51:25

    R&D.

  451. 51:25

    R&D. And the second reason was?

  452. 51:27

    Um, to provide liquidity to the employees.

  453. 51:30

    The employees.

  454. 51:30

    Yep.

  455. 51:32

    I, I, I think it's a, it's a nice and healthy way, and I think, yeah, the, the e- ego part, we don't talk about, and the identity, and es- especially the, the closer you are to, to tech ecosystems where a lot of people are raising, it, it will be part of it.

  456. 51:45

    As closing, I, I wanted to ask you about the way you have a remote culture. These days, I'm seeing it, especially for companies that do anything with AI, may that be building AI infra or, or, or just AI products, a lot of them prefer in-person, having a HQ oftentimes in SF or wherever your headquarters is, may that

  457. 52:05

    be London or somewhere else, because you often f- the, these companies often find that they have faster iteration. Uh, it's just fewer layers cut in between and, of course, speed is, is very, very important.

  458. 52:15

    You have started full remote, and you're still full remote. How is it working, uh, and what kind of quirks or like, or at [REDACTED:username] ways have you found to, to make this work better?

  459. 52:28

    Yeah. I think... So the, the company started in, in '23, so sort of like on the, on the, on the cusp of COVID, where a lot of companies were just remote.

  460. 52:37

    Um, the Shopify infra chain was remote since the very, um, very beginning-

  461. 52:41

    Yeah

  462. 52:41

    ... 'cause it was very difficult to get them all to move to Ottawa. Um, and

  463. 52:47

    so it was natural to me. It was like, okay, I think there's kind of maybe two cities where you can build a database company fast, and that's San Francisco and, and, and maybe New York.

  464. 52:57

    There are maybe other cities-

  465. 52:58

    Yeah

  466. 52:58

    ... right? But that's like kinda where it's been done.

  467. 53:00

    Yeah.

  468. 53:00

    And so if you don't wanna do that, I think you have to go all in on, on, on some distributed model. And so we've tried to figure out what does that distributed model mean for [REDACTED:username]?

  469. 53:09

    It doesn't mean the absence of in-person. We get everyone together twice a year in, in some, in, in some location. Uh, earlier this year we were in, in Banff, right?

  470. 53:17

    And then we were in Mexico City and so on. So it's like that's, that's not that uncommon. Um, but one of the things that we're, we, we've been trying to do is we have this concept called campfires, and the concept of the campfire is that when a couple of people just sort of randomly congregate in a place,

  471. 53:32

    you call it a campfire, and you encourage as many people as you want to come and join. So for example, this week is a [REDACTED:username] campfire in San Francisco 'cause I'm here for this conference and a bunch of other things.

  472. 53:42

    And so everyone is invited to come. Like, we're gonna go meet customers, right? We're gonna put on dinners for our customers and things like that, and we just make a thing out of it and, and spend time together and, uh, we encourage everyone to come.

  473. 53:54

    We've also gone to the extent now of, um... We wanna encourage that, but not everyone, not everyone needs to go to the campfire all the time. Some people just wanna, you know, lock in and hacks into tent, and that's great.

  474. 54:06

    We have people that just make it to the offsites twice a year, and otherwise they're home, they're with their families, and they don't, they don't spend time on an airplane.

  475. 54:13

    Um, fantastic. Like, that is completely compatible with this model. And there are other people at the company who are on a plane probably every two weeks. Um, we had someone the other day where they saw a campfire happening in New York, and everyone was dialing in from a meeting room in New York, and she had so much

  476. 54:29

    FOMO that she took an Uber straight to the airport in Ottawa and flew to, [laughs] flew to New York to hang out with the team, right? And I think that's fantastic.

  477. 54:37

    Um, and we've also introduced these things where, um, if you, if you, uh, if you do a cur- conference talk or a blog post or something like that at [REDACTED:username], something a bit extracurricular, we give you a Turbo credit.

  478. 54:50

    And a Turbo credit allows you to upgrade your next flight to business class, which again encourages spending time together with the team. Um, and now, I mean, Turbo credits are probably gonna take on a life on their own.

  479. 55:03

    Someone was talking about doing a central bank and doing interest rates on the Turbo credits, um, and doing a betting market on the Turbo credits. And so, like, this might take on its life on its own.

  480. 55:12

    Um, and, uh, you, you know, if you, um, if you're at a conference like this, there's some of the, our engineers here who are just wanna interact with customers and be on, like...

  481. 55:22

    And standing on a, like, expo floor all day is quite taxing. And so if you do that for two days 'cause you wanna do it, oh, you get a Turbo credit, right?

  482. 55:30

    And so it's just, like, these fun little things that we try to do to, to, to encourage people to meet if they wanna meet.

  483. 55:36

    Thank you. Well, uh, uh, um, in this session, uh, what I found very interesting is [REDACTED:username] is a m- so many AI companies are using you as an infrastructure layer.

  484. 55:44

    But in this conversation, we managed to talk very little ab- about AI and a lot more about engineering principles, pushing, being curious, and the human connection, how important it is for people to work together, to trust each other.

  485. 55:57

    So just thank you very much for that. So let's give a big round of applause for Simon.

  486. 56:00

    Thank you so much.

  487. 56:02

    This was great. Thank you. [upbeat music]