← All AI Engineer talks

AI Engineer World's Fair 2026

State of Data

About this talk

Sean Cai examines how AI training-data markets are shifting from static annotations and final-state records toward authentic enterprise workflow and process data. He distinguishes Type 1 from Type 2 data, links application maturity to task verifiability, and argues that benchmark gaming and evaluation-harness differences obscure real model capability. Finance-task examples illustrate the need for robust rubrics, deterministic verification, and durable enterprise data infrastructure.

Chapters

  1. 0:00Inside AI training-data markets
  2. 2:38Type 1 and Type 2 data, compute, and process capture
  3. 5:52Verifiability, GitHub, and enterprise workflow signals
  4. 8:29Benchmark incentives and harness-dependent evaluations
  5. 10:14Finance-task benchmarks and rubric analysis
  6. 14:00Enterprise infrastructure and closing remarks

Talk transcript

  1. 0:00

    [on-hold jingle] Good to, uh, give this talk, and, uh, I just wanna say right off the bat, this is probably gonna be a little different from what you've seen so far at AI Conference.

  2. 0:20

    I'm here to not deliver an agenda on any company's behalf, but just to expose a lot of alpha in data markets. Um, but also just tell you what's really going on behind the scenes, um, in a, in a very murky landscape where nobody seems to know how Mercora, Handshake, and a lot of these folks actually produce data.

  3. 0:38

    So, you know, q-quick reframe before we start. When people hear data markets, they picture Scale around twenty nineteen. They picture these rooms of annotators in like Manila labeling images, and that's, that's real, and it's maybe, you know, ten to fifteen billion dollars a year per lab, but it's kind of the least interesting part.

  4. 0:55

    Um, the models work now. What's scarce and badly priced is the sort of data that takes them from generalist competence into real expertise. We knew this since, uh, twenty twenty-four when Scale got acquired, and we were, um, spamming a lot of GPQA datasets.

  5. 1:09

    So look, um, let me start with this framing. Data is to the white-collar revolution what coal and iron basically is to the Victorian age and the information age right now.

  6. 1:20

    Um, I have this piece called The TAM isn't a vertical. It's all of labor. And so the supply chain is basically just doing what every sort of industrializing supply chain does.

  7. 1:29

    It unbundles. Two years ago, one vertically integrated giant like a Mercora or a Surge or a Scale AI, if you're guys-- if you guys are unfamiliar, they're massive data companies.

  8. 1:39

    They had to do all of it, but that was because that's the only way that unit economics worked in an immature market. Today, um, that's increasingly not the case.

  9. 1:48

    Specialists out-compete the giants at a lot of steps, uh, sourcing the people, building environments, uh, designing rewards, running evals. The fragmentation, I would say, is pretty permanent. It's not transitional.

  10. 2:01

    And, uh, quality increasingly does not scale linearly with quantity, which leads to a sort of cottage industry in data land right now, where you have labs literally mandate, uh, vendor diversification on the scale of like twenty to thirty different vendors because they inherently distrust their ability to scale quality with quantity.

  11. 2:19

    So to orient you, you know, here's, here's the sort of argument as it developed this year. Um, January, we saw a lot of industrialization and unbundling. You'll hear me come back to this mechanism I call the Antikythera mechanisms, which is this sort of bespoke systems that translate messy business context into evals.

  12. 2:38

    Uh, increasingly important in labs' hungry quest for real-world seed data to set seed ends. Um, I also explored this concept called Type 1 versus Type 2 data, uh, contrived versus non-contrived for the more researcher types in the audience, and why the GPQA sell playbook that works in twenty twenty-four for data acquisition kind of falls apart on long-horizon,

  13. 2:57

    non-verifiable work. Um, so, you know, going through all these topics, all of which are online, uh, I'll als- I'll actually just skip to the more interesting part, but I'm bringing this up just here in case any of the particular topics I talk about are, are particularly interesting and you wanna dive in more.

  14. 3:14

    So model improvement is a function of three inputs. We all sort of see this, uh, commonly expressed as the compute data and talent. I put together this very, uh, like rudimentary sort of chart online in one of my pieces in a re- very rudimentary, uh, equation just to express the fact that, you know, if there's any sort

  15. 3:32

    of imbalance in this compute data and talent, then you start seeing, uh, sort of inefficiency in producing, uh, what I call like generalized AI model performance, right? But also, uh, we, we see a sort of inefficiency in, in CapEx spend that i-- that is a sort of result of this equation right now.

  16. 3:49

    Um, exponentially increasing CapEx spend, but sort of like, um, AI revenues are, are sort of falling far behind. Uh, data is sort of the underfunded leg here. It's the one that sort of turns a generalist model into a real expert.

  17. 4:02

    It's the one that's, uh, actually, I believe, quite lacking in this equation, and thus because of the imbalance penalty parameter concepts that I'm expressing, um, presents a whole opportunity.

  18. 4:15

    But, you know, going back to what I said at the start, what is data actually? So most people picture state-based data, the rows in an ERP, which is like the final output or a saved file.

  19. 4:26

    That's kind of the twenty twenty-three like next token prediction model, um, and it's mostly personal data wrapped in privacy law. What is actually available nowadays is process-based data, which is the trajectory, the reasoning tra-trace, the sequence of decisions.

  20. 4:41

    Um, so it-it's kind of what gets a professional to get from a blank page to a finished work output and sort of delineates how the work gets done. Um, so on top of that, it's a quality axis.

  21. 4:51

    This is the vocabulary I'll use all talk. Um, type, Type 1 data is a sort of pure capture of real workflows like GitHub commits or session replays with minimal reward shaping by non-experts.

  22. 5:02

    And Type 2 is contrived data where you sort of hire experts, you sit them in an arbitrary setting, you have them manufacture examples. Type 2, the right place to start models when they're reading at a first-grade level.

  23. 5:15

    I think anybody in the world could teach them, uh, as a third-grade teacher. But Type 1 is what gets you from twenty to eighty percent, so to say, 'cause the, the, the realism is inherited from the work itself.

  24. 5:25

    And the structural reason why it matters is the data is kind of the most depreciable asset there is. Um, a dataset is only available insofar as frontier data markets as the frontier moves.

  25. 5:35

    So the only durable supply of it, technically, is a live business you partner with, not a dead startup's code bases like so many data companies out there are buying today.

  26. 5:43

    The dirty secret of the industry, though, is that everybody sells, um, Type 2 and bills it as Type 1

  27. 5:52

    So now let me talk about the central axis of how you can think about the sort of verification and deciding which application layer domains are, are, like, maturing first and why.

  28. 6:01

    Uh, you know, there's a reason why Claude Science came out so far, um, you know, now in the future after models got good before a lot of other application layer advances.

  29. 6:11

    So Jason Wei is a researcher online whose blog is great. You should read him. He has this law, it's called Verifier's Law. The ease of training a model to do a task is proportional to how verifiable the task is.

  30. 6:22

    So I break verifiability into three axes here. And once you have them, a huge amount of this market stops being mysterious. First, asymmetry of verification. How hard is it to decompose a task into checkable steps?

  31. 6:33

    Veracity of verification. How much consensus is there about what correct even means? And then thirdly, proliferation of verification. How often does the real world sort of hand you fresh examples of verified work?

  32. 6:44

    If you think about why coding is the first mature AI app layer market, that's really no accident because we were blessed to have something called GitHub from Web 2.0, which solved all three of these at once.

  33. 6:55

    Unit tests, they sort of give you objective decomposable correctness. The community agrees on what working code means generally, and there are effectively infinite public examples with these commit messages as basically free reasoning traces.

  34. 7:08

    So they score high, high, and high on these three axes. Now, look where the money is trying to go now. Biology, security, taste, finance, healthcare, law. These sit pretty low on veracity, pretty low on verification, and the verification examples are sort of like locked all in these enterprise workflows no Web 2.0 system ever captured really, or very

  35. 7:28

    sparingly few. Uh, that's why, you know, probably maybe you've been approached by Mercora to buy your Slack logs or your Jira logs as of late if you work [chuckles] at an AI company.

  36. 7:37

    And, uh, that's the whole game, right? It's, it's predictive. It's not just descriptive. Classify any professions like tasks on these three axes, and you can sort of tell which markets mature most.

  37. 7:46

    So it's no surprise that after code, we went to search, and after search, we went to finance, and after finance, we went to healthcare and law, and after healthcare and law, well, I'd say cyber, biological, uh, and scientific discovery, and maybe even taste, which is probably the most unverifiable out of all these.

  38. 8:04

    So verification is a sort of bottleneck. Um, let me talk about why the industry sells a lot of snake oil here and why most benchmarks you see are, are sort of quietly fake.

  39. 8:13

    Um, the dominant type two recipe is, like, let's hire domain experts. Um, let's have them chat-- use chat to generate plausible tasks. Let's have them solve those tasks, and let's cherry-pick the ones where the model diverged, and let's package them as a hard North Star benchmark.

  40. 8:29

    And then this is the perverse part. Let's sell the data to hill climb that same benchmark. It's basically, um, and maybe some of you guys here in SF have heard this a little too much.

  41. 8:38

    It's Goodhart's law with a profit motive, basically. The moment your measure, like, becomes a target and then the target is set by people who aren't true to the main experts, it stops sort of measuring anything real.

  42. 8:47

    And so the whole market, uh, I like to call it, sits in a sort of fog of war.

  43. 8:52

    Labs, vendors, and enterprises, they're all kind of guessing which data actually improves the model. Nobody can see this clearly. There's a structural tell.

  44. 9:00

    Contrived benchmarks, they only ever test a single isolated in-distribution question. They can't test whether a model sort of sustains, like, correct reasoning across a long dependent episode because these tests were all sort of meant to be solved in isolation.

  45. 9:12

    Uh, in May, I spent a lot of time in this, um, increasingly, uh, in a lot of Anthropic blog posts, but also like a lot of other researchers have noted that cross-harness differencing and cross-infrastructure differencing is the primary cause for a lot of benchmark divergences in performance.

  46. 9:28

    A lot of you are probably very familiar with like the frontier suites of the world and the deep suites of the world benchmarks. Um,

  47. 9:36

    there are issues that all of these benchmarks have, uh, related to that specifically. In particular, just high false positive and false negative rates that are not immediately obvious when you look at the above head stats on the benchmark.

  48. 9:52

    So a single benchmark number under a single scaffold is like basically one sample from a distribution whose basically width nobody me-measured. And that's why there's so much what I call benchmark psychosis today.

  49. 10:03

    It's like a-- It's a pretty noisy sample. So, um, you know, really briefly on an example I took, I, I always have like-- I basically have an internal version of VALS AI and a lot of private benchmarks.

  50. 10:14

    Um, so I, I took three finance tasks from three real-world vendors, um, like an ARR waterfall reconciliation, an LBO valuation memo, and a sort of like long-short pair trade for hedge fund, uh, trading.

  51. 10:26

    So it's long-horizon, non-verifiable finance tasks, relatively robust like deterministic verifiers paired with LLM as a judge, uh, applied correctly. So, um, when you actually run these, I think [chuckles]

  52. 10:41

    you'll, you'll notice very clearly Opus 4.8 is worse than 4.7 on a lot of these rubrics. Uh, you'll, you'll start noticing things if you actually do ru-rubric analysis at the 4.7 to 4.8 step over-engineered self-reflection.

  53. 10:53

    Um, you'll, you'll notice that, uh, GPT 5.5 and Opus 4.8 score within three points on the same task. They fell in like exactly opposite directions, whereas GPT nails the arithmetic, um, but Opus nails the, uh, the methodology but loses arithmetic, which is to say, uh, it's, it's clear if you actually do very good agnostic benchmarking, um, you,

  54. 11:14

    you can tell where the post-training directions for a lot of these teams went. And subsequently, it informs, I think, like a lot of the data that, that comes for this.

  55. 11:24

    Um, so look-- you know, going back to the start, a single benchmark number is a sample from a distribution nobody measured, and it's basically like taking a Swiss Army knife and using the screwdriver bit to cut cheese and like concluding the knife is broken.

  56. 11:36

    Um, [chuckles] so, um, you know, the, the receipts from the last slide, this, this goes into an RL environments report that is sent to labs. Um, but no single leaderboard number ever shows you this part of, um, you know, data which some of the most sophisticated RL environment companies, some of which are actually here at this conference, um,

  57. 11:58

    would, would show you this. So, um, how do we read sort of which domain is next instead of chasing hype? And you can actually use this as a proxy data markets as an upstream indicator of what next application layer products the labs will come out with.

  58. 12:15

    In, in January, uh, Anthropic was spending a lot on cybersecurity data from new vendors.

  59. 12:21

    And in March and April, they were spending a lot on biological data, um, from, from associated vendors. And what do you know happened like two to three months after?

  60. 12:30

    Well, Mythos and Cyber and Claude Bio/Life Sciences today. Uh, so if you want to use this checklist I use, classify the professions tasks into three axes, apply the long horizon bar, the, the ones real labs use, um, enforced step length, heterogeneous tool calls that aren't interchangeable, state transitions that genuinely constrain future actions, mandatory failure recovery.

  61. 12:55

    Uh, you'll notice a pretty mixed bag, uh, of, of how long horizon is even defined from the vendor perspective in terms of specs. Um, and then three, you want to look at the raw data for five signals.

  62. 13:05

    So sequential decisions against like one entity, an inferable action-- uh, expert action per step, outcomes recorded by independent parties. Um, but moreover, just economically available high wage work. So, uh, su-subsequently, one, one counterexample to keep you honest, like, uh, robotics, the modality is not settled, um, Ego versus Teleop versus Umi.

  63. 13:26

    But moreover, I, I find just generally a huge degree of unsophistication within robotics data vendors today. A huge degree of unsophistication. Y- [chuckles] uh, I don't know. Some of you guys are probably robotics researchers in a crowd.

  64. 13:38

    I don't know how many times like people have come up to you and they're like, "Oh, here's like a hundred thousand hours of like iPhone video from my friends in India.

  65. 13:45

    Do you want to buy this for Ego data?" [laughs]

  66. 13:48

    Um, so, uh, in, in, in that sort of domain, your vendor's, uh, choice is entangled with an unsolved research question. At the end of the day, you have to realize like RL environment companies, if they actually succeed, they're more so research accelerators.

  67. 14:00

    It's a boutique industry. If it's venture scalable, it's because the infrastructure they're building agnostically in-house helps an enterprise application layer use case rather than assuming that data markets today will stay as they will forever.

  68. 14:14

    So don't die on a modality hill, basically.

  69. 14:18

    So, so one last exhibit just from my work. This is the map under everything I just described. Look, the share of the white-collar work is sort of on the vertical axis.

  70. 14:25

    Task horizon is on the horizontal. Short horizon is a l- is very addressable right now. But you think about this long tail on the right side, this deep dependent long horizon work.

  71. 14:33

    That's where the real economic value and data build out sort of both lives. And this threshold line that comes-- kind of moves rightward every time somebody builds a real world data pipeline.

  72. 14:44

    Um, now I want to talk a bit more about the model layer. Um, it changes who needs to build what. So historical fact, no pioneer of an infrastructure technology has actually held more than ten percent of the market in the long run.

  73. 14:57

    Um, I'm not saying th-this applies to Anthropic, but you know, just those who forget history are condemned to repeat it. Railroads built company towns. They charged tyrannical rents. They got nationalized as soon as the automobile moved in.

  74. 15:09

    AWS and Google, they consolidated the infrastructure layer and still never captured the application layer. And right now, OpenAI and Anthropic are carving out these fiefdoms. But like all these pressures from anti-distillation export appeals, enterprise exclusivity, they're basically the equivalent of railroad rents in their erode, right?

  75. 15:27

    Like the automobile in some cases has already arrived, like GLM five point two surpassing G-GPT on a lot of real world rubrics, um, is, is, is pretty hard proof that a lot of app layer companies can decouple themselves from the model layer.

  76. 15:42

    And because models differ on efficiency and modality, they're not fungible like electricity. So it's not, it's not exactly leading to nationalization, but it's definitely not heading to durable lock-in either.

  77. 15:52

    And the whole question on what to build next hinges on whether a general enterprise can decouple models from foundation model labs. Um, but luckily, you know, w-we, uh, for a lot of data companies, this is actually where they're headed.

  78. 16:06

    I came here to give a talk on data markets. I'm here to tell you that the successful data companies nowadays are all pivoting to enterprise.

  79. 16:16

    I maybe shouldn't say this in a public setting, but Mercor and Handshake, people don't notice an incredibly large amount of their revenues are enterprise now. Enterprise in ways you wouldn't expect a data business to do.

  80. 16:29

    Once enterprises stop renting a lab's intelligence and starts owning their own, you need an entire abstraction layer that doesn't exist yet. There's like five jobs, serve and route small models targeted by cost latency and performance profiles.

  81. 16:41

    Um, and, uh, two, manage your RL data sets instead of like, um, a cross-base model migration. So when you swap to a new open source base, you rerun post-training automatically instead of starting over.

  82. 16:55

    Three, the Antikythera mechanisms that I talked about before. Um, and in, in the interest of time, happy to talk more about like emerging infrastructure needs afterwards if you want to.

  83. 17:04

    So let me bring it all together. I think data companies all realize that they have to be Neo labs. Data businesses do not stay data businesses because the durable value sort of accrues to the services and app layer of actual work.

  84. 17:15

    So two takeaways. You know, if you're a researcher, stop outsourcing your definition of realism to the same vendors you buy your evals and tasks from. Um, that's kind of just letting the task-- test writer grade the task.

  85. 17:25

    And if you're a builder, your moat's not the data, it's the sort of pipeline into real world work, plus the infra to keep retraining on it as the models improve underneath you.

  86. 17:34

    I'll close on this. Uh, I'm building, um-- I'm working on something new. [chuckles] This is the first public announcement of it. I'm building Antikythera mechanisms. I'm working with a lot of real world companies to monetize their data assets, but moreover, help implement RL as a service with a lot of the enterprises in the world, mitigating a lot of

  87. 17:52

    the pitfalls of a lot of companies I mentioned. If you guys want to talk about it afterwards, I'm on Twitter. I always write a lot on Twitter and Substack, and, um, I'm around afterwards too.

  88. 18:03

    Thank you. [audience applauding] [upbeat music]