← All AI Engineer talks

AI Engineer World's Fair 2026

Verifiable Environments for AI in Biology — Kenny Workman, LatchBio

About this talk

LatchBio co-founder and CTO Kenny Workman explains why rapidly growing single-cell and spatial-biology datasets create opportunities for scientific AI agents. He presents SpatialBench, a benchmark of verifiable spatial-analysis tasks using deterministic graders, and describes how expert review uncovers ambiguous task specifications and problematic grading thresholds. He then introduces SpatialBench-Long for realistic, multistep biological investigations and argues that rigorous benchmarks can drive improvements in scientific AI models.

Chapters

  1. 0:00Why biological data needs AI agents
  2. 3:17LatchBio's path from scientific infrastructure to agents
  3. 6:41SpatialBench and deterministic evaluation design
  4. 10:16Human verification and ambiguous benchmark tasks
  5. 12:11SpatialBench-Long and realistic scientific workflows
  6. 15:45Biology-specific model limitations and the benchmark flywheel

Talk transcript

  1. 0:00

    [upbeat music] Thank you to the organizers for having me. Uh, I'm one of the co-founders and CTO at Latch.

  2. 0:17

    We are basically a vertical AI lab for benchmark and agent engineering, um, hoping to motivate and explain exactly what that means today. Um, starting directly with motivation for agents in, in bio generally.

  3. 0:33

    Um, many people in my domain are familiar with this curve, uh, but this is basically the log linear curve of data generated over the years in, in biology. And the reason I'm bringing it up, it will become directly important to the kinds of things we want to do in engineering.

  4. 0:46

    Um, this curve is driven by a very small handful of experimental classes. Uh, one is called single-cell biology. Um, this is where we split up cells, break them apart, measure their RNA.

  5. 0:58

    The second is spatial biology, which will become the focus of the next segment of the talk. Same thing as single cell, but you get spatial resolution. You can look at how RNA is spread out geometrically over a tissue.

  6. 1:08

    And the third thing is proteomics. It's a broad category of different techniques. They measure proteins, um, less, less abundant in ordering, like less, less data volume generated relative to the other two, but still important.

  7. 1:21

    Um, you guys are technical, and I always think it's good to ground things somewhat quantitatively, but, um, these are really big numbers, and the experimental data from these techniques is growing quite rapidly.

  8. 1:33

    Um, almost greater than any other domain of science other than particle collider, um, machines. Single-cell experiments can yield two to six terabytes per run. Spatial runs can yield seven terabytes of run, proteomics a few hundred gigs.

  9. 1:47

    Um, and the only reason I bring this up is to say, hey, like, m- the output of a single experiment can exceed what a scientist can safely store on a consumer laptop in many cases.

  10. 1:58

    And the laws driving how the, the molecular capture works point to rapid, uh, gains in this throughput over the coming years.

  11. 2:05

    Um, one thing I like to do, uh, is when I read a new paper, is decompose it and align it to this framework because it will become important in a second.

  12. 2:13

    Uh, modern biology research is centered around those experiments. You basically choose a model, biological model, not the kind of models you guys are used to. You generate data from that model, you process the data, you creatively think about the results in the context of prior literature, and you make a claim.

  13. 2:30

    Almost all modern experiments, papers that you see published follow this loose structure with a lot of nuance. Um, all, all, all that to say is they be- they become something of a, a panning experiment.

  14. 2:40

    You're looking for a signal using measurement in a sea of noise. Um, and so this is building up to the claim that like code and SWE, data analysis scaffold to agentic biology.

  15. 2:50

    It becomes this executable substrate that we can use to train things. It induces a natural way to benchmark and climb capability. Um, I've written about this a lot at this blog, uh, link here.

  16. 3:02

    There's a lot more depth to this claim, so I wouldn't take it at face value, but it's something to look into. All, all you can take away from this is like, just like code provided a verifiable substrate for complex software tasks that are not inherently verifiable, uh, data analysis might do the same thing in bio.

  17. 3:17

    So how did we get started? We were originally a data tool vendor for biotech and pharma, um, where we started five years ago out of Berkeley. I'm [REDACTED:age]. We started when I was [REDACTED:age].

  18. 3:29

    We stored, transformed file data from large experiments as a service. Uh, we tried to build products, explored lots of things. Over the last two years, um, we started moving away from biotech and pharma and more towards the people who build those kits I was talking about.

  19. 3:45

    Packaged the software into kind of white-labeled things that they s- provided the scientists themselves, um, help them analyze their data. And then over time, there became this strong interaction, uh, with the agents, uh, using the infrastructure components as tools in the loop context you guys are familiar with.

  20. 4:03

    Except in our domain, the tools can take days or weeks. I'm serious. Um,

  21. 4:09

    what started to happen around last summer is agent prototypes, uh, started to work. So we took, uh, coding models. Uh, to our knowledge at the time, they were not seriously post-trained on any tasks in biology to this point.

  22. 4:21

    And we started to build products that look a lot like all the other agent products. They have a chat interface for you to ask questions to, and they build dashboards and dispatch operations to external compute.

  23. 4:32

    Um, the kinds of things these-- that this, this agent would do is take large file data from the types of experiments I was talking about earlier, say like, uh, tissue bisec-- uh, biopsy from cancer with a spatial measurement.

  24. 4:45

    And then the, the scientist is just iterating with it to get at some question they have, you know, maybe between a malignant and non-malignant part of the tissue, what kind of genes are being overexpressed.

  25. 4:58

    Um, but what was fascinating is e- even though it was pretty bad, it showed the early signs of working. And, uh, it became clear to us at this time that agentic biology might look a lot like code.

  26. 5:08

    I actually lifted this slide from Anthropic's, um, cloud science announcement yesterday. But we-- just like, you know, you had this like kind of faulty, silly engineer that became better, and then a- as it improved in capability, you could dispatch work to teams of them that work together.

  27. 5:24

    The same pattern will probably emerge in science, and products and harnesses will emerge to orchestrate, uh, work and abstract it so teams of s- teams of agentic scientists can take on capability.

  28. 5:36

    But we needed focused post-training, uh, 'cause at the time and still now, um, frontier models cannot be trusted to do real work. They're missing some capability between knowing biology and writing code, and this is exactly extracting scientific insight from real-world data.

  29. 5:49

    Unlike code, it-- which is one constituent component of this work, uh, it also involves data analysis and domain reasoning, scientific reasoning. We thought spatial biology was a good place to start, so we started building agents.

  30. 6:00

    This is a technical greenfield. We had many existing customers. Uh, it's also just a beautiful example of measurement drives progress. You can actually see biological phenomena play out. Um, you can look at a developing mouse embryo.

  31. 6:13

    Um, and we had to get in the guts of how the data was captured and analyzed to build good agents. Um, I'm not gonna get into this in detail, but I put up this tree of, uh, different capture technologies in spatial biology to highlight the diversity of things that exist.

  32. 6:28

    They really span advances in chemistry, optics, semiconductors, physics. Each branch, uh, is induced from decades of cumulative work to figure out how to measure a, a type of molecule.

  33. 6:41

    As, as a specific example, one technique we work with is called sequencing-based spatial, and it's where you take a slide of little beads with clumps of DNA attached to them that fuse the RNA inside of a tissue section.

  34. 6:54

    So biologists can, like, lay like a chunk of a tumor over it, and then it'll capture all the RNA in it, and then let you know with precise, uh, geometric resolution where the RNA was.

  35. 7:04

    It's cool stuff. The data, when it comes out, ends up looking like a big matrix of numbers in a large high content image. Um, you have to take it through a sequence of steps, uh, to get to the end thing that you want.

  36. 7:16

    These steps are highly variable, especially across technology types, tissue, disease contexts. Um, there isn't a lot of consensus in the field, uh, for each step, so we really need a measuring stick to understand if the agents we're building were doing scientific work.

  37. 7:30

    Uh, the existing benchmarks we saw at the time did not measure the tasks relevant to this category of work. Um, they mostly measured things in a Q&A setting, so like what would you do in a kind of academic way, or they weren't sufficiently focused on the experiment type.

  38. 7:43

    This is an actual screenshot of, um, Anthropic's model card at the time that we built this benchmark. So we built one, it's called SpatialBench, um, last December. There's one hundred and forty-six problems.

  39. 7:55

    They spanned the, the different kits I talked about or attempted to, and then they spanned all those different tasks that I talked about as well. So the thing that we found at this time, and still to an extent is true today, is the, the grading of these end outcomes in biology, uh, is too sparse because the models

  40. 8:12

    are pretty bad. So you have to break things up into manageable chunks to get some semblance of verifiability, and that's kind of induced by, um, sticking to these little components of like that DAG, that analysis DAG.

  41. 8:26

    Um, getting data to a state where it would exist right before a scientist or theoretic could do work on it, and then, um, figuring out what the ground truth would be in that context.

  42. 8:34

    So a single evaluation kind of looks like one or more data nodes, again, like a matrix of numbers, a high content image, something like this. A task prompt carefully describing some scientific goal, configuration for a grader, and then a deterministic grader.

  43. 8:47

    It's like a Python function. If you guys notice, this looks a lot like SWE-bench. We borrowed a lot of the early ideas and tried to extend them as much as possible.

  44. 8:53

    Evaluation ends up looking like this, just a lot of JSON. And we ended up identifying properties of like what we thought good biological tests were. Um, a little, little different from code, and we've built on these over time, but they still hold up.

  45. 9:06

    They gotta be verifiable. You have to be able to check the success condition with the function. Um, nothing's changed there. We'll get into some rubric stuff later, but still holds.

  46. 9:15

    Durability is particularly important. Science does not admit clear ground truth. Um, if you are lazy with your ground truth construction of the task, a possible valid analysis path can come with a correct answer, um, and you'll fail it, uh, incorrectly.

  47. 9:29

    So you gotta make sure you're reasoning about something that's somehow invariant across analysis paths. And then obviously, we're, we're working with agentic stuff here. You don't want the model to answer the question in one turn.

  48. 9:38

    You want the conclusion to require interaction with the data and not some memorized knowledge. In practice, that's pretty difficult. We learned a lot about what models could do and which ones to use in specific context for this category of work for our customers.

  49. 9:50

    And we thought, "Hey, this is pretty cool. Let's, let's start to improve and learn more about this benchmarking problem." So we jumped to human verification and long-horizon extension. I'm gonna quickly breeze through these.

  50. 10:00

    So human verification is incredibly important in science. Science does not admit clear ground truths. Uh, after watching trajectory data from multiple rounds of model releases, circa like January to March of this year, um, we really realized a lot of our assumptions were pretty bad.

  51. 10:16

    Um, and in the absence of like a canonical answer, uh, having a bunch of scientists grade each other's work ended up being like the best proxy. So I'm gonna look at one, one issue to highlight exactly what I'm talking about, um, is problem ambiguity.

  52. 10:32

    A task might ask an agent to split a gene list into two groups of activity, microglial activation, oligodendrocyte, inflammation. Just like biological categories of things. Score the cells, find neighboring oligodendrocytes around some region using an appropriate radius.

  53. 10:48

    Compute a Spearman correlation at two time points. As you can probably clearly deduce, the original pro-problem statement creates a host of open choices. How do you split the gene list?

  54. 10:57

    How do you count what inflammatory genes are? It's like some ambiguous word. How do you normalize the data? Um, what, what, what the hell is an appropriate radius? How do you pull the counts within the selected radius?

  55. 11:08

    Um, these, these are all problems that pointed to tasks that were bad, that only became revealed with human verification. Another issue is just like a lot of people in bioinformatics canonically have used like numerical thresholds to QC stuff.

  56. 11:22

    Just like completely arbitrary stuff. Um, cool thing about evaluation like coding is it forces you to reason about things more rigorously than you would when you're doing the thing yourself.

  57. 11:31

    If you have to teach a machine to do it, uh, you, you might be picking out some structure that's more important or more durable than what you were doing if, if you're just doing it on your own.

  58. 11:39

    So we just found a lot of these numerical thresholds to be like bad. Uh, I'm not gonna get into this. Um, after two rounds of human attempts, we produced a verified subset of the benchmark.

  59. 11:48

    Uh, we, we published it. That was fun. Uh, and then we also tried to increase the time horizon. So I want to be clear, the frontier of knowledge is still not quite there with biology.

  60. 11:59

    Like the labs are r- starting to catch up with the post-training, but we kind of want to stay ahead. Um, so we built a benchmark that we thought would recapitulate like really difficult true work.

  61. 12:11

    Um, so we built a SpatialBench-Long. Um, real biological tasks are messy. They use lots of different experiment types. They use the whole workflow, they don't use little chunks. They are, uh-- tasks, every step is interpreted against experimental design or contextualized with some prior literature and the original goal of what you're doing in the first place.

  62. 12:30

    Um, so we built a bank of these tasks that are really trying to simulate the results sections of entire papers or the kinds of decisions you'd make in practice in industry to make a go, no-go decision on a drug program.

  63. 12:41

    Um, these ki-- these tasks took like a week for a group of three people to make each. Taught us a bunch of stuff. Um, an example is can an agent reconstruct a like metastatic niche in a tumor?

  64. 12:54

    Um, if you have like a tumor biopsy and a bunch of metastatic biopsies from like where it metastasizes and spread across the body, can it use like both the genetics, mRNA of the metastatic lesions and the tumor to like find the part of the tumor that initially seeded the metastatic growth and let it spread?

  65. 13:09

    From that, you can figure out like, hey, what parts of the tumor are more like genetically fit? Which ones actually cause problems? And construct targeted medicines to nip them in the bud.

  66. 13:18

    Like for example, this is one of the benchmark, uh, evals in the long horizon set. None of the models get this right, um, but they're getting there. As you can imagine with these long horizon extensions, uh, verifiable rewards at the end are like somewhat uninformative.

  67. 13:31

    So we, we're starting to play with rubrics, uh, constructing these choke points. If you can imagine like the set of analysis paths as inducing some sort of tree, um, there are nodes that are invariant with respect to, you know, different paths, and you can use these to build rubrics, um, using knowledge of how the tasks work.

  68. 13:49

    Uh, we, we-- we're playing with these. We noticed that, um, they're associated with the verifiable outcomes, uh, which is exciting, but they're loosely correlated numerically, um, making us not fully have confidence in them for things like RL or benchmarking.

  69. 14:06

    Um, a lot of-- lot more work to do here still. We, we still strongly believe the verifiability structure is what's going to carry in-intelligence, uh, a bit, a bit longer.

  70. 14:19

    And so these days, uh, excitingly, we've been expanding, um, from this initial spatial focus. Uh, really cool to see the frontier labs and community adopt these benchmarks organically. Um, we had this interesting position by like, you know, building and shipping products, uh, early and kind of playing with the coding agents, so I think we just had a

  71. 14:38

    early advantage. But the benchmarks are now in like the recent Anthropic Model Cards and, uh, this is a picture from yesterday, Eric just showing the benchmarks, um, at the Claude Science launch.

  72. 14:49

    They don't tell us this happens. They just like do it, um, and you like read about it, and it's cool. Uh, we published a bunch more papers, um, beyond spatial to other omics classes, so other experiment types, single cell epigenomics, so RNA, and then the bit above the, the DNA, and then long horizon extensions of these things.

  73. 15:09

    Um, and then we're starting to index and measure the very gnarly, complex landscape that is drug discovery. We just put out our first benchmark on, um, preclinical pharmacology for small molecules, and then systematically biting off pieces of the program landscape from discovery to development to translation, strati-stratifying it by, um, therapeutic types and experiment types.

  74. 15:33

    We just acquired a, uh, company building a biosecurity to form a s-biosecurity team, and then we just put out a collaboration with American Wetware and a surveillance company called Aclid,

  75. 15:45

    um, some new work. Uh, the first was released this morning. I don't know if you guys have been hearing, uh, refusals kind of suck in biology right now. If you ask Fable basic questions about the like mitochondria, it'll-- won't answer.

  76. 15:58

    Uh, it's kind of stupid. So I mean, this is just like an evaluation problem. There's a lot, there's a lot more nuance to this, but I just use that because people tend to recognize it, uh, where we build routine tasks that simulate the kinds of things a scientist would ask for, and then more sinister red team tasks,

  77. 16:14

    which are supposed to look innocuous, but have some structure that is bad. Like, "Hey, I want to clone a gene into a bacteria," and I'm telling you it's GFP.

  78. 16:23

    It's like a glowing protein. But in reality it's like a, a toxin, um, or, or could be used to bootstrap a virus. We found that the routine tasks like drastic-- get drastic-- or fused drastically more frequently than the red team tasks, which is, uh, not great.

  79. 16:40

    Um, we're aggregating a lot of these results, uh, essential resource along with all the pre-prints and a lot of evals and trajectories for you guys to check out. Um, and I actually did okay on time.

  80. 16:51

    Let's go. Uh, and that's it. So we, we are kind of like-- I don't even-- I hate the word lab, but we're kind of like a research lab for [chuckles] bio, and we do research and deployment of these agents.

  81. 17:02

    So we, we still have like a lot of customers. Um, we work with the kit manufacturers, and we use that to inform what kinds of things we make benchmarks for.

  82. 17:10

    We try to get the labs to compete, um, on the benchmarks because then it makes the models better at our products, and it's a-- it's been a pretty rewarding flywheel.

  83. 17:18

    A lot of growth, and we're hiring aggressively across engineering and science. So if you're interested in this work, please find me afterwards. Thank you. [outro music]