← All AI Engineer talks

AI Engineer Europe 2026

How to Leverage Domain Expertise — Chris Lovejoy, Notius Labs

Read the talk

Building a Domain-Native AI Organization

Domain expertise improves an AI product only when someone can turn judgment into changes. Oracle, Evaluator, and Architect roles offer three ways to organize that work.

From a talk by Christopher Lovejoy

How does clinical expertise become a better product?

How do you take what a doctor knows and build it into an AI product? Christopher Lovejoy encountered this problem repeatedly after training at Cambridge, working for several years in the NHS, and moving into AI in 2018. His work included clinical AI at Tandem and US prior authorization at Anterior. Across these companies, the challenge was to turn domain expertise into a differentiated product, rather than simply have knowledgeable people nearby.

The system for incorporating domain insights matters more than the sophistication of the models or pipelines. That is Lovejoy’s starting thesis. The last mile is understanding the particular workflows, use cases, and nuances of the customer being served. He reports that his previous talk reached roughly 100,000 people across platforms; the recurring follow-up was organizational: which experts should a company hire, and where should they sit so their knowledge changes the product? A domain-native AI organization is built to answer those questions.

Slide showing a prior talk thumbnail about building an LLM-native expert system, a social post, and the question “but how should I build my org to enable this?”
“But how should I build my org to enable this?”
0:310:39
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:31 · section reference included

The opportunity depends on operational judgment

Winning in vertical AI is, in Lovejoy’s framing, an organizational problem. He distinguishes three models for incorporating expertise: Oracle, Evaluator, and Architect. Each gives domain experts a different relationship to assessing and improving the product.

The commercial attraction is that AI can address work previously performed by people, expanding beyond software budgets into labor spending. Lovejoy uses an approximate comparison of a $50 billion vertical SaaS market with a multitrillion-dollar labor market. The related Bessemer analysis supports the broader opportunity, but uses public-company market capitalization and GDP spending shares; those measures do not establish his approximate market-size figure.

Execution is the harder part. Gartner reports that at least 50% of generative AI projects were abandoned after proof of concept by the end of 2025. Lovejoy proposes insufficient understanding of the workflows being automated—and of how experts actually perform them—as one contributing explanation. That is his diagnosis, not a cause measured by the cited statistic. His view is that frontier models are already good enough for many opportunities; organizations need to operationalize expert judgment around them.

Three recurring mistakes turn that gap into a staffing problem:

  • Missing or late expertise: nobody with the necessary judgment shapes the product early enough.
  • The wrong expert: the credentials do not match the actual task.
  • Poor organizational placement: the right person is present but cannot effectively influence the work.

These lead to three practical questions: why do you need an expert, who should that person be, and how should you use their expertise?

2:242:32
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

2:24 · section reference included

Who assesses quality, and who improves it?

Choosing between AI approaches requires a company to recognize good output. That recognition depends on judgment, and domain expertise supplies the relevant judgment. Sometimes it is formal: doctors for healthcare, lawyers for legal work. Sometimes it is informal expertise developed through doing a task and understanding its users. This does not necessarily require a new hire. Someone already inside the company may perform the function and need the authority and organizational support to do it well.

The three roles differ in where they act on a simple loop: assess the product’s current performance, then improve it.

RoleAssessmentImprovement
OracleExpert judges the outputExpert changes the application
EvaluatorExpert defines and measures qualityEngineers act on the findings
ArchitectExpert designs the assessment systemExpert designs automated improvement

The Oracle embeds expertise directly, the Evaluator makes quality measurable, and the Architect builds a system that learns from use and user interactions.

An Oracle performs both sides of the loop. They inspect outputs, use the product, examine traces, and identify what is going wrong. They then make changes themselves: often editing prompts, but also adding documents, tools, or other sources of domain knowledge. The important feature is the short path from recognizing a problem to changing the application.

An Evaluator makes assessment a repeatable organizational capability. They define meaningful metrics and build a way to collect them. Depending on the product, that might mean a customer-derived north-star metric, clinicians reviewing a sample of medical outputs, or LLM-as-judge evaluations. Engineers predominantly implement improvements using those findings, with the expert helping them understand what works, what fails, and what needs to change.

An Architect steps back from individual assessments and fixes to design a system that performs both with limited human intervention inside the loop. The expert’s contribution is knowing how to structure automated assessment and improvement. Lovejoy points to another talk for the underlying improvement mechanisms; this framework specifies who designs that capability, rather than an implementation recipe for it.

4:434:46
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

4:43 · section reference included

Choose the role from the bottleneck

The selection process starts with measurability, then asks about capacity and speed:

  1. Can meaningful metrics capture performance? If quality depends primarily on taste and cannot be adequately measured, use an Oracle.
  2. Can one Oracle cover the work? If not, divide responsibility among experts who own different subsets of outputs.
  3. If quality is measurable, is manual iteration fast enough? If engineers can respond to the findings at the required pace, an Evaluator is sufficient.
  4. If manual iteration is too slow, can improvements be automated? That creates a need for Architect capabilities.

The distinction is operational: a measurable output makes an evaluation system possible, while an engineering bottleneck can make automated improvement necessary.

Small startups commonly begin with an Oracle because one person can inspect the product and immediately improve it. Oracle → Evaluator → Architect is not a mandatory maturity ladder. Change the arrangement when the current process breaks at scale, and only when the next approach is feasible. Moving toward evaluation requires a meaningful metric; moving toward architecture requires both an iteration bottleneck and identifiable methods for automating improvement. Another valid direction is simply more Oracles with distinct areas of ownership.

8:488:51
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

8:48 · section reference included

Granola: one Oracle can remain central

Lovejoy introduces Granola, the AI meeting-notes company, as having recently passed a billion-dollar valuation. His example of an Oracle is Jo, whom he describes as a writer and journalist who joined as the first employee and wrote all the prompts. She researched what makes a good meeting note through extensive reading and conversations with hundreds or thousands of users, becoming the primary gatekeeper of AI quality.

Her work closes the entire loop: examine the notes produced by the current product, decide what needs to improve, and directly revise the prompts. The arrangement fits because there is no objectively perfect meeting note. Human taste remains central, and the product’s concentration on one core output makes a central review-and-improvement function workable even as the company grows.

That does not mean Granola lacks evaluations or contributions from other people. Lovejoy says the company also runs evals and has internal tooling that lets others contribute to prompts. The Oracle model describes where primary judgment and direct iteration remain concentrated, not an absence of supporting infrastructure.

10:5110:59
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

10:51 · section reference included

Tandem: distribute ownership across clinical contexts

Tandem’s medical AI scribe also generates notes, but from doctors’ consultations. Its first domain expert, Roy, was a doctor with McKinsey experience. In Lovejoy’s account, Roy initially reviewed medical notes and updated prompts himself: the same Oracle loop. As the product scaled, one person could no longer cover the work.

Tandem responded with decentralized Oracles. Additional doctors took responsibility for subsets of the product, and the platform grew support for a long tail of prompt customizations. Specialties, countries, note types, and use cases all introduced differences that a single centrally managed prompt could not adequately address. Doctors familiar with those contexts could make the corresponding changes.

Medical expertise was essential, but it did not remove subjectivity: there was still no single perfect medical note. A local expert could own a customer relationship or specialty in a particular location, understand its requirements, adjust the prompt, and make that version live for that context. Lovejoy describes thousands or more prompt variations available across different contexts. Here, scaling expertise meant distributing direct control while providing the platform support needed to keep those variations usable.

12:1912:41
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

12:19 · section reference included

Anterior: scale assessment, then improvement

Anterior’s prior authorization product addresses a different clinical workflow. A doctor requests a treatment or scan—for example, an MRI—and the request goes to an insurer, where nurses or doctors assess whether it is appropriate. Lovejoy describes himself as Anterior’s first technical employee: he built the initial code and prompts, reviewed the outputs using his clinical judgment, and changed the implementation when a decision was inappropriate.

More customers, demand, and variation made that personal review loop insufficient. He moved into Evaluator work: defining metrics and failure modes, building a review dashboard, and hiring clinicians to review outputs. Their assessments produced performance information the engineering team could use to make changes. Assessment had become a system rather than one person’s available attention.

The next bottleneck was fixing the problems. Organizations interpreted policies and rules differently, so manual engineering iteration could not keep up with the variation. Lovejoy describes needing to learn at the edge, where those customer-specific interpretations appeared. That pushed his role toward designing methods for automated improvement, with the implementation details again deferred to another talk.

The progression made sense because, in his account, the relevant decision quality was measurable against medical evidence. The product’s outputs were approval or escalation for clinician review, not approval or denial. Determining whether the chosen action was appropriate still required clinical reasoning. Variation in policy interpretation then created a reason to adapt from usage rather than rely entirely on centrally implemented fixes. Finally, the earlier Oracle work supplied an advantage: firsthand knowledge of failure modes helped determine what the evaluation system should assess and what the improvement system needed to change.

Anterior prior authorization case study slide featuring Chris Lovejoy, his responsibilities, and four reasons the Oracle-to-Evaluator-to-Architect progression worked.
Anterior’s progression from Oracle to Evaluator to Architect, including methods for automated improvement.
14:1414:25
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

14:14 · section reference included

Hire for the actual task and the adjacent work

For an Oracle, relevant expertise means direct experience of the use case. A medical qualification alone may not qualify someone to shape a medical coding product. Have they done medical coding? Do they understand its process and failure modes? The same specificity applies to legal and other professional work. Prompting and context engineering are helpful but comparatively learnable. Depth of insight, attention to detail, and customer communication matter because the expert must discover requirements and translate subtle quality judgments into product changes—the kind of detailed work illustrated by Jo at Granola.

An Evaluator needs domain expertise plus data science intuition: choosing useful metrics, collecting them reliably, and making them actionable. The supporting skills depend on how the evaluation system operates:

  • Statistics: analyze measurements as review volume grows.
  • Industry connections: recruit qualified reviewers from a relevant network.
  • Leadership: manage the review team.
  • Product management: translate findings into effective collaboration with engineers.

These skills support ownership of a measurement system, rather than merely the ability to score individual outputs.

An Architect additionally needs experience with LLM-powered products and the available ways to improve their performance. The earlier skills remain relevant, while engineering ability helps the expert implement changes or steer their development. Domain expertise remains the baseline: technical knowledge is useful because it gives the expert more ways to turn that judgment into an improving product.

17:0017:06
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

17:00 · section reference included

Give the expert authority and room to grow

Define a principal domain expert: one person ultimately accountable for AI quality, with authority to make decisions. This avoids consensus by committee, where shared responsibility leaves nobody clearly responsible. It also gives someone the time and space to understand performance deeply enough to make informed decisions quickly.

Give that person ownership beyond an advisory role. They should be in the room when product decisions are made. Having an expert review outputs and hand the results to someone else captures only part of their value; the larger opportunity is to let them shape the system through which the organization measures and improves quality.

Lovejoy describes an unnamed company with two senior clinicians but no clear principal expert. Final authority was ambiguous, their roles were largely advisory, and progress on the wider quality-improvement system was slow. Both clinicians left after roughly 12–18 months. He suspects insufficient ownership contributed to their departures; the organizational loss was also the accumulated context they took with them.

The third principle is to hire for breadth without expecting one person to possess every adjacent skill. Domain expertise is the foundation; seek as much relevant breadth as practical, then pair the expert with complementary specialists. If statistical expertise becomes necessary, for example, a statistician can collaborate with them.

Breadth creates room for the role to evolve. Someone hired solely for narrow domain knowledge may struggle to take on evaluation-system or architecture responsibilities if the company later needs them. Preserving continuity is valuable because Oracle work builds detailed knowledge of the product: which failures recur, which distinctions matter, and what should be measured. That accumulated understanding makes an experienced Oracle well placed to become an Evaluator.

Three-column slide recommending a principal domain expert, ownership, and hiring for breadth as the role evolves; supporting text mentions avoiding consensus by committee and former Oracles becoming Architects or Evaluators.
Define a principal domain expert, give them ownership, and hire for breadth.

The practical staffing recommendation is to establish a principal expert early, give them ownership, and usually begin with direct Oracle work. As the product and organization change, that person may enable decentralized Oracles or move toward Evaluator and Architect responsibilities. The appropriate direction depends on the use case, scale, and bottleneck; the organizational playbook is still developing. Lovejoy closes by offering the slides on his website.

19:3519:57
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

19:35 · section reference included

Resources

From the talk

Read the complete timestamped transcript
  1. 0:00

    [upbeat music] Okay, so welcome everybody.

  2. 0:17

    Uh, hi, my name is, is, uh, Christopher Lovejoy, and I'm gonna talk about how to leverage domain expertise to build better AI products, and that the way I believe you can do this is by building what I call a domain-native AI organization.

  3. 0:31

    So I'm gonna talk about what that looks like. [sniffs] Um, brief background about, about me because, um, this is relevant. So I started out my career as a medical doctor.

  4. 0:39

    I trained in the University of Cambridge and then worked in the NHS for several years. And I then, uh, in twenty eighteen moved into kind of AI space, [clicks tongue] uh, training and building models and working with various organizations, including Tandem, uh, which is the largest, um, clinical AI product provider in the, uh, in the UK in terms of

  5. 0:57

    adoption. [sniffs] Uh, also Anteria, I was the first employee, uh, and it's a square back startup performing, um, prior authorization in the US and then also, uh, various other startups as well. [sniffs]

  6. 1:11

    And my, um, kind of challenge at all of these different companies was we're building some kind of product that bakes in domain expertise, how can you, how can you do that?

  7. 1:21

    How can you leverage that in a way that then builds a differentiated AI product?

  8. 1:30

    So I talked a bit about this, uh, at the last AI engineer conference in San Francisco. Um, I shared my kind of thesis, which is that the system for incorporating domain insights is more important than the sophistication of your, of your models, of your pipelines. [sniffs]

  9. 1:43

    I talked about the, the last mile problem, which is this challenge of getting your product to really understand the specific nuances of the workflows, of the use cases, uh, of your, of your customer that you're serving.

  10. 1:54

    Um, and, uh, this talk was seen by a lot of people. About a hundred thousand people, um, saw this across different platforms. Many of them reached out to me, and the kind of most common question that they had was, "Okay, but how do I, how do I build my organization?

  11. 2:09

    Like, I, I'm kind of on board. I get, I get that [sniffs] domain expertise is important, but how do I build my organization to actually enable this? To like, what kind of domain expert should I hire?

  12. 2:18

    Where should I put them in my organization to enable this to, to take place?" [sniffs]

  13. 2:22

    So that's what I'm gonna talk about today.

  14. 2:24

    Um, and you know, I've heard people say that winning in vertical AI, uh, you kind of wanna get the best model, and actually, I don't think this is true.

  15. 2:32

    I think that fun-- [coughs] I think that fundamentally, [clicks tongue] uh, winning in vertical AI is an organizational problem. [gulps]

  16. 2:44

    And so I have this, this framework based on different, um, organizations I've seen and, and worked at, and, and how they're baking in domain expertise, [clicks tongue] uh, which is that you can do it in three main ways.

  17. 2:55

    You can have your domain expert as an oracle, as an evaluator, as an architect, and I'm gonna talk a bit about m-- uh, more about those in more detail.

  18. 3:04

    But just stepping back, uh, why, why do we care about vertical AI?

  19. 3:08

    Ultimately, it comes down to the fact that vertical AI is a big opportunity. Uh, so a lot of VCs are talking about, you know, this is the next big thing.

  20. 3:15

    Uh, there's a lot of startups raising big amounts of money on this potential, uh, this promise. And as Bessemer pointed out, vertical AI, uh, you know, we had vertical SaaS, and that was a, you know, fifty billion or something dollar market.

  21. 3:26

    But actually, now AI's moving into the kind of labor force, and that's like a multi-trillion-dollar market.

  22. 3:33

    But we've not yet really seen, um, success at scale. And according to Gartner, about fifty percent of all generative AI projects were abandoned last year. [sniffs] And I think there's many reasons for this, but my take is that one of the core reasons for this is that we're often building AI products and AI systems without really having, uh,

  23. 3:51

    you know, a deep kind of understanding of exactly what workflows we're automating, um, and you know, exactly h-how the kind of, how the domain experts would, would, uh, perform these, these kind of processes. [sniffs]

  24. 4:04

    So just to reiterate, my belief is front-end models are good enough, but the gap is now how do organizations operationalize the expert judgment around them?

  25. 4:11

    And the three most common mistakes that I have seen are, firstly, not hiring domain experts or hiring them too late. Secondly, hiring the wrong kind of domain experts, and then finally not fitting them into your organization appropriately, not leveraging them, uh, correctly.

  26. 4:25

    And so this maps to three questions, which I'm gonna address in this talk, which is, do you really need a domain expert? Um, what, what's the why? If so, who do you need?

  27. 4:34

    And then how do you leverage them? So I touched on the first one, and then this Oracle Evaluator Architect framework I'm gonna go through answers those second two.

  28. 4:43

    So do you really need a domain expert?

  29. 4:46

    My take is that the answer is, is yes, and it's because appraising AI quality is something that's very important to do in your company. You want to be able to make decisions between different approaches based on the kind of output that they give, and your company needs to have a sense of what good AI quality looks like.

  30. 5:04

    And that ultimately requires judgment. And the best kind of judgment involves some kind of domain expertise that you can then bake into your, into your company and into your product.

  31. 5:14

    And this domain expertise could be kind of like specialized domain expertise. It could be you're building like a healthcare or legal product, and you therefore need doctors or lawyers and like bring their domain expertise in some way.

  32. 5:24

    But it doesn't have to be. There's also kind of informal domain expertise, um, and I'll give some examples later. Um, and it's worth saying also that,

  33. 5:32

    you know, I'm, I'm, I'm not making the pitch in this talk that you should go out and hire a domain expert per se. I mean, maybe that makes sense, but in many cases, you might already have somebody in the organization that essentially performs this function, and you want to basically kind of like empower them or, or understand

  34. 5:45

    how best to organize, uh, your organization around them. [sniffs]

  35. 5:51

    Okay, so then who, who do I actually need? I'm gonna go into the framework now. So I mentioned these, these, these three models for incorporating domain expertise. The first of these I call the Oracle, which is that you have your domain expert who directly embeds their domain expertise into your actual application.

  36. 6:07

    Um, and I'll talk about what that actually looks like in, in a second. The Evaluator performs a different kind of role. They define how you measure quality. Like, w-what is it that matters?

  37. 6:17

    What are we optimizing for? And once that's defined, they set up a system where that can then be measured, and you get that data that you can then ultimately use.

  38. 6:25

    And finally, the Architect builds this system that automatically improves itself and, and bakes in more domain expertise and just learns from being used by users, uh, and those interactions. [sniffs]

  39. 6:37

    So if we consider, like, a very simplified, um, view of what does it mean to build an AI product, there's kind of two stages, right? Like, you assess how, how's my AI product doing right now?

  40. 6:45

    Like, what is the state of play? And then how can I improve that?

  41. 6:50

    So in the Oracle model, both of these steps are being performed by the domain expert. So the domain expert is looking at the AI outputs, they're playing with the product, they're looking at the traces, and they're seeing, okay, how's it doing, where's it going wrong, uh, what could be improved?

  42. 7:03

    And then they're going in and they're improving it. And it, you know, might be in a simple case, they're just, like, kind of tweaking prompts, I think is, is often, um, the mechanism here, but it might be that they're also baking in some kind of domain expertise, um, you know, via adding documents or, or tools or, or

  43. 7:17

    something else. Um, there's a few different levers, but fundamentally, they're doing both sides of the equation. Now, this changes when you think about the Evaluator. Now, the domain expert is still assessing, but assessing is actually more complex.

  44. 7:28

    So they are defining metrics that you can store some kind of objective way of quantifying performance, and they are then, you know, building that system. It might be that you're gonna get the metric from customers.

  45. 7:40

    The customer uses it. There's some kind of user metric that you're gonna, you're gonna use and, and uses your north star that determines, you know, really what you care about.

  46. 7:46

    It might be that you want to actually hire out domain experts to perform some kind of reviews. So in a clinical context, you might wanna hire clinicians who review a subset of all of your AI outputs and determine how you're doing.

  47. 7:57

    You might do things like LMS Judge. Um, this can, this can really kind of scale in different directions depending on your use case. But actually, the improvement is now done, you know, more by engineers.

  48. 8:07

    The domain experts can still contribute, but predominantly their role is understanding the state of play, understanding what's working, what's not working, and then working with the engineers to actually then go ahead and, and make those changes. [sniffs]

  49. 8:17

    And then the final, which is the Architect, here, kind of stepping back, the, the domain expert is actually designing this, a system that does both. So the idea is that there's, there shouldn't be too much human in the loop in this actual process in the middle, but the domain expert here is creating the ability to do that,

  50. 8:31

    um, and, you know, leaning on the different mechanisms that are available to, to kind of have this automated improvements. And I spoke about this before, so there's a link at the bottom.

  51. 8:39

    Uh, I'm not gonna go into detail exactly what the system looks like, but there's, there's different, uh, different levers that you, that you can lean on. So then that's the framework.

  52. 8:48

    Who do you therefore need for your specific organization?

  53. 8:51

    I kind of came up with this, like, rough guide to understand based on your specific use case and also your specific scale what makes the most sense. And the first question, I think, to ask yourself is, can I measure performance in metrics?

  54. 9:05

    Is there some kind of objective thing that I can, I can measure and that that's really meaningful, or is it something where actually taste is a bit more important?

  55. 9:12

    Um, if it's not the case that you can measure it, then you want to take an Oracle approach. And the question is, is one person enough to do it maybe at my current scale, or maybe also certain types of products are actually very suitable to just have one Oracle.

  56. 9:26

    And I'll give some examples to make it more concrete. Um, whereas if not, then you might have multiple who maybe handle different sub-subsegments of your, of your kind of AI outputs.

  57. 9:36

    Now, assuming that you can measure things and you want to measure things, the follow-on question is, okay, w- is manual iteration fast enough? Can I just have an engineer who makes some changes, um, and that, then that's fine.

  58. 9:45

    I can, like, adapt to customer needs and, and, and learn. If so, you only really need the Evaluator. You need your domain expert to basically assess what is quality, like how are we doing right now, feed that into engineers who make the improvements.

  59. 9:55

    But actually, if that's not fast enough and you, you kind of don't wanna be relying on this, this human iteration, then you want some approaches that an Architect would come up with.

  60. 10:05

    And it's worth also calling out that within an organization, this can evolve over time. So a common starting place is, is an Oracle. Particularly if you're in a startup and you're kind of small in scale, your domain expert probably should be the Oracle to begin with.

  61. 10:15

    But then they can progress, and in some cases, they wanna progress one way, in some cases another way. Um, there are certain criteria for when it's actually necessary and also if it's possible.

  62. 10:26

    So it's necessary if things are currently breaking at, at your scale. Um, and then depending, you might go in one or two different directions. But you can only actually move on to an Evaluator if there's some kind of objective, uh, metric you can measure.

  63. 10:38

    And likewise, as I mentioned, if manu-manual iteration is too slow and you can identify methods to automate improvements, then you might, uh, progress to a kind of an Architect model.

  64. 10:47

    All right. L-let's make this more tangible with some case studies.

  65. 10:51

    So Granola, uh, for those of you who don't know, is a, is a company that, uh, generates AI meeting notes, um, recently, uh, kind of passed a billion in evaluation.

  66. 10:59

    Um, and my, my friend at Granola, Jo, has a background as a, as, like, a writer and journalist, and she joined Granola as the first employee. She wrote all, she wrote all of the prompts, and she did extensive research, like many, many hours reading papers and, um, talking to hundreds, thousands of users to understand r-really what it

  67. 11:18

    is that makes a good meeting note. Uh, and, and her role has kind of been to be this primary gatekeeper of AI quality.

  68. 11:24

    So question to the room, of the three models I've outlined, like which, which would you say this, this sounds like?

  69. 11:32

    Yes, exactly. So what she's doing is she's doing both sides. She's, she's kind of assessing the outputs, looking at the meeting notes that the current version of the product generates, and then she's making those, uh, improvements herself directly and, and doing this iterative loop.

  70. 11:44

    And in Granola's case, my argument is that this makes sense because firstly, there's no objectively perfect meeting note. We can't just say, "Okay, this is, this is the best note."

  71. 11:52

    Um, so ac- it's a bit more about human taste than kind of really baking those sorts of things in. And then because with meeting notes, um, Granola, that is the main kind of core output of their product is this meeting notes, so actually it's amenable even at scale to have this kind of direct human review and improvement

  72. 12:08

    loop. Um, so even as they've scaled, Jo is still kind of largely playing this role. There's, there's nuances. They also are running like evals and, and have, um, kind of built out some other internal tooling to help other people contribute to prompts.

  73. 12:19

    But fundamentally, uh, this is a process of, um, kind of working as an Oracle. So let, let's consider a, a, a second use case. So Tandem, uh, have this medical AI scribe product, uh, which basically listens to a doctor's consultation, and again, basically generates meeting notes, but this time they're medical, so that introduces some nuance.

  74. 12:41

    And their first domain expert that they hired was someone called Roy. His background is he's a medical doctor, and then he went to McKinsey. And what he did is he reviewed the medical notes, he updated the, the prompts.

  75. 12:51

    So he's following this kind of Oracle approach. But then as that scaled, it became impossible for one person to do this.

  76. 13:00

    So what they did was, uh, they went the approach of a decentralized Oracle and hired out, um, various other doctors to, um, basically do this in kinda like subsets.

  77. 13:13

    And so they updated their platform to support this kind of long tail of prompt customizations. The challenge was that they're serving many different specialties, many different countries, many different types of notes, many different use cases.

  78. 13:23

    So they kind of needed the ability to have doctors from those different countries, specialties, et cetera, to be able to make these, uh, tweaks and changes.

  79. 13:33

    And this, this worked well for them because you need medical expertise, um, hence needing domain experts. There's some subjectivity. There's no, like, perfect medical note, in the same way there's no perfect meeting notes.

  80. 13:42

    There were these variations I described. And then because there's so many customizations, you need many different domain experts, each who might own a particular relationship or a particular subset.

  81. 13:50

    Like, you might have somebody who, uh, is doing, you know, some subset specialty in some, uh, geographical location who will work with the customers there, understand what they need, tweak the prompts, and then that prompt version is live there.

  82. 14:04

    But then they have many other, you know, like thousands or more of different variations on prompts that are available in different, different places. Okay, third and final case study, and excuse the self-reference here.

  83. 14:14

    This is, um... This is, is it... I, I can use it as an example because I'm obviously very, uh, familiar with this one. Uh, so I worked at Anteria, and we were-- our first product was prior authorization.

  84. 14:25

    And prior authorization, for those who are not familiar, is basically a process in the US where you go to your doctor, your doctor requests some kind of a scan, let's say requests an MRI.

  85. 14:33

    It then goes to the insurance company, and the insurance company has nurses and doctors inside, uh, who will determine, okay, should this treatment be approved? Like, is it appropriate?

  86. 14:41

    Um, and so at Anteria, I was the first technical employee. I built the initial product, the, the prompts and the code. Um, and then I was reviewing our outputs.

  87. 14:51

    I was kinda putting my doctor hat on, clinically assessing is this, is this an appropriate, uh, decision that's being made? Um, then updated the prompts and, and the code to handle that.

  88. 15:00

    But again, that didn't scale, and as we had more customers and we, we had, uh, you know, more, more kind of demand, um, and more different variation in, in what we were serving,

  89. 15:11

    it was then necessary for me-- for, for that role to evolve. And so I then defined metrics and failure modes. Um, I built a review dashboard to enable clinicians to come in and look at certain outputs and hired clinicians to perform that process.

  90. 15:24

    So we sort of scaled that out, and that gave us these metrics and information on, on performance that we could then use to collaborate with the engineering team to make these changes.

  91. 15:32

    But again, this also, um, didn't scale from the fixing point of view. So having the kinda like manual iteration of these engineers didn't work because, um, there were many different variations in how these organizations are ultimately, uh, yeah, kind of like, um, uh, where was I going with this?

  92. 15:51

    Uh... Yeah, the, the, the ki- the kind of... There's a lot of variation in how they interpret their policies and their rules, um, so you kind of need a mechanism to learn that at the, at the edge.

  93. 16:03

    Uh, and that was why I then needed to kind of evolve more into an architect approach and, um, design methods for automated improvement, which I talk about in another talk here.

  94. 16:14

    And so overall, this made sense for us to go from, um, Oracle to Evaluator to Architects because in our case, the AI quality is clearly measurable. Um, the AI output is either approval of the prior authorization decision or escalation for a clinician review, so it's either correct or incorrect based on the medical evidence.

  95. 16:29

    You need clinical reasoning to determine that. There's a large variation in how prior auth rules are interpreted. Uh, that, that was the point I was trying to make here.

  96. 16:35

    So therefore, because of this kind of variation, you need a system that is able to adapt dynamically and learn from the usage such that then you can, um, can, um, solve that.

  97. 16:45

    And then also, uh, this progression, so understanding the failure modes, uh, as an Oracle and, like, really kind of having a deep understanding of the way in which the AI is performing and the way, the ways in which it's failing is very helpful for deciding, def-- uh, defining how to assess, uh, and improve things.

  98. 17:00

    Cool. So who do I need from my, my organization, and what skills should they have?

  99. 17:06

    So to think about what kind of skills you might wanna have for each of these different roles, it's helpful to kind of come back to what are they doing?

  100. 17:12

    So the domain expert is kinda looking at the outputs, tweaking prompts, um, making changes, improving the product. So the core skills that they require is relevant domain expertise, and by relevant here, I particularly mean direct experience of the use case.

  101. 17:25

    Uh, so, you know, just being a doctor might not actually be sufficient. Let's say your product is medical coding, you, you know, have you got experience of medical coding?

  102. 17:32

    Do you know what that, you know, involves? Do you know where that can go wrong? Um, so yeah, you kinda wanna think more granularly than just, oh, I need a, I need a doctor, I need a lawyer, I need, you know, whatever domain expert you might need.

  103. 17:42

    And then other relevant experiences, you know, things like prompting, context engineering. These are all, like, nice-to-haves. I mean, I think these are a little bit more learnable, um, but, but are helpful.

  104. 17:51

    You know, depth of insight, attention to detail. I mentioned the example of, of Jo Grano- uh, Jo at Granola. She, like, really went into the details, um, and, uh, that was very valuable.

  105. 18:00

    And then things like customer communication to understand, you know, what exactly do your, your customers want.

  106. 18:05

    For the Evaluator, if they're doing something differently, they're, they're, um, building up this kind of system to understand the performance. So there, I think, as well as domain expertise, you really want this data science intuition because I fundamentally see, like, a lot of this as being kind of data science skills.

  107. 18:21

    You're defining metrics that you care about. You're figuring out a way to build a system that collects those metrics and, and makes them usable, and th- these are kind of-- these are all data science skills fundamentally.

  108. 18:30

    Um, it's also helpful to, you know, have some of the skills that we mentioned from before. You might want statistical skills if you're doing this at scale, and you're really kind of wanting to analyze some of these metrics.

  109. 18:39

    Industry connections are helpful if you go down the line of building out a team that does these internal, uh, reviews. It's helpful to be able to hire from that network.

  110. 18:46

    It's helpful to have leadership experience to then manage that team, and it's helpful to have product management experience because you're then gonna help feed into, uh, engineers making these improvements, and you want to be able to, um, collaborate with them effectively.

  111. 18:57

    And finally, what do you need if you are a kind of like architect profile? Um, similar, obviously, the domain expertise is kind of a given, and then you also want experience working on LLM-powered products.

  112. 19:09

    Uh, so you want to know what are the kind of levers that you can lean on to ultimately improve performance. Um, and then it's also relevant, you know, as I mentioned, everything kind of from before, and then engineering skills can also be helpful because you can lean on those to even, like, implement or, or steer the, the

  113. 19:22

    development of some of these, uh, improvements. So third and final question: How should you leverage them? Okay. You found the perfect domain expert. Maybe you already had them, or maybe you have, have, you know, found someone you wanna hire in.

  114. 19:35

    Um, I think three principles that I, that I, uh, yeah, I, I, believe are important. The first is there's a lot of value organizationally in defining a principal domain expert, and what I mean by this is a single individual who's ultimately accountable for the quality of the AI performance and makes-- can therefore, like, make decisions.

  115. 19:57

    Um, so this avoids, you know, kind of consensus by committee where it's everybody's kind of responsible, so nobody's truly responsible, and, and you move a bit more slowly. Um, and then if that, if that's that individual's res-responsibility, they can really invest the time in deeply understanding how the AI is performing, um, and that can then, you know,

  116. 20:14

    ultimately feed into better decision-making. Uh, but it kind of gives them the, the time and space to really focus on that, and I think there's a lot of value in, in that from a, from a, um, speed point of view.

  117. 20:24

    The second is, I think it's important to give them ownership. You, you ultimately don't want to treat them as just kinda like a consultant or, or somebody that you come to for advice.

  118. 20:32

    You want them to be in the room when you're making these decisions, um, and that's because you wanna be able to leverage them to build a differentiated product. Uh, if they're not in the room, then it's, it's much harder to do that.

  119. 20:44

    And, um, yeah, I think, you know, kind of concretely, if o-o-one of, one of my arguments is you could just have a, a, a domain expert who kind of performs review of the AI outputs and kind of gives that to somebody else.

  120. 20:55

    But what you actually wanna do is you wanna go beyond that and, like, kind of build out a system around that such that organizationally you're able to measure things accurately and improve things accurately and, uh, you know, that, that's an important part of, of building a differentiated product.

  121. 21:10

    Um, so I've seen failure modes here where, for example, uh, one company I was at, uh, had the kind of two different senior collisions. Um, so neither of them were really a principal domain expert.

  122. 21:18

    It was kind of a bit ambiguous, uh, who had the final say. They kind of weren't given that much ownership, um, in, in the process and were a bit more kind of advisory, and essentially the progress was just very slow, uh, in terms of actually, like, building this, this wider system for, for, you know, improving the AI

  123. 21:33

    quality and the performance. Um, and then what ended up happening is actually both of those two individuals after about, like, twelve, eighteen months kind of left the company, I think in part because they maybe didn't feel they had the ownership.

  124. 21:43

    And, uh, that's obviously a loss for an organization because they have a lot of relevant context in their head that, that, you know, you kind of wanna be building on top of.

  125. 21:50

    And the final point is hiring for breadth. So we saw that, you know, all these skills listed on the previous slide. I mean, I think it's a, it's a big ask to, you know, try and find somebody who has all of those skills.

  126. 21:59

    So my general recommendation is hiring somebody who has the domain expertise kind of ha- is, is the base. It needs to be a given. But then as many of those skills as possible, um, and you can always pair them with somebody who then kind of complements that.

  127. 22:11

    So let's say, you know, you end up needing a statistician, and this person doesn't have stats experience. You can find a statistician, and, and they can kind of collaborate there.

  128. 22:18

    Um, but I think the failure mode here is that you, you hire somebody who is maybe just a domain expert. And I don't wanna, like, trivialize being a domain expert.

  129. 22:26

    I mean, obviously that, that's, that, that can be a big thing, but, um, that can then make it hard for that individual to grow from being an Oracle to Evaluator to a, to an Architect if that's necessary.

  130. 22:35

    Uh, and then, you know, organizationally, maybe you need to bring somebody else in, and that, that's, um, kind of less, less desirable, I think, than having somebody who's kind of with you the whole way through.

  131. 22:43

    And also actually having, having experience as kind of in this Oracle role in the company, you have a great insight into how the AI performs. You're very well-placed to then become an Evaluator.

  132. 22:52

    You know what kind of things you should be measuring, what kind of failure modes to be avoiding, and, uh, yeah, it's, it's kind of very complementary to, to have somebody here who evolves there.

  133. 23:01

    So just to summarize, um, I think assessing the AI quality in your product requires domain expertise that can be formal or informal. There's broadly three ways that I've seen domain experts, um, kind of be baked into the organization.

  134. 23:14

    The first is directly bringing domain s-expertise, uh, into the application, and this is kind of the Oracle role. Also defining quality so that you can measure and improve it, which is the Evaluator role.

  135. 23:24

    And then designing systems that improve and learn over time, so this Architect role. And the appropriate approach really depends on your specific use case. It also depends on your current scale.

  136. 23:34

    And each of these approaches requires different adjacent skills. And my advice is hiring a principal domain expert early is very effective organizationally. You want them to have the relevant domain expertise, and you want this breadth of adjacent skills.

  137. 23:46

    You wanna give them ownership. You wanna probably start them as an Oracle in that kind of role. Um, and then as your product evolves, as your organization evolves, their role should evolve along one of the, the axes that make sense.

  138. 23:57

    Maybe it's kind of decentralized, but they're playing a role in enabling the decentralized Oracle. Maybe it's through this Evaluator-Architect spectrum. Um, but it's up to you to kind of, like, figure it out.

  139. 24:05

    And, and ultimately the playbook on how this stuff works is still being figured out at the moment because we're still early on in this, in this journey of building, uh, AI-powered products with domain expertise.

  140. 24:16

    Um, so thank you. Uh, yeah, I put the, the slides on my, on my website if you wanna, um, download these slides. I think maybe do one question. I'm conscious of time.

  141. 24:24

    Uh, but yes. Thank you. Thank you for your attention. [audience applauding] [upbeat music]