AI Engineer World's Fair 2026
From Ambient Documentation to Clinical Intelligence
About this talk
Abridge engineer Chaitanya Asawa explains how ambient clinical documentation can expand into context-aware clinical intelligence and decision support. Clinician testimonials illustrate reduced documentation burden, while product demonstrations show voice-driven clinical-trial eligibility checks and chart preparation. Asawa traces his background at Vicarious and Glean and argues that healthcare AI requires rigorous evaluation because errors carry serious consequences and clinically valid answers can be almost as difficult to verify as to generate.
Chapters
- 0:00Audience introduction and clinician testimonials
- 2:08From AI startups to Abridge's clinical-intelligence mission
- 9:28Voice-driven clinical trials, chart preparation, and documentation
- 11:42High-stakes clinical decision support and the generator-verifier gap
- 20:56Closing: healthcare AI at scale
Talk transcript
- 0:00
[upbeat music] Thank you so much for everyone being here. We're gonna get started in a second.
- 0:16
Um, but before we get started, I am curious, how many of you currently work in the healthcare industry in some shape or form? Oh, that's amazing to hear. Uh, how many of you are clinicians by training?
- 0:28
Okay, a couple. How many people in the room are engineers?
- 0:31
Okay, awesome. Um, and then how many people have heard of Abridge before?
- 0:37
Okay, awesome. Uh, well, I'm gonna let you hear actually from our users to start off on a little about Abridge. [upbeat music]
- 0:53
Full day of twenty two patients, out by four thirty PM, notes done. That's nice.
- 1:01
When I think about Abridge, I think the thing that comes to mind is it's really a cornerstone of how I practice medicine today. Um, there's just no way, um, I would do a clinic or see a patient without using.
- 1:14
I can be present throughout my clinical encounters. I don't have to think about, um, "Oh, wait, did I get that? Do I need to write that down?" Because I know Abridge has my back and has everything ready for me.
- 1:25
Abridge makes me feel free, 'cause I can really look at a patient, really listen, and not have to be thinking about, "What do I need to put in the computer?"
- 1:34
Full day of twenty two patients, out by four thirty PM, notes done.
- 1:37
I get to go home and protect my family.
- 1:38
I don't have to think about, "Oh, wait, did I get that? Do I need to write that down?"
- 1:40
I can really look at a patient, really listen, and not have to be thinking about, "What do I need to put in the computer?" [upbeat music]
- 2:08
Our marketing team produces really good videos, and so they always hype me up. Um, but the goal of this talk for me, and I know that we have a lot of engineers in the room, my g- my goal is to talk about healthcare as a domain.
- 2:21
At least I felt in the past that there was a lot of stigma around maybe the technical problems weren't as interesting in healthcare. And it is true in some ways, there's some parts of healthcare that might not be as tech forward.
- 2:31
A lot of things run on fax machines, for example. Um, but I wanna give exposure throughout this talk of two things. One, Abridge's journey from clinical documentation to clinical intelligence and what that looks like.
- 2:45
Um, and then two, I wanna expose you to some of the technical problems we work on, um, and that have to be th- that are truly frontier AI produ- uh, problems that have the highest stakes.
- 2:57
A little about me. My name is Chaitanya. You can call me Chai. Um, my career has always been in AI companies and startups. I first started as, uh, in research engineering at a company called Vicarious, uh, which its goal was actually to develop AGI, but they took very different methods.
- 3:11
They wanted methods inspired by neuroscience and probabilistic graphical mo- uh, models, um, and they concretely worked on robotics. Uh, then I started working at this company called Glean, 'cause I faced this problem in my workplace itself.
- 3:23
How, uh, like, information scattered all over the place, context is everywhere, and it's so key to decision-making. And Glean was building basically the ChatGPT for your workplace. I was there about six and a half years as one of their earliest engineers, as we went from ten people to over eleven hundred people and work with some of the
- 3:39
largest companies all over the world. Um, but that j- journey was amazing. I love that product. I love that company. I love the people there. Um, but I've actually always really been interested in healthcare.
- 3:50
I remember a decade ago, cover of Nature magazine was AI to detect skin cancer. I was like, "Wow, is someone interested in AI?" That was amazing. At the same time, I went to the hospital for something, and I remember seeing, oh, coming back, it was a really minor thing, but I remember coming back and looking at the
- 4:05
bill, and I was like, "I'm not really sure what exactly I paid for." Um, and so there's these known problems in healthcare, and I'll talk a little about some of them, access to care and cost, and then we had these amazing solutions, so I was like, "Why don't we bridge these things together?"
- 4:18
Um, and remember, this is about a decade ago. Um, and so I actually started this seminar where I invited speakers who were physicians, researchers, entrepreneurs to, to, uh, to talk about the space.
- 4:28
And what I learned was while there was really, really cool technology, very little of it made its way into the clinic at that time. And so again, my, my journey went a different way.
- 4:38
But as I, as I peeked my head out ten years, uh, ten years later then, actually our technology has gotten better than ever, as everyone knows, and the AI wave as it's taken over the whole world has also influenced healthcare.
- 4:50
And as you, as you saw towards the end of that video, Abridge, in the matter of two to three years, got its way into three hundred of the largest health systems, uh, in the United States, Kaiser, Mayo, John Hopkins, Sutter, and so forth, and maybe, maybe you've visited some of these hospital systems.
- 5:06
And once you're inside the hospital systems, you realize there's so, so much more you can do, and I'll talk about that journey that we've had. I specifically work on, um, lead our engineering teams for clinical decision support and our agentic experiences that the technology has now enabled and how we can bring that to healthcare.
- 5:25
But first, maybe, maybe some of the problems that inspire us as a company at Abridge. One, one of the things that we've noticed, uh, or many people, econo- economists have noticed over the past few decades is actually in many other industries, you actually see the cost of a good go down, and that's because the productivity has increased.
- 5:43
But in healthcare, we actually see administrative costs have only gone up over the past, uh, few decades, um, and productivity hasn't necessarily increased. Um, and a lot of our problems in healthcare we solve with labor, but even that we cannot actually keep up.
- 5:59
So there's this bit of this, like, productivity pa- uh, paradox you might have heard of, like Baumol's cost disease, and it's part of-- partially because technology, I think, hasn't fully touched healthcare as much as it's touched other industries to increase that productivity.
- 6:12
A few other problems, you know, hospitals are shutting down, margins are actually razor thin for many health systems. Of course, some patients have massive, uh, medical, uh, debt. And then we actually-- and the problem that's-- another problem that's very near and dear to our heart is that we hear all the time that doctors are burnt out, and
- 6:30
they actually often don't recommend it as a profession to, uh, to their children.
- 6:36
So what we started as, as a company was working on clinical documentation. So the idea here, if you're not familiar with it, is at, at the end of every patient visit, the, the clinician must create a note.
- 6:48
Uh, our typical format is a soap note that has a couple different formats, like what's the chief complaint of the patient and a few other sections, and then what's the assessment plan?
- 6:57
What do we do with this patient? You have to do-- write this after every single visit, and there's some different variations on this de-depending on specialty. Um, typically clinicians end up often doing-- it takes like two hours a day to write, just write these notes, and you often do it what's known as pajama time after work itself.
- 7:16
Uh, and that's a common source of clinician burnout, spending all this time outside of work, and it's not the most fun part of the job. However, these documents are actually extremely high stakes because these clinical notes are often used as a basis of billing, but also, uh, which is of course, financial things are high stakes, but also
- 7:33
have clinical impact. And the reason for this is because these prior no- these notes are used for the next clinician, or as you switch health systems, they use-- they provide context to the clinician of the, uh, patient's longitudinal medical record.
- 7:47
So it's actually really high stakes to get this right. Um, we started here because we-- it's a known pro-- uh, pain point that's existed for many, many years. But finally, the technology's caught up to do really, really high quality medical notes that's actually personalized to the clinician.
- 8:03
This was an amazing wedge into healthcare industry, which has typically been technology reticent because it re-led to... They actually care a lot about getting these notes high quality and right.
- 8:13
It led to higher doctor satisf-- uh, sas- provider satisfaction and such, they could actually see more patients. Um, and it can actually help create a higher record, um, that helps prevent, uh, as it relates to billing, auditing and other reasons.
- 8:28
So we started there. Just this product alone scaled to three hundred ho-hospital systems. But I wanna show you a little about where we're going next. And, um, and the core thesis of the company is that every area in healthcare is, everything is around the conversation.
- 8:42
So we started over here with the co-co-conversation to clinical note. Everything else is downstream of that, whether you-- it relates to billing, whether it relates to things like cl-clinical trial matching or whether it relates to clinical decision support.
- 8:57
It's all about the conversation, that sacred doctor and patient conversation, and we've just built all this administrative machinery around that. But how can we bring it back to that conversation and actually automate some of that, uh, administrative machinery?
- 9:13
So to give you a tactical example of what this looks like, um, and I'll play this video. [coughing]
- 9:26
Of where, where we're going from here.
- 9:28
We've been building a solution that allows the physician to interact with Abridge directly by using their voice. Hey, Abridge, [beep] is Nathan eligible for any clinical trials?
- 9:41
He may be eligible for the Abridge HF study. He remains symptomatic despite maximal therapy, and most screening criteria are already met. But an updated echocardiogram is needed to confirm his ejection fraction and complete eligibility assessment.
- 9:53
All right. Please order that echo for him.
- 9:56
Done. Confirmatory echo ordered.
- 9:58
And then when I'm done for the day, I can just ask Abridge, "Hey, Abridge, [beep] can you prepare my charts for tomorrow?"
- 10:06
And Abridge is working for me.
- 10:10
Pause it right there. One of the things that you'll, you'll notice is that we are thinking about how to, uh, revolutionize the entire visit, uh, for a clinician. From pre-visit, how, um, earlier in the-- we have suggested discussion topics.
- 10:23
Here's things that you can talk about, uh, with your patient, whether they're clinical or more, uh, billing related. Um, we have, after the visit, we actually create everything for you.
- 10:32
The patient visit summary, the actual clinical note, and we actually penned orders, as you might have seen. We're able to use-- W-we are able to do this all by reading all this context.
- 10:42
We have access to all of the EHR context, so we know everything about the patient, the, the prior labs, the prior notes. We have access to the live conversation between the doctor and the patient.
- 10:52
That's where the quote unquote debugging happens in healthcare, where you learn about what the patient is facing now. Um, and then we have access to world's medical literature that we can ground and clinical guidelines that we can ground all of our work in.
- 11:05
So, do, do, do. I wanna, I wanna s-switch now, given the context of where we're going as a, as a product, I wanna switch into some of the key technical and engineering problems we face and inspire you on some of the, what I think are f-very much frontier AI challenges.
- 11:27
So if you've ever worked on a agentic product before, the-- regardless of vertical, um, the three KPIs that tend to matter are quality and latency and cost. In healthcare, I feel that we're actually playing on hard mode for all of these three KPIs.
- 11:42
This is a high stakes scenario, especially when you're doing something like clinical decision support. You have to be right because the downside is extremely high when you're wrong. When I used to work at Glean, you know, while I love that product, I could be wrong and it would have been fine.
- 11:55
Maybe we answered a question incorrectly. But in healthcare, if we answer something incorrectly, there's actually consequences, and we entirely lose our trust. So quality needs to be absolutely high, and I'll talk a little about how we keep that bar high.
- 12:07
And then latency and cost also really matter for us when you're live in the conversation. You can't, uh, with latency, you can't act on information too late, and you have to act on the-- also at the right time for it to be useful.
- 12:20
And then finally, cost at the scale we're, we're doing this at.
- 12:25
Um, and, and, and as an interesting aside, it actually relates to, uh, our-- We have a motto inside the company that our goal is to save lives, save time, save money for, uh, for the hospital system and for the healthcare industry as a whole.
- 12:36
And it actually-- I think it's funny that it really maps to the three K-KPIs that you care about in any agentic product.
- 12:44
So talking a little about quality, how do we keep that bar high? For us, we really treat evals as the l-operating system, the life's blood of the, of the company.
- 12:54
This starts from internal benchmarks and offline evaluation. Before we develop any product, you know, whether we're talking about clinical trial matching, clinical note, clinical decision support, coding, we start with a robust set of internal benchmarks.
- 13:08
This is pre-deployment, and then we test that against, you know, things that we've actually seen in the wild. Then we al-- have a staged rollout. We know that we need to, uh, make contact with realit-reality.
- 13:20
Not everything offline will perfectly represent what happens in practice, and so we slowly roll it out. Uh, maybe it starts with the alpha s- group of clinicians that we trust and under-- they understand the stakes.
- 13:31
We roll out to beta. Maybe there's AB testing at scale. And then even after it's fully rolled out, you always need continual, uh, monitoring. Again, the stakes are really high, and you cannot get away with just being, like, a prototype that you just ship out there and be like, "Yeah, I mean, I tested it on a few
- 13:45
cases, and it works." How we do this is we always have expert-calibrated LLM judges. So we have clinicians embedded throughout the entire company. The clinicians are d-domain experts, but not all of us are clinicians.
- 13:58
I'm not a clinician. So how can we, the rest of the company, still move fast is by encoding that clinician judgment into LLM judges. You know, I think a really great evaluation system has a property that it reflects the behaviors that you want in your product.
- 14:14
At the end of the day, we are making a product for clinicians, and so who best other than our clinicians to actually create our judges that represent what they want?
- 14:21
And those judges, once you have that, create a feedback loop so that anyone, whether you're a clinician or not, can actually, uh, hill climb and learn from that. We also have a lot of online signals, whether how you're editing the clinical note and your typical thumbs up, thumbs down, uh, and star ratings and other free-form text.
- 14:40
So this is a general framework we use for our, for all our products. I wanna deep dive into the product that I work on, which is clinical decision support.
- 14:50
So to give you an example of what clinical decision support looks like, and specifically, we're building something novel, which is contextual clinical decision support. Maybe a provider asks a question like, "Hey, does this patient meet the criteria for febrile neutrophenia?"
- 15:04
Um, and so what we have to do here is actually a lot of context is underspecified in this question. We're-- So the first thing we're gonna do is we're actually gonna pull from the EHR data previous, uh, previous lab values.
- 15:17
Then using that context, we're gonna, uh, use the, uh, c-- uh, we're also gonna use the live conversation, and we're gonna use clinical guidelines and medical journals to com-- use as the reasoning sources for combining all this context to actually answer the provider's question.
- 15:34
Now-- But I wanna focus on evaluation. Again, the stakes are really high here. We, we really can't get this wrong. So how do you tell whether or not an answer is correct?
- 15:42
And sometimes I, I f- uh-- And this is a case where the generator and the verifier gap is really small. What I mean by this is, in some problems in AI, such as like Sudoku, it's really, really hard to generate a solution to Sudoku, but it's extremely easy to verify it, uh, once you do have the solution.
- 16:00
And that makes, uh, that makes it much easier to hill climb against, uh, and build evaluation for. But in a case like this, the generator and verifier gap is really small.
- 16:08
If I had a really, really good generator, uh, verifier, then that would just be my generator itself. So how do I create a reference that isn't just a language model itself and ground itself so I have trust?
- 16:21
So what we do is we tackle this by having many, many different signals. Uh, we have a clinical quality judge, which I'm gonna dive deep into, and then we tackle from, uh, we have many signals from a boundary and adversarial judge.
- 16:33
We have a clinical safety judge. And we also have judges that re-represent product as-aspects like tone and style, and that matters a lot as well for AI products. So all of these are different signals that try to get a piece of this, like, really, really hard-to-measure problem and guarantee it in the way that we want the product
- 16:49
to be. So diving into the clinical quality judge. So again, I said the v-- generator–verifier gap is really small here. So what we need is we actually need human references to tell are, are we generating the right thing?
- 16:59
But you can't just create a human golden response because there is a lot of variability in the potential responses. So what we did is we took a lot of real clinical cases.
- 17:09
We had independent physicians create a rubric. So this rubric said elements of what we wanted in the response. So it's not, "Here's the exact response," because again, there's many infinite possible responses, but a re-- a good rubric elements that what a good r-response would look like.
- 17:22
And then we had a separate physician that actually adjudicated it, brought these two independent rubrics together, created a final rubric, and we actually had a fourth clinician do QA on these rubrics.
- 17:32
Once you have these rubrics, and here, here's a sample rubric, what it looks like. Here's actually a question. And then, I mean, there, there's more context in the case itself.
- 17:38
You have a rubric of what are the elements that a response should look like. Now, we can actually have an LLM judge that compares our agent's responses to these rubric elements and does some s-- uh, basic semantic match to tell, hey, is our model performing well as we continue to hill climb, whether it's our agent architecture, our
- 17:55
models, or search ranking algorithms. I wanna talk a little now about cost and latency, two other really, really hard problems for us. So we do this, as we said in that intro video, we do this on the-- live in the conversation, and we do on the run rate of a hundred million medical conversations a year.
- 18:13
How do we do this in a way that doesn't really break the bank for us? So one, one place that this problem comes up, uh, or is, is actually in generating the clinical note.
- 18:22
So when you're generating the clinical note, there's many different sections to it. There's a history of p- uh, present illness, past medical history, and there's the assessment and plan.
- 18:30
So one of the core insights for us is, rather than, say, using a foundation model to generate all of this, is we can actually break, decompose this problem into simpler, smaller workflows.
- 18:41
Healthcare is actually many specific workflows. You don't need, you know, Fable Five, uh, to actually solve all of your, uh, clinical notes. We, we don't need frontier-level intelligence for every problem, so we actually post-train a lot of smaller models for different problems, such as different-- Actually, even to the granularity of different sections in the clinical note.
- 19:00
And that lets us use much smaller mo- uh, models because it's a m- more specific problem and at, uh, at much cheaper cost and latency. And we have this data flywheel that we have this unique dataset of a hundred million medical conversations a year.
- 19:14
And as far as we know, no one else has such a h- uh, large dataset. So our key insight is having a right to tr- win in training models.
- 19:22
There are problems where the quality's already maxed out, and so you should train models then to reduce quality and latency. But there are other problems where the quality isn't maxed out, and people say, "Oh, the frontier model would just steamroll you."
- 19:33
Our key insight is we can actually potentially beat the rate of change on the frontier model if we have the right, uh, to win by having the right data that they may not have and the focus on a problem that they may not be focusing on, and that lets us still maximize quality.
- 19:49
Another, uh, problem that I'll quickly touch on is in-visit orders. So, uh, do- doctors really aren't big fans of pending orders, but often they'll mention orders during the visit itself, medication or non-medication orders.
- 20:01
So what we have is while we're listening in the visit, as the, uh, clinician says, uh, order, we actually queue it up in the background, uh, le- and let them, uh, actually sign it off in the EHR.
- 20:11
But you can imagine if we did this in a very naive way, like every few seconds are just listening for orders, that would really break the bank. Um, and so a lot of our tricks are like, how do we find the right events in the conversation to actually trigger heavier models that will actually do the order matching?
- 20:26
Because you need to match the order, not just is the order said, but does it match and reference the system's orders that are approved by the system and are relevant to the conversation.
- 20:34
So we have a number of different gates that are cheaper and faster that let us trigger actually larger models and hand off to them for actually doing the end-to-end work.
- 20:46
Um, but the last message, this was of course a very quick talk, but the message I wanna leave you with is healthcare is a domain that needs frontier AI and actually puts it to the test at, uh, higher stakes.
- 20:56
In the past, I as an engineer myself was wary about working in healthcare. Does, uh, does, like, does healthcare technology actually work? Well, Abridge has proven this at scale for sure, and I was-- hopefully I gave you a taste of some of the frontier problems that we work on.
- 21:10
So thank you so much. My name is Chaitanya again. You can, uh... And feel free to connect with me at Twitter or LinkedIn. Thank you. [audience applauding] [upbeat music]