AI Engineer World's Fair 2025
Case Study + Deep Dive: Telemedicine Support Agents with LangGraph/MCP
About this talk
Stride AI practice lead Dan Mason presents an interactive workshop on an agentic telemedicine system developed with Avila Science to support patients through multi-day, at-home treatment while preserving clinician relationships and human escalation. He explains configurable treatment blueprints, persistent patient state, messaging, adaptive schedules, virtual operations associates, LangGraph orchestration, MCP-connected tools within a VPC, and evaluation and response-validation techniques.
Chapters
- 0:15Workshop introduction, Stride, and the Avila telemedicine case study
- 12:33Treatment workflows, virtual operations associates, and LangGraph state
- 26:56Messaging, treatment timing, and resumable patient journeys
- 50:59Blueprint integration, agent judgment, and patient-context handling
- 1:21:50Scoring prompts, MCP gateways, and VPC-local tools
- 1:37:39Prompt inspection, response validation, evaluation, and closing
Talk transcript
- 0:00
[upbeat music] Okay.
- 0:15
Um, hey, everybody. Thank you so much for coming. Uh, really appreciate you being here. Um, this, this is a great show. I love this show. Um, I was here last year as an attendee, um, spoke in New York, uh, at the, the New York summit in February, and I'm, I'm really thrilled to be back.
- 0:28
Um, so this is very much a show and tell. I think I, I, I said this in the Slack channel, so if anybody's not in the Slack channel, feel free to join it.
- 0:35
Um, there's a couple links in there that might be helpful to you. Um, it is workshop LangGraph MCP agents, if anybody needs that. Um, but fundamentally, uh, I'm, I'm just here to walk through some, some really interesting work, um, that my team's been doing around, um, building agent workflows for, uh, a healthcare use case.
- 0:53
Um, and, uh, this is, this is very much, like, the way we did it. Um, I'll get into some details about that. It's not the only way to do it.
- 0:59
Um, and I'm hopeful that somebody in this audience might look at this and be like, "That's dumb. You should do that better," and, you know, please raise your hand and tell me.
- 1:07
Um, but, uh, but it's been really fun to build. I'm really, really happy with the results and, uh, really excited to show you guys, um, what it's all about.
- 1:14
Um, okay, so I will go into presentation mode.
- 1:21
All right, here we go. Okay, uh, first just a very quick couple things about Stride. Um, that's me if anybody needs my LinkedIn, but, um, there's al- a couple other places you can find that.
- 1:30
Um, we are a custom software consultancy. Um, so what that means in practice is, uh, whatever you need, we'll build it. We have been doing a whole lot of AI stuff.
- 1:40
Um, this kind of falls into a few specific buckets. Um, we use a lot of AI for code generation. Um, we, we have a couple of both products and services we've built to do things like, uh, unit test creation and, and maintenance.
- 1:52
We've done a bunch of stuff around, uh, modernization of, of super old, dumb code bases. Um, you know, things like, uh, you know, early 2000s era .NET is one of the things we specialize in.
- 2:03
Um, but what I'm gonna show you today is really, uh, what we do around agent workflows. So the, the idea with, you know, this agent workflow stuff is really just that, you know, it's something that could be done with traditional software, and, and this thing I'm gonna show you was done with traditional software in its first run.
- 2:18
Um, but we have rebuilt it with an LLM at the core to make it more flexible, um, more capable and, you know, ultimately, um, just a lot cooler. Um, so, uh, really excited to, to show you guys more about that.
- 2:29
So I'm gonna start with a little bit of grounding. Um, I'm gonna do a case study, um, which is very brief, but it'll give you a sense of kind of the problem we were trying to solve and, and, and how well I think we solved it.
- 2:39
Um, and then we'll go as, as deep as we all wanna go in terms of, of how it works. Um, let me ask this up front. If you have questions, please raise your hand.
- 2:47
I will try to notice, and I'll try to get to you. Um, there are some mics we could pass around, um, but it probably is better for you just to shout it out, and then I'll repeat it, um, into, into my mic.
- 2:55
Um, and, and fundamentally, like, I have no idea if this is two hours' worth of material. It probably is, but, you know, please, uh, keep me honest. Um, I'll talk about anything that is relevant to this that you guys wanna talk about.
- 3:06
So, um, the client here, uh, is Avila. So Avila Science is, uh, a [REDACTED:gender]s health, um, uh, sort of institution, which is trying to help with, uh, the treatment of early pregnancy loss, um, otherwise usually known as miscarriage.
- 3:21
Um, what, what this is though is specifically a treatment where what happens is that, you know, you experience the event, you end up at the hospital or at a clinic.
- 3:28
They send you home with medicine, right? [coughs] The medicine is something that then you have to administer yourself, um, at both a very traumatic time for you and your family, at a time when, you know, you need to keep track of when you s- are supposed to do things.
- 3:40
It can be really challenging. Um, there's other use cases beyond this in terms of chemotherapy, when people have trouble remembering what day it is, you know, let alone what they're supposed to be doing, right?
- 3:48
You know, there, there's a variety of treatments that this is relevant for. [coughs] Um, but Avila in particular has a system that they use to help people es- essentially administer these telemedicine regimes at home, right?
- 3:58
And, um, that system is text message-based. So everything here I'm gonna show you is essentially a text messaging-based, uh, engine with, you know, some, some core business logic, um, that helps people stay on track.
- 4:10
It answers their questions. It checks in on them, you know, to make sure that the treatment went well. They still have a doctor relationship. This isn't replacing the doctor.
- 4:17
It is simply helping to get people through this treatment without a doctor's direct support, uh, at least a lot of the time. Um, so a few disclaimers up front.
- 4:25
First of all, I'm gonna show you a whole bunch of stuff here that is, is the, is the client's actual code. Um, thank you so much to my client.
- 4:31
Thanks to Avila for being so open with this. It's really awesome. I'm really happy to be able to show you as much as I'm gonna show you. Um, I have redacted a few things.
- 4:37
Um, I think what's left is still, you know, very much gonna give you the character of the whole thing and, and an idea of how it works.
- 4:43
Um, Stride, we are custom software people, so we built custom software. I, I don't wanna hide that part, right? It's possible to do a lot of this stuff with off-the-shelf tools, um, but there were some specific requirements this client had that made it better, frankly, to build a lot of it custom.
- 4:58
Um, so we did. But we did the best we could to use, you know, big swaths of off-the-shelf, right? So you'll see a lot of LangGraph, LangChain, LangSmith, you know, a bunch of other things like that, you know, very much in here because we do believe that that adds value and that it fundamentally makes the system a,
- 5:11
a lot more explainable. Um, yeah, and you know, there are also some constraints in terms of how it's hosted. This is, you know, at least partially intersecting with, with patient data and various things like HIPAA and, and other privacy requirements.
- 5:23
Um, the other thing, as I, I started with earlier, there's no right way to do this, but this one does work for us. And, and I think you'll see as we walk through it some of the choices we made.
- 5:32
Um, you know, there, there are definitely other ways we could have plugged the tools together. I think there's definitely other ways we could have done this workflow. Um, but, but we like how this came out.
- 5:39
It, it preserved some of the things that we really knew were important to our client and that kind of reserve, um, you know, a lot of human judgment, you know, as opposed to sort of taking the LLMs entirely at their word.
- 5:48
And this is very much a hybrid system with humans very much in the loop. Um, and again, as I mentioned before, um, I would really love it if you guys looked at what we're doing here and said, "That's dumb."
- 6:00
That-- Or, or, "Have you thought about this?" Right? Because A, you know, this is a project which we've only been working on for a few months, but, you know, things have already evolved.
- 6:07
That's the way it is in, in AI. [laughs] So, um, I'm certain, and I know of a handful of things where, you know, we could replace some of the choices we made with newer, more modern choices.
- 6:16
Um, and at the same time, you know, there may be cases where, you know, I, I'm genuinely not using LangGraph right. I, I would love if someone raises their hand and tells me that.
- 6:23
So please do, uh, have that in the back of your brains.
- 6:27
Cool. Um, really briefly on the, the stack that we used and, and on the team that we built. So the first thing, again, there's a lot of LangChain in here.
- 6:35
Um, that is not because other frameworks can't do this. It's not because we couldn't build our own. The number one reason we went with this is because of how easy it is to explain the system to other people, right?
- 6:44
You know, if, if you look at... And I'll show you the LangGraph stuff in particular. It, it was straightforward to go into our client, you know, on, on a very early day in the project and say, "Hey, this is how this thing works.
- 6:53
You can see it goes from here to here. There's loops here. Like, this is where we're doing our, um, you know, our, our evaluation of the process, and here's where humans come in."
- 7:01
It was very straightforward to do that. Um, and I think it would've been a lot harder with something that was less visual and, and frankly, just less well-orchestrated. So, so we're happy with this, right?
- 7:09
There, there are some trade-offs to the, the, the LangChain tools, but they're, they're mostly things we can live with. Um, we are using, uh, Claude in the examples that I'm gonna show you here.
- 7:18
Um, but the, the core code that we wrote works with Gemini, works with OpenAI. Um, there are a few reasons that we think Claude is, is better for this.
- 7:25
I'll get into that as we go. Um, but you know, there's, there's no model-specific stuff really happening here. This is almost all just tool calling and MCP and, you know, other things that are, are pretty portable across most of the models.
- 7:36
Um, the stack overall is not just the LLM piece, right? So the LLM piece is Python in a LangGraph container. Um, and then the, the other piece, right, the piece that is a text message gateway and a database and a, a, a dashboard, which I'm gonna show you pretty extensively, is Node and React and MongoDB and Twilio.
- 7:54
Um, and the whole thing is hosted in AWS. None, none of that has to be that way. That's just what we picked. Um, you know, the, the main reason we picked AWS was for, you know, th- this has to support multiple different regions.
- 8:03
We had to be able to deploy stuff, you know, entirely in Europe in a couple of cases, right? And so we needed to make sure that we had, you know, a, a decent set of, of, you know, cloud connections that we could work with.
- 8:14
Uh, evals. So I will show you the, the eval system that we built. Um, we were not able to use, or at least... I shouldn't say not able. We chose not to use the stuff, um, entirely off the shelf from LangSmith.
- 8:25
This is partly because I didn't really wanna be fully locked into them. I wanted the data to live there. I wanted to be able to see, you know, the current system in LangSmith.
- 8:32
Um, but I wanted to have something separate. And it turns out that some of what we had to do to make the evals, um, you know, fundamentally, uh, functional required a lot of pre-processing.
- 8:40
So we built an external harness that essentially pulls data out of LangSmith, processes it, and then runs things through Promptfu. Um, and one of the reasons we picked Promptfu, if anyone's ever worked with it, um, they have a very flexible, uh, they call it an LLM rubric.
- 8:54
And, and so this is an LLM-as-a-judge. You basically describe how you want the, the eval to work. Um, you feed the data in, and you know, then it gives you a separate sort of visualization for that.
- 9:02
So we, we ended up very happy with it. It's not the only way to do it at all. It was definitely, you know, the, the, the thing that fit best for, for us.
- 9:10
Uh, the team. So there were, uh, and, and still are, um, two software engineers, one designer, um, and me. And, and I'm just-- I, I would not call myself a software engineer, that's why I didn't include myself in that pool.
- 9:21
You can imagine there, there being two software engineers kind of maintaining the core system that has the, the gateway and the, um, dashboard and the text message stuff, right?
- 9:30
And, and the database. Um, I maintained and, and built basically everything on the, the LangGraph side, right? So imagine this as being two separate systems that talk to each other through a well-defined contract.
- 9:41
Um, and that two, that, those two software engineers understand roughly how my code works, but they really weren't maintaining it. You know, it was, it was almost entirely me with, uh, AI friends.
- 9:50
Um, and on that note, so everything I'm gonna show you is the code that I wrote. And, and I wanna be very clear, [laughs] um, I haven't been a real software engineer in a long time.
- 9:58
I do have an engineering background. I spent seven years out of college, you know, hacking on mobile apps.
- 10:03
Um, I took fifteen years off and went to be a product person, and for about two years now, I've been back. But what that really means is just that, you know, essentially the stuff you're seeing, right, or the stuff that I'm gonna show you is mostly, you know, code that I wrote with Cline.
- 10:16
That's, that's my personal favorite. Um, and so there's a bunch of options here. I like Cline best of all these options. Um, you can use anything you want. The code isn't actually that complicated.
- 10:25
Like, I would estimate, and I, I haven't actually counted, but there's probably a few thousand lines of Python, and there's a few thousand lines of prompt. It's, it's about equal, right?
- 10:33
So I vibe coded the Python, and I mostly hand-coded the prompt. Um, not a hundred percent, right? But, but that's, that's the way to think about the division of labor here.
- 10:42
Um, and for that matter, just, I mean, any of these tools can be great. The main reason I picked Cline was just because, you know, we did not need, uh, you know, sort of a hyper-optimized, you know, um, like twenty dollars a month flow.
- 10:53
Like, I've spent a lot more than twenty dollars a month on tokens. That's, that's just the way it is. Um, you know, the, it was worth spending the money to just have sort of the best available context of the model at any given point.
- 11:01
Um, Cline is a very good way to do that.
- 11:04
Okay. And there's a little bit of, of sample code. So I did mention, and this is in the Slack channel as well, if you wanted to follow along with any of this, you could sort of do it by standing up your own little LangGraph container with MCP.
- 11:14
You're more than welcome to do that. Um, everything I'm gonna show you, though, is proprietary client code, so I, I obviously can't send you those links. So if you'd like to, um, feel free to fire it up.
- 11:23
Um, we have two hours, which is a really long period of time. If, if you're interested in spending a little time at the end of this actually working with some of this real code, I'm, I'm thrilled to do that.
- 11:31
Um, so feel free to get yourself, uh, ready in the meantime.
- 11:35
Okay. Just a couple things up front just to make sure we're level set in terms of, of sort of the terms and kind of the way that we're talking about this stuff.
- 11:42
So, um, I do like the LangChain definition here of, of agent, um, basically just because, you know, you'll see what we're doing here is using an LLM to control the con- the, to control the flow, right, of, of this application.
- 11:55
That is literally what, what this is. Um, and I like this, and, uh, I don't know if Chris is here. He, he was at, uh, the last event in New York.
- 12:02
I like this as a way of, of sort of justifying the way that we tried to architect this system and, and why, right? So the idea of, of a-agents in production, right?
- 12:12
You have to know what they're doing, you have to know, you know, that they can do it, and you have to be able to steer, right? If you only have a couple of these things, you end up with, with bad outcomes, right?
- 12:22
And so I, I, I just like this framing of, if you're capable but you can't tell what it's doing, it's dangerous. If you know exactly what it's doing but you can't control it, it, it does weird stuff and you can't help.
- 12:30
Uh, please, go ahead.
- 12:31
A bunch of us are looking for the Slack channel.
- 12:33
Oh, uh, let me find that one more... It's right here, actually. Workshop LangGraph MCP agents. Got it? Okay.
- 12:42
Thanks.
- 12:42
No problem. Okay, but, uh, and so transparency with no control is frustrating, and control with no capability is useless. I, I just love this framing. I think this is exactly the thing that we were trying to solve for.
- 12:51
We needed something that was able to do the job, clear about what it was doing, and that was steerable by humans in a really obvious way. So, um, with that said, I'm gonna start with a case study.
- 13:02
Right? And this is gonna be a little weird out of context, but hopefully this will [laughs] give you a, a sense of, of what we were trying to solve for.
- 13:07
So the idea here was that there's, there's an existing product, right? So there was a product out there that was essentially having, uh, you know, humans manually push buttons on a console that would enable, uh, a text message to go out, right?
- 13:20
So you would read what the patient had said. They could say, "I took my medicine at 3:00 PM." They could say, "I'm bleeding and I don't know what's going on."
- 13:27
Like, "Am I okay?" Um, they could ask other sorts of questions about the treatment. And a human would have to go into a piece of software and click, you know, a button that accurately reflected sort of where in the workflow somebody was, right?
- 13:39
Because, you know, you can model a lot of this out. You know, imagine there being fantastically complicated flowcharts of all the things that can happen during a medical treatment.
- 13:47
Um, so the Avila team had built this, right? They realized, though, that essentially to scale the human team to be able to serve a lot more patients w- was prohibitive, right?
- 13:54
They, they needed too many people clicking too many buttons. They also realized they couldn't really scale the system to new treatments, right? Which was something they wanted to do, that this isn't the only regimen that you needed to support.
- 14:04
They had other ones. Um, and so the idea is that either they were gonna rebuild the legacy software to be more flexible, or they were gonna essentially rebuild it to, to, to use a, a different kind of decisioning at the core.
- 14:15
And, and when they were looking at doing this, you know, LLMs had, had started, I think, become capable enough to, to handle this kind of, of work. Um, so what we built, what we did is we built for them a workflow and essentially a piece of software that connects to it that enabled them to do new treatments
- 14:31
flexibly, right? So this idea of essentially defining a blueprint and a knowledge base is the way that we, we thought about this. Um, and, and essentially medically approved language, right?
- 14:39
So one of the reasons that you had humans pressing buttons instead of typing text messages [coughs] is because this is medical advice, right? You know, you, you are not, um, you should not at least be giving medical advice, um, that differs substantially from, from this approved language, right?
- 14:51
There's reasons that this stuff, you know, is said the way that it's said. Um, you know, and, and doctors have, you know, similar limitations. Um, we also built a self-evaluation function, which I'll go into tremendously, um, in a second.
- 15:03
Uh, it-- we wanted to make sure that we caught essentially situations that were complicated, um, and surfaced them for humans, right? Because we wanted to have a human in the loop.
- 15:11
But we were trying to raise up the existing folks who were really just operating the system and, and clicking all those buttons to be supervisors of, uh, of agents that were doing that instead, right?
- 15:19
That, that really was the, the model that we were working with at its core. Um, saw a question over here. Yeah.
- 15:24
And you may have said it. Were these operators medically trained?
- 15:27
Uh, so the question is, are these operators medically trained? There is a physician's assistant who essentially leads the operations team. So the way that you can think about it is that, um, she would be escalated to whenever something came up that was outside of the blueprint, right?
- 15:40
So if you had a situation where they're just like, "I'm really not sure what to do here," a Slack goes out to that channel with the physician's assistant in it, who would then give medical advice.
- 15:47
So, you know, a- again, this is one of the reasons that it was hard to scale, right? Because, you know, you only had one of those people on this particular team.
- 15:53
Thank you.
- 15:53
Sure. Um, and so, you know, to, to sort of jump a little bit ahead, but hopefully you'll see why this is in a minute, this roughly, and, and again, we're, we're still doing the measurement, right?
- 16:03
We're still trying to figure out exactly what, you know, capacity has, has gone up to. Um, we think it's something like 10X. We think that they can surface roughly 10X more people with this new approach.
- 16:14
Um, now it's not free, right? We have to build the software. We have to pay for the tokens. Um, tokens can get expensive. But if you think about, you know, just the scale issues involved in scaling up a team of people and again in building the software to be more flexible for more treatments, um, we think this
- 16:26
capacity increase is, is very, very much warranted and very much the thing that, that, you know, solves, solves the problem. Um, and you can do new treatments and new workflows without writing more code, right?
- 16:36
That was the single biggest thing about this, and, and you'll see what we're doing here is largely Google Docs, right? And, you know, we have some more advanced techniques to, to manage those things and in version over time.
- 16:44
But, but we're talking about being able to support whole new treatments and whole new workflows without going back to the code, right? That's hugely valuable to these guys.
- 16:52
Question.
- 16:53
Um, you mentioned velocity increasing like 10X. Is there any measurement about quality of care?
- 16:59
So question was, uh, velocity increases, is there a quality of care measure? Um, short answer is it's early, right? I mean, this is still a system that's, you know, in progress.
- 17:07
It is being used with real people, but it's still very, very much early on that. The way that I think we're looking at it is that there would be some combination of the operators being the ultimate arbiter, right?
- 17:16
They're gonna be able to see these conversations and determine and as they approve them, you know, as they review them, like, "Hey, is this mostly getting it right?" And then there are, there are sort of existing kind of CSAT, you know, level measures that you can apply to the people who are on the other end of the
- 17:27
treatment.
- 17:29
So 10X sounds a little low. Is that because the operators are still approving everything that comes out right now?
- 17:33
Uh, so they're not approving everything that comes out. A- and I agree, the 10X is kind of, it's an order of magnitude, not a, a precise measure, right? But I think in this case, you'll see a couple of cases that require approval, right?
- 17:43
And sort of why. Um, but the approval also is very quick, right? So I... The, the argument is that you probably only see one of every ten exchanges, and when you see it, it takes you roughly as long as it took the last time to just push the button, right?
- 17:54
Which, which was the thing that they were already doing. So that's kind of why we've benchmarked it there.
- 17:59
All right. Um, so let's get into it a little bit. So, uh, th- this is just a snapshot of, of what this looks like in LangGraph. I'll show you the real thing in just a minute, and it's actually evolved a tiny bit since I took this picture.
- 18:10
Um, but really what we're talking about here is, um, the, the, the people who operate the system today, we call them operations associates. So what this is really doing is introducing a virtual operations associate.
- 18:20
That operations associate is going to assess the state of essentially a, a conversation interaction with a patient, um, determine what the best response is, both in terms of, uh, the text message you might send, um, the questions you might ask, the actions you might take, because some of this is about maintaining essentially a state for that patient,
- 18:38
right? You know, you are, you are at any given point trying to figure out, um, when is this person taking their medicine? When did they take their medicine? Um, you know, what medicine do they have?
- 18:47
Um, what time is it, um, for them? Which is actually more important than you, than you may think. Um, all of this has to be maintained, right? By the system.
- 18:53
And so the virtual lawyer is doing all of that work, and then it's passing essentially its proposal, right? It, it, it basically comes up with, "I think this is what we should do," and it passes it to an evaluator agent.
- 19:03
So there's a live LLM-as-a-judge process separate from the evals, which, which we'll get to. But the live LLM as a judge is essentially saying, "Okay, given this thing that just happened, um, here is our assessment of A," you know, how right the LLM thinks it is.
- 19:17
Um, that's frankly very challenging. LLMs are very hard to convince that they're wrong about anything. But, um, it also is looking at the complexity, right? So even if the LLM believes it's made all the right decisions, you can have it impartially say, "Well, I changed this and I changed that, and I'm scheduling a bunch of messages.
- 19:31
That's complicated. Maybe a human should look at this." Right? So that's actually a lot easier to, to implement. Um, and both of these things are calling tools. The tools are a mix of MCP, um, and so there, there's sort of two versions of MCP here.
- 19:45
I'm gonna show you one which is basically just looking at local files, just so I can show you all the stuff in my environment. Um, but there's also MCP going across the wire to the, the, the larger software system and keeping all this stuff in the database, right?
- 19:56
So there's, there's a mix of those two things. Um, and the rest of the tools are about maintaining state because as a conversation is happening, the, the LLM needs to know, you know, essentially, "Well, I made this update and that update, and here's the current state that I'm working with," and it has to be able to sort
- 20:11
of m-manipulate these things in real time. That is not MCP. That's not going to a database anywhere. Like, this is happening entirely in sort of the live thread, and then once it finishes, then it gets pushed out and, and essentially saved away.
- 20:23
Okay, um, again, we'll get into a lot more of that. I did wanna spend a minute on the s- on the system architecture, right? And so I realize it's a little small.
- 20:30
Go for it.
- 20:31
When you mentioned about the system state, I, I heard before that LangGraph has a context or some state object built into LangGraph itself. You mentioned that you use tools.
- 20:41
Are you talking about separate things or the same thing?
- 20:44
Uh, question was about how the state is managed in LangGraph.
- 20:46
Yes.
- 20:46
So, um, short answer is this may be one of the things where I'm not, not doing it optimally, by the way. But, um, with LangGraph, there is a state object that we load essentially when the request comes in from a, a JSON blob, right?
- 20:59
We keep it alive inside the, the graph run. It is not directly accessible to the model, right? The, the-- at least not the way that we're doing it, right?
- 21:07
So you'll, you'll see actually as we get into this that you can see all the state coming in in LangSmith, right? I can see, like, hey, this is the whole thing that was, was loaded.
- 21:13
I still have to repeat that in my first message to Claude, right? It doesn't actually show up, um, you know, in the same place. And then I call the functions.
- 21:21
That state will evolve in terms of what's inside the graph run, and then when it outputs, it's the, it's the Python code, not the model, which essentially takes all that state and then, uh, serializes it and, and sends it out.
- 21:31
Um, so you'll, you'll see how it works, but, like, I-- that's generally one of the things that I'm not sure I'm doing right.
- 21:37
Anything else? Yep.
- 21:38
Yeah. Somewhat related. You have one node for that virtual agent that's going back and forth with tools, right, and updating the state.
- 21:44
Yep.
- 21:45
I'm assuming the reason you haven't in some form hard-coded out that business logic more into separate nodes is-
- 21:51
Yep
- 21:51
... precisely 'cause you'll lose the workflow agnostic nature for the next time. Is that sort of the notion there?
- 21:57
Yeah. So the question is why the, essentially the virtual A is one n- one agent and not, you know, a, a sort of a, a pre-coded sort of version of here's how I administer the specific treatment.
- 22:06
Yes, the, the reason I think we kept it simple is because we did not wanna be super treatment specific in how the architecture worked. But I-- you could imagine doing, you know, a set of slightly smaller, you know, better tuned agents that were, you know, um, kind of taking care of elements of the task that was still
- 22:20
pretty generic. The main reason I think it's not optimal to do that is, is caching. Um, and this is another question where, um, you know, I, I think I'm doing this right, but there are a lot of variations here.
- 22:30
Um, caching the entire message stream is easier with, with either one agent or with sort of one agent doing most of the work. Um, we're, we're using Claude. Claude has very explicit caching mechanisms.
- 22:41
Um, and every time I switch the system prompt, I think the cache blows up. And so fundamentally changing the agent, um, identity does that. So that was one-- that's one reason we chose that.
- 22:50
It's, it's certainly not, you know, a, a hard and fast forever choice.
- 22:55
What's the, uh, duration of the, uh, like care? Are we talking like months or-
- 23:03
So-
- 23:03
Like how many messages between-
- 23:05
Yeah. So, uh, this use case, the early pregnancy loss, um, it tends to be a treatment which takes, I think, three days, uh, end to end to administer most of the time, and then there's a check-in after that, right?
- 23:14
So imagine that probably within a week, the entire interaction with that patient is done unless they come back and just have questions later on, right? You know, there, there are some variants of this where you take a pregnancy test after six weeks, right?
- 23:24
And, and so that's all fine. Um, the message history is preserved, but the computation that happens to generate each message is not, or at least not in, not in sort of the state that we save.
- 23:35
So, like, you know, the most complicated, uh, conversation I've seen was something like a hundred and fifty texts. It's a lot in terms of, you know, a human keeping it in their brain.
- 23:43
It's not that bad for an LLM, right? So-- but it's, it's that level.
- 23:49
All right. Um, so, uh, again, just to point out where the lines are here, right? So I s- I kinda got off on a tangent. The top box is what we're gonna be looking at here today, right?
- 23:57
It's really a Python container with access locally to these blueprints, this knowledge base, right? We are also then maintaining some stuff over across the wire in this blue container.
- 24:07
That's really where the, the dashboard I'm gonna show you is. It's where the text message gateway is. Um, and it is where we're gonna be moving, I think, a lot of that context, right?
- 24:13
All the, the blueprints. Like, all that stuff really should live kind of in the more durable software container. Right now it lives, you know, close to, to the Python.
- 24:20
Okay. Um, so let's get into it Um, so the first thing I'll do here is just to show you, uh, kind of at a high level what the software looks like.
- 24:29
So, um, this again is, uh, the, the, the co- the console, the dashboard, right? The thing that, that the operations associates, the humans, are gonna be looking at. Um, and I'll show a couple things here just to, to give you the, the sort of b- uh, baseline, right?
- 24:41
So the first thing here is this Needs Attention. So the current system basically has this Needs Attention flashing all the time. Every time a text message comes in from any patient, th- this thing is going off, right?
- 24:51
You know, so there... And there's, you know, hundreds of patients, thousands of patients in the system at any time. So, you know, this Needs Attention used to be something that multiple people were having to stare at constantly, right?
- 25:00
Just to make sure that they caught everything so that they got messages in a, in a reasonable time. Now, Needs Attention is really, you know, just sort of one thing at a time, right?
- 25:08
And if I look here at the conversations, there we go, um, you can see that the top one here actually needs a response. I'll get to that in a minute.
- 25:14
But at any given point, right, this is my test environment, you know, I've got a handful of these conversations kind of already, already queued up. What I can see here, if I click into these things, is essentially...
- 25:23
I'll just go back to the beginning here for the, the whole message history, and I'm gonna toggle this rationale on. Um, what you're seeing is the entire conversation. Is that readable?
- 25:32
I'll just see if I blow it up a little bit. Is that a little better?
- 25:35
Yeah.
- 25:36
Okay. Um, so the idea here is, uh, the agent is named Ava, right? That's the personality that people are interacting with. Um, this language is all coming out of these blueprints, right?
- 25:47
That I'll show you. And so this first message is just an initial message sent by the system essentially just to kick things off. So imagine someone is-- they have a package of medicine in their hand, they scan a QR code.
- 25:56
They put in their phone number, they get this text message, right? And then they start talking. Um, so you can see here, the kinds of things a patient is gonna say are, you know, free-form text, right?
- 26:06
You know, this, the... I mean, they could say yes in any number of ways. The old system used to have literally different buttons for yes. Like, "Yes, I have the medicine.
- 26:15
Yes, I heard you. Ye-" I mean, it's like there's all sorts of variants, right? And because you, you did have to respond differently depending on what those things were.
- 26:21
What we're able to do here is really just take, you know, these free-form answers, interpret them, and then essentially provide a rationale for why you would say a given thing at a given time, right?
- 26:31
So this is equivalent to if you were doing this with a human and you asked the human, "Well, why did you say this?" The LLM can provide this kind of, of context.
- 26:38
So this is Claude looking at the history here, and I'll, I'll show you what this looks like in LangSmith, which will make it a lot more obvious. And then saying, "Okay, here, here's the next thing that I should say.
- 26:46
And my confidence that I should say it is 100%," right? It's, it's usually very confident, [laughs] right? But, but the point is, this, this whole process is largely gonna go along in an automated fashion, right?
- 26:56
You don't usually need humans involved because this is a very straightforward thing. They have their medicine. The next thing I need to know, and this is a very interesting part of this treatment, I need to know what time it is.
- 27:06
These are text messages. We don't know anything about these people. For a variety of reasons, it's kind of good that we don't know much about them, right? We don't wanna have to deal with all of the stuff around provider confidentiality and, and patient data, right?
- 27:15
So one of the things that we need if we're gonna go through this longitudinal treatment is to figure out what time it is for them, and then essentially pull out that data and figure out what their local time is, right?
- 27:24
So in this case, I was in Eastern Time when I answered these questions. This is all me doing this, you know, from my laptop. Um, I tell it what time it is.
- 27:31
It calculates an offset from UTC and says, "Well, I guess you're in Eastern Time," right? And then it sets this over here and it says, "All right, from now on, I know that my patient is in Eastern Time unless they tell me otherwise."
- 27:40
And they could come back and tell you otherwise, [laughs] right? That's something the old system really didn't have a good way to do. Um, but if the patient comes back and says, "I'm on a plane, it's actually 7:00 for me," we just update the time zone and move on, right?
- 27:50
This is a very flexible system that way. Um, then we get into this over here and we say, "Okay, now that I know what time it is, I'm gonna ask them if they've started their treatment," right?
- 28:00
And, you know, there is a blueprint, right? Which we'll get to, you know, that essentially just has, you know, the, the medicine that they're gonna take and a very specific way to take it, right?
- 28:07
The, the protocol. Um, the patient says, "Well, no, I wanna take it soon." You know, the Ava says, "Cool. I'll text you when we're ready." And then it gives, you know, a regimen, which in this case, this is an SVG that we are stapling times and dates on top of, right?
- 28:20
So, you know, fairly straightforward. We're doing this in software. The LLM's not doing it. The LLM is actually just passing along the instructions. You know, it says, "Send the step one image and provide, you know, like, this date and this time," and, and we substitute the rest of it in.
- 28:33
And this goes out as an MMS, right? So this is, this is a text message. Um, and so we provide this. The patient, you know, says... You know, in this case, we're, we're talking, you know...
- 28:42
Again, this is all the LLM reasoning through this, right? You know, I, I, I am sending this immediately because it's actually within sort of the 35-minute window that you've told me that I have to send these things.
- 28:51
This is all business logic that the LLM is, is interpreting pretty much on the fly. Um, and then I have these reminders, right? I didn't get back to it, so this is an important part.
- 29:00
It sent me this thing, and it thought that I was gonna take it at 5:45. I didn't text it back, right? This is partly because I was main- maintaining the system myself and I forgot, so I had to come back in the next day and catch up.
- 29:10
Um, so it sent me an automated reminder because it scheduled one when it sent the first message. So part of this is the LLM only gets called when the patient says anything.
- 29:18
So if they don't, you know, you have to make sure that you stay engaged, right? You don't do this overly. Like, we don't try to bother people beyond one or two reminders.
- 29:24
It's their treatment. Um, but this bump sort of functionality was really important to the client, right? So we built it in. Um, so you can see here, I came back the next day and I said, "Yep, sorry.
- 29:33
I, I did take it." You know, Ava confirms that I completed step one, and what it does is it sets this thing called an anchor, right? And it says, "Okay, you know, the patient was going to take it at 5:45.
- 29:43
They confirmed that they did, and so now, you know, I can refer back to this. I know that this happened," right? And if the patient had then said, "Oh, no, I screwed up.
- 29:50
I actually haven't taken it. I'll take it today," we just change the anchor. We update everything, right? So this is a system that humans used to have to do.
- 29:56
If a patient came back and said, "I didn't take my medicine," you know, a human has to go in and manually update all the times and all the scheduled messages, and it, it was, it was a big pain in the butt.
- 30:05
Um, yeah, please.
- 30:06
So to implement this functionality where a patient does not report back when they are supposed to report, your, uh, Python container is maintaining state per conversation?
- 30:17
Not exactly. Um, the way that we do state, and I'll, I'll spend a lot of time on this, but the way that we do state is really just that with any given message from the patient, right?
- 30:25
This entire system only kicks off when the patient sends a message. Um- What we do is we say, "All right, given this state, what is the best response?" And that response could be, "I change some of these anchors, I update their treatment phase, I s- I schedule a bunch of messages."
- 30:40
All that state is preserved so that the next time they write in, then, you know, we have that state to go on. Um, but again, we're not checking, right?
- 30:47
We- there's no polling going on in the system where we're saying, "After three hours, did the patient text me back?" We, we don't do that. We depend on the scheduled messages essentially just to nudge the patient.
- 30:57
Um, if they choose to not say anything for three days and they come back after three days, we just pick up where we left off. Um, again, this is a choice.
- 31:03
This is the way the client wants it. It's, it's, uh, uh, intended to be low enough touch that it doesn't bother people, but high enough touch that it doesn't lose track.
- 31:11
Thank you.
- 31:11
Sure. Um, I'll pause here, actually. Any other questions so far? Um, I, I realize I'm going through a lot. Yes.
- 31:17
Do your anchors have to be sequential, or can your answer come in at any point in the treatment plan?
- 31:22
They can. G- great question. So the question was, do the anchors have to be sequential? Um, or like, do you have to go through these one step at a time?
- 31:27
So one of the great things, one of the best things about this system, is that I could have, and, and I'm happy to try this w- when we go a bit later, I could have basically said, "Oh yeah, I already took the first pill, and I'm, like, in the middle of taking the second pill," you know, as,
- 31:38
like, the first thing I say to, to Ava. And she would be like, "Okay cool, there's an anchor, here's the next thing." It, it skips ahead, and it doesn't force you to go through this prescriptive part of the blueprint.
- 31:47
Whereas the old system, you know, at least nominally did, right? Like, you could, you could kind of skip ahead, but this automatically does it. You know, part of the instructions are, don't ask the patient a question they've already answered.
- 31:56
Like, period, right? Like, that's annoying. Don't do that [laughs]. Um, so, so yes, that's, that's very much in there.
- 32:02
Anything else? Yeah.
- 32:03
Does the LLM itself have a concept of the internal state machine that is kind of determining all of this? Or is that kind of outsourced to the actual software stack and then the LLM just kind of makes it look better?
- 32:15
Yeah. Yeah, so question is, does the LLM have an internal representation of the state? Um, kinda sorta. So you'll, you'll see, um, when we get into the, the, the actual back and forth with Claude, um, in LangSmith, you as a human can see kind of where it starts, right?
- 32:28
So every thread is gonna show, like, all right, here's the incoming state. We repeat it essentially to Claude. Again, that's just one c- dot I've never made, managed to connect with LangGraph, right?
- 32:36
So we basically have to serialize the state and say, "This is your, this is your starting point," but then the LLM has that in its window. And then, you know, it's gonna cause changes to the state.
- 32:45
It'll call functions that update the state. It can always ask again. It can say, "Well, what's the current state?" You know, it can go back and, and retrieve it.
- 32:51
Um, but in the context of that one, from when the patient responded, you know, to when I actually come up with my response to them, that whole thing is gonna be in its memory at one moment.
- 33:01
Did you ever run into issues with state being too big?
- 33:05
Um, so the, the short answer, uh, is... The question was, um, do we ever run into the state being too big? Uh, generally speaking, because of the way that we're kind of compressing and serializing at the end of the conversations, it doesn't ever get so big that it can't finish its job of responding to one situation, right?
- 33:21
You know, like, patient said this, now I'm gonna do this. We have considered having longer running threads where you kind of pick up in the middle, and you've already...
- 33:29
Y- you can reload sort of the entire previous conversation. That does get weird, right? Especially with older Claudes, you would get it forgetting to sort of call tools the right way and it'd have all sorts of JSON errors, right?
- 33:38
We have a bunch of retry logic in there to kind of compensate for that. Um, so that's one reason we kept it short. We make it so that we basically throw everything out and restart when the patient gets back to us, in part because blueprints could change, right?
- 33:50
Uh, you know, a bunch of things could ch- could change in the meantime that might end up with, with weird states.
- 33:55
So on the management of states and taking decisions what to do next-
- 33:59
Yeah
- 33:59
... and so on, this is 100% LLM driven or there's some softer, like, uh, logic around it as well?
- 34:07
It, it's 100% LLM driven. Uh, sorry, the question was, uh, is, is the, is the steering done by software, right? Any, any of that steering. The answer is really no, it's not, um, except for when it surfaces to a human, right?
- 34:17
And so when it goes to a human for approval, the human can use English and basically say, "Yeah, change that word to that," and, "That message shouldn't go out," and, you know, whatever.
- 34:26
So, like, we, we actually, as part of the flexibility part, we are not building any software that manages the state. We just want you to talk to the LLM to do it, right?
- 34:35
We think that's a better practice, right? It means, like, you know, you as a human just have to talk to it, and you don't have to figure out how to flip all the bits on this new console.
- 34:44
Do you have any RAG, uh, system? And, uh, second, if a, a patient going off a typical journey, how do you de- detect that and in- intercept?
- 34:52
I'm sorry, what was the first question?
- 34:53
Uh, do you have any, uh, RAG system in-
- 34:56
Or REX? I'm sorry, I don't understand.
- 34:58
RAG, like, uh, retrieval.
- 35:00
Oh, oh, got it. Sorry. So, um, question was, is there a RAG? Uh, no there's not, and, and it's actually just because what we really did is we just came up with a structure for the documents that was self-referential.
- 35:11
So you read a very small document which says, "Here's the treatment," right? "If you need to read for this phase, go to this file," right? "If you need to read for this phase, go to this file.
- 35:18
If you have a question that doesn't fall underneath any of those things, here's a CSV with a bunch of questions and answers." We didn't do it as RAG in part because we didn't believe that e- either we could do a really good job of getting all the right information into the window, like we didn't think we'd be
- 35:32
reliable enough about that. We just want to give the entire document. They're not that big. Um, and because th- these, this is Claude, right? It's, it's got a big enough window that we could just put the entire thing in there, you know, for, for most treatments.
- 35:42
Um, so we chose to do that. What was your second question, though?
- 35:44
Uh, if the patient going off a typical journey, um-
- 35:47
Yeah
- 35:48
... how do you detect and intercept?
- 35:50
Right. So the question is if the patient goes off track. So we, we have this idea of a blueprint, but then there are plenty of cases where the blueprint may, um, you know, not fully answer whatever the patient is, is, is bringing up.
- 36:01
Um, like one example is the blueprint is very much about asking questions, right? So you will say, "Have you taken your medicine yet? When do you plan to take your medicine?"
- 36:09
The patient will say, "My stomach hurts." Okay. So yes [laughs], your stomach hurts. You didn't answer the question. What we do is the patient typically will get an answer to their question.
- 36:19
So one of the principles is always answer the patient's question, right? We don't ever want to leave them hanging. But then ask yours again. So the idea is that at any given point, we can answer anything that they need, and as gently as we can, we'll try to pull them back onto the blueprint so that we understand
- 36:32
where they are in the treatment. Um, it's an, an exact science.
- 36:39
Someone that goes off or completely off the track
- 36:42
Well, the LLM does that effectively by knowing that it's supposed to keep people on the blueprint, but having an escape hatch for the knowledge base, essentially what we call it, right?
- 36:49
Triage or knowledge base, you know, whatever you want to call it. Um, so, you know, we, we don't have an explicit bit sort of flipped in the system that will say, "This patient is off track."
- 36:59
We just kind of know roughly where they are in the treatment, and if, if they want to ans- if they want to ask a bunch of questions, we'll, we'll just answer them until they, they are satisfied.
- 37:07
Okay, uh, anything else? Yeah.
- 37:09
So if you were building it again today-- Or actually, the question is twofold.
- 37:14
Yeah.
- 37:15
What drove you to actually use LangChain? And second fold, if you were building it today-
- 37:20
Yeah
- 37:20
... would you still use LangChain?
- 37:22
So question is, why did we choose LangChain, and would we still? Um, I, I will be very candid that the main reason that I chose LangChain is that I had personally gotten pretty comfortable with LangGraph as, as a, a demonstration of these concepts, right?
- 37:34
It, it's not that Crew-- I mean, we did a lot of AutoGen work back in the earlier days, right? You know, I've, I've done a little bit with CrewAI.
- 37:40
All of those frameworks can functionally do very similar things. LangGraph was the absolute best at explaining to people who were not neck-deep in this stuff how it worked. Um, and because there was a path to production from there, I didn't feel a need to, to re-platform and change all of it.
- 37:54
We certainly thought about it, right? We considered, well, what if we didn't do this in LangGraph? What would we gain? And the-- but the answer is you still have to implement observability in certain ways.
- 38:01
You know, you don't necessarily get, you know, the support that you might get from LangChain if you end up in a place. Remember that we're also doing this for clients.
- 38:07
We're not gonna be there forever. Um, leaving them with something that they can call, you know, somebody to, to support is also a helpful aspect. So I think-- I, I, I don't think I'd do it differently.
- 38:17
I think it's really just that, you know, ultimately, you know, we're getting pushed, a-all of us, in the direction of using the native model tools for this, right? You know, OpenAI has the Responses API, which lets you define tools.
- 38:29
Claude has its new stuff, right? Like, I don't really want to be locked in. Um, I, I, I am to some degree locked into LangChain now, but I, I prefer that honestly to being locked into the models.
- 38:39
Um, these are, these are not performance-intensive things we're doing in terms of the software, right? Like, you know, I don't care that LangChain is sometimes a little slow. Um, I would rather have the optionality.
- 38:50
So you said you're not using RAG in such documents. As the documents scale, how, how does your system be able to, like, fetch those in a deterministic way if it's not RAG?
- 39:00
Uh, it, it is just that they have-- Uh, sorry, the question was about, um, if it's not RAG, how do we fetch documents? The documents refer to each other.
- 39:06
So y-you'll see that we have an overview.md, right? This is all in Markdown. Um, there's an overview.md that tells you what other documents are involved in the treatment, right?
- 39:15
There's some of the prompting which says you can always request a triage overview, right, to, to try to handle problems. Um, and it'll be there, right, regardless of what the treatment is.
- 39:25
So it, it is very much just a document management thing. Um, RAG, the main issue is just that I, I don't think, and, and, you know, this will probably be more obvious as we get into it, right?
- 39:35
I don't think that you could really design a RAG which would pull back snippets of everything in sort of perfectly relevance wa- relevant ways. You really do kind of need to understand the shape of the whole treatment, right, to, to make a good decision, right?
- 39:46
Otherwise, you're just gonna parrot whatever particular snippet the RAG happened to bring back, and then the, the logical has to be in the RAG. It makes more sense and it's more transparent, I think, to do it this way.
- 39:54
Maybe you'll get to this in the state management later on. Are anchors predefined in blueprints, or are they determined by the LLM?
- 40:01
They, they are mostly predefined by the blueprint in that we say as part of the overview, you know, the concept of an anchor is that it is a thing that happened or a thing that will happen, and here are the examples for this treatment, right?
- 40:11
This is the thing that will happen or did happen in this treatment. Sorry, that was the question about the anchors. Yeah.
- 40:16
So when you, uh, you, you mentioned that you, like, compress the conversation and you keep track of all that. Is that in anchors, or is that another part of the product?
- 40:24
Uh, no. So the, the state is essentially, um, you know, we, we call it, for reasons that only an engineer could love, we call it a schedule document, right?
- 40:31
The idea is that for any given patient, there is a schedule that they're on, and the document snapshots their current state at any given point, right? And it's a versioned database, so we could go back in time and we could see what their document was three days ago.
- 40:43
Um, but it has at any given point the messages that have been exchanged, any unsent messages that are scheduled, and enough state about their treatment to fill out this view.
- 40:52
Thank you.
- 40:52
Yeah. Uh, yeah.
- 40:54
On the MCP front, do you have any, like, security layer, or is it just internal communication?
- 41:00
So in this case, all of this stuff is locked away, right? So I mean, just to, to go back to this diagram for a second, um, this entire thing is all behind, you know, AWS's, um, VPC, right?
- 41:10
So, like, there is no external access to the LLM, period. The only things it can talk to are essentially its own documents, you know, in, in local files and to the, the blue box.
- 41:19
So, you know, there, there certainly are vectors, but the vectors would be through the text messages, right? Not really through anything else.
- 41:26
Uh, yeah. Oh, I'm sorry. A bunch of people. You first.
- 41:29
Yeah. Um, just a question regarding, I guess, two-- it's twofold. There's one is, like, how are you assessing the confidence rate from the model's response? And the second-
- 41:36
Yeah
- 41:36
... is how are you safeguarding against prompt injection or malicious behavior?
- 41:40
Yeah. Well, so, uh, the question was about prompt injection and, and generally sort of steering. Um, I, I mean, the, the, the basic answer is just that w-you could definitely try to trick the model by sending weird texts, right?
- 41:52
And, and we do that as part of our, you know, sort of internal red teaming. Like, we have the entire team of operations associates who have been spending, you know, weeks and months trying to trick this thing.
- 42:00
Um, and granted, they're not trying to trick it from a reveal proprietary personal medical data. You know, I mean, there, there's things like that. We also obscure a lot of that medical data.
- 42:08
So the things that get in- get to the yellow box do not include phone numbers. They do not include anything other than the patient's identified first name. Um, so there's a lot of, there's a lot of that data that's kept only in the blue, which is a lot easier to, to defend against.
- 42:20
Um, so yeah, we, we, we very much do obscure the, the, the patient. We don't obscure the treatment, right? The treatment is fully visible to the LLM. Yeah. Cool.
- 42:30
Yeah.
- 42:30
Uh, you mentioned human-in-the-loop. Uh, how did you create evaluation of asking the correct answers based on the book?
- 42:38
Yeah. Uh, hold that thought. I will get to that very, very shortly. Um, we're back there.
- 42:49
Uh, sorry, is this a question about unclear instructions?
- 42:52
Yeah.
- 42:52
Um, so, uh, when, when the situation is ambiguous, the LLM is told to look at the blueprint and pick the best possible answer. Now, if you don't believe, or the, the LLM, you, if you don't believe the, the answer is perfect, um, you should say so, right?
- 43:08
In the rationale. So if I go back over here, this idea of the rationale, if there is uncertainty on the model's, you know, point of view, it can say, "Well, I picked this blueprint response, but I'm not sure that it's right."
- 43:17
In practice, it's not great at doing that, right? But that is the idea. And then the evaluator is also gonna look at this and say, "Well, did you actually pick either the exact blueprint response word for word?
- 43:27
Did you adapt it? You know, does this seem right to you?" Like, we're, we're trying to at least give a little bit of a layer before we get to humans.
- 43:33
And then hopefully we, we can trap situations like that and say, "Well, this is a complicated situation. A human should, should take a look." Um, it is not an exact science though.
- 43:40
Like, that's generally just true with this stuff. Sorry, you in the back.
- 43:43
Uh, I'm curious about the scale. Like, uh, how, like what's the load like how many messages can it handle in production? And if you had any issues with API calling on the model in that area.
- 43:54
Yeah. Um, so the question was just about load and scale. So, uh, the... look, the really short answer is that this, this system exists, right? There's an existing version of it that is humans pushing buttons.
- 44:02
Um, that scale is, you know, again, let's say thousands, not millions of patients. Um, this opens up the possibility of doing more treatments, right? That's how we would get sort of additional patient scale.
- 44:12
You can also sell this to new hospitals, new clinics, things like that. Um, so part of this is to get the scale to be larger. Um, we have not run into scale issues with, you know, just the, the conversations with Claude.
- 44:22
You know, the software that we're building would scale much, much larger than thousands of users, right? You know, the, the text message gateway might actually be the, the, the biggest bottleneck.
- 44:29
So it's, it's-- honestly, it's a problem we wanna have. Um, go ahead.
- 44:33
So, um, you keep saying Claude. Did you guys select the LLMs because it was what the client had access to? Or was there, like, a specific reason why you're going with Three Bot or whatever you're using?
- 44:43
Yeah. Uh, so the question was of model selection. Um, when we started this, right? And I think, you know, let- let's assume that we kicked this project off, you know, late last year, early this year, right?
- 44:52
Um, we had to make a choice, and our main criteria were, it had to be a steerable model that we felt pretty good about, you know, transparency-wise. Um, you know, one example, just, just to give you a, a specific one.
- 45:02
o4-mini is pretty good at this workflow, but it won't show its reasoning. Um, like, I mean, that's just one example. And, like, it's not a deal breaker. Like, we can still see the rationales, like, there's some pieces of it.
- 45:11
But I like being able to go into LangSmith and seeing the whole conversation, right? That, that really helps me out. Um, we needed, you know, again, flexible hosting, but I mean, all the clouds kinda do that.
- 45:20
Frankly, we didn't wanna deal with Microsoft, and we kind of preferred AWS to Google. That, that [chuckles] was kind of how we got there. But you know, uh, you can do this anywhere.
- 45:27
It really was just, we had to pick a horse, and we largely have not regretted it. A- and in part because we built enough flexibility where if I wanna switch, I, I still can.
- 45:38
So my question is about, um, sensitive data.
- 45:41
Yeah.
- 45:43
Are these messages being sent through a phone carrier? And sort of a little bit more about that. And then part B is, um, how are you using the data that you collect from people to enrich, to enrich the model, if at all?
- 45:58
Yeah. Uh, so the question was about sensitivity, uh, of data through the text carriers and also about, uh, using the data to learn. Um, I'll do the learning first.
- 46:06
Um, we don't. We, we do not take any of the responses and, and do anything to the models other than when we see situations that we as humans have evaluated and found wanting, um, we can tweak the prompting and the guidelines, right?
- 46:17
But we are not putting this in any sort of durable form. Like, ultimately, you know, we believe the right model here is the provider interaction. If there's a provider involved, that sticks around, right?
- 46:26
The provider knows that you interact with the system. They can have, you know, whatever records they need. Um, otherwise, you know, we forget about you when your treatment is done.
- 46:33
We think it's better that way. Um, on the, on the, the sensitivity question, yes, there is sensitivity involved. And at the same time, again, there's prior art with these products, right?
- 46:43
There are existing systems which essentially take, you know, text messages in and, and provide medical advice. Um, we're just trying to stay within the guidelines of that. And again, that's one reason why we don't want the LLM actually to have any data that is not explicitly required just to do decisioning, right?
- 46:57
It doesn't need anything beyond that to, to make a good decision. Okay.
- 47:02
Running in office-
- 47:03
Yeah
- 47:04
... uh, so every response is one hundred percent.
- 47:07
Mm-hmm.
- 47:07
How is it determining that that's one hundred percent?
- 47:09
Yeah.
- 47:09
And what about a situation where that's not eighty-five?
- 47:12
Yep. Um, sorry. Hold that thought too, because I will get to that in just a second. Um, let me move on. Uh, please, like, bring these questions back up.
- 47:18
I just wanna get a little bit further so we can see some other, some other cool things about this. Um, I'm gonna move on from this flow just because you can imagine that this is going over a period of days, right?
- 47:27
There's another step here, step two, where there's, you know, more medicine being dispersed. Um, and then, you know, ultimately we're gonna get to the end, right? And, you know, essentially, did you complete this?
- 47:35
And then, okay, great, you know, this is what's gonna happen to you. You know, you're gonna see some bleeding. Um, and then we have this check-in, right? So imagine that this now is, you know, a full, let's say, three or four days later, right?
- 47:45
After the, the treatment has begun. Um, you know, we check in. You know, the patient gets back to them or not, right? Remember, some of these patients will just be like, "I'm done.
- 47:52
I don't really need to talk to this thing anymore." But if they do, right, we continue with the treatment. We don't bother them. We just let them sort of resume where they left off.
- 47:59
Again, we have these rationales, you know, we have these questions. And then what I want to do here is just to show you briefly, um, sorry, I gotta zoom back out so I get the full phone number, um, what it would look like to interact.
- 48:09
So if I go here into my sandbox, um, imagine that normally this would be a text message. Um, so, you know, I would be doing this on my phone.
- 48:16
Um, but here, you know, I can answer this question. If I had any pregnancy symptoms before, have they decreased? It's like, "Yes, uh, they have decreased."
- 48:28
Okay, so I post this message. Now, what's gonna happen from here is thinking. So none of this is instant. And so now what I wanna show you is what this looks like in LangSmith.
- 48:37
So, um, you can see here a couple of things. Um, one is that this, this is now spinning. Um, so this thing that I just asked it is now in active processing.
- 48:44
I'll show you what it looks like when we're done. Um, but I will give you just a brief look at, um, I think this is probably a useful one here.
- 48:52
Um, what this actually looks like in terms of processing the state. Um, so I'll blow this up a little bit and make it a bit bigger. So, um, what you can imagine, this is using Sonnet 4, um, is that every time a message comes in from a patient, this is what I get, okay?
- 49:08
I get this description of, you know, everything that's going on here. I can see this is an Avila patient. I can see the thread that we're currently executing, right?
- 49:17
Because you may need to resume these threads if you need to give feedback. Um, I have this idea of I'm in the three-day check-in phase, so that's the blueprint that I'm gonna read.
- 49:25
Um, and then I have a couple of things. I have these anchors, right? Which, you know, you could see, I think this is exactly what you saw before, um, you know, f- in, in that same patient.
- 49:32
Um, these are all defined as, you know, actually a mix of UTC and, and, uh, Eastern timestamps. Um, that's one of the [chuckles] problems that's hard to eradicate. Um, getting LMS to deal well with time is really tough.
- 49:42
Um, but then I have this entire message queue, right? And this is the compressed state of the conversation to date, right? This does not include every message that Claude sent itself while it was thinking, right?
- 49:52
That part is contained in these individual LangSmith threads. I could go back and I could look at this if I needed it. Um, but what I'm doing is I'm compressing and basically saying, "All I really care about is the actual messages that went back and forth."
- 50:02
I want these rationales because I wanna be able to review them, right? That helps me understand the decisioning that's going on here. Um, you know, I want these confidence scores so I can go back and look, you know, what did it think at any given point.
- 50:12
And again, I'll show you one where the confidence was low. But these things can go on a little ways, right? This is probably, I don't know, twenty, twenty-five messages.
- 50:18
Right? All of this goes in as initial context in the window. Right? So if you had a hundred and fifty messages, all hundred and fifty of them are gonna potentially go in.
- 50:26
Now, we do have a, a, a function where you can optionally set it to compress and say, "Well, just show me the last fifty," right? If I need to request more, I can do that.
- 50:33
There's a way to do it. Um, but I don't need to have the entire thing in the window. Um, so I get down here. This is the last message from the patient, right?
- 50:39
So the question was, "Did you notice blood clots?" I said, "Yes, a few." Right? You know, that was, that was what I as a patient said. Claude is now gonna start processing this thing, right?
- 50:49
So imagine, you know, this all being basically pasted into, you know, a, a Claude window and then having it go through this process and, and call tools. So it starts by looking at directories that it's allowed to view.
- 50:59
Again, this is a version where it's got the blueprints kind of all local and, and it's, it's talking to them this way. We have another version where it talks via MCP over to the, the blue box, right?
- 51:07
The, the larger system. Um, so it figures out what directory it has. It reads these basic ones because these need to be read in all cases. So these guidelines, right?
- 51:15
The idea of how do you do your job, right? The idea of what the confidence framework looks like, the overview of the treatment, right? You know, those sorts of things.
- 51:21
We read those up front. None of these is very large, right? And so you read all this stuff, you know, it, it comes into the, the window. Um, and then, you know, essentially it reads those descriptions and it says, "Well, I was told as part of this that I have to read the current blueprint for this current
- 51:34
phase," right? So I read that file individually. So a bunch of these early calls are just about setting up the context. This is not the only way to do it, right?
- 51:41
I, I mean this, this is the way that we've chosen to do it. Again, we chose not to do RAG for a couple of, you know, reasons around we just did not think we could get good enough results, and because this is honestly easier to interpret, right?
- 51:50
You can sort of tell what it's doing. Um, I get to the blueprint. The blueprint, and you, we'll, we'll see more of these examples in a second. But the blueprint is basically this kind of structured, bulleted list, right?
- 52:00
Here's all the stuff that you might need to say to somebody, right? And, you know, here's what you do when, you know, the user says a certain thing. This isn't actually that prescriptive, it's just structured, right?
- 52:11
This isn't a- an if/then statement, right? It- it's kind of like that, but it's not an actual if/then statement. So, like, this format, you know, is one that we iterated on and got to a point where we actually get really good results.
- 52:22
Um, but, you know, it wasn't a hundred percent obvious this is the way to do it up front. Um, you know, we started with charts. Um, and so now you get to this point where now you can see, okay, now I gotta look at these, you know, uh, conversations.
- 52:33
I gotta figure out what's been going on here. And so you can see here, even though I passed in the state, it has a function to list messages. And so it basically says, "All right, well, now that I sort of know what's going on, let me see the last five messages," right?
- 52:43
And you can see here it's gonna start sending, you know, a bunch of these in. Um, and so it does that. It looks to see if there's anything scheduled.
- 52:49
There's not, right? And so now it says, all right, this is, this is sort of the point where Claude does its little explaining thing. I understand what's going on.
- 52:57
The patient's in the three-day check-in phase. I already asked about bleeding and cramping. I, I asked about blood clots, and the patient, you know, basically just said, yes, they have blood clots, and so I'm, I'm just gonna keep on going, right?
- 53:08
And it goes to the next question about pregnancy systems. This message comes directly from the blueprint, okay? And, and I'll show you, uh, in a Google Doc form in a second what that looks like.
- 53:17
Um, so it schedules it. It says, "You should send this message, you know, as, as soon as you want to." And then we get over to this evaluator flow, right?
- 53:24
And the evaluator says, "All right. I'm gonna look at this situation. I'm gonna look at everything that requires confidence scoring," right? That new message is the only thing. It's, it's the, the only thing that just happened.
- 53:33
Um, and I'm gonna send it immediately. This is just a, a timestamp for immediately. Um, I then get this kind of report, right? And the way that we set up our framework, um, and I'll show it in code a little bit clearer, is, you know, do we know what the user is saying?
- 53:48
Do we know what to say? And do we think that we did a good job? Again, this is a tough one [chuckles], right? Um, generally speaking, the LLM, you know, says at all times, "Yes, I know what I'm doing," and, you know, like, "Buzz off."
- 53:59
Um, but what I can also do is I can say, all right, then there's a bunch of cases in which if I set an anchor, if I updated the patient's data, like maybe I changed their time zone offset, maybe I changed their name, right?
- 54:09
That's a weird thing that, you know, if it happened, you'd probably want a human to look at. Um, do I-- am I sending multiple messages? Do I send a...
- 54:15
Am I sending duplicate messages accidentally? Do I have reminders for things that have already happened? All of those things would deduct from the score and cause a human to get involved, right?
- 54:25
That, that's part of how we do this, is to combine, does the model think it's okay, right? That's this top part. And then overall, is there a weird circumstance that I should try to catch, right?
- 54:33
And that I should try to, to, to show people, uh, to show a human for review. Um, in this case, nothing came up. I update the confidence. It's confidence of a hundred percent.
- 54:41
Um, and then essentially the virtual OA, you know, as, as a final thing, it's very hard to get Claude not to summarize itself. It, it does. Um, it basically just says, "Here's everything I did.
- 54:50
I'm good." And then if you go down here to the bottom, this is the output state. So this output state says, "Well, I have a hundred percent confidence" again, its version of it, "that I did the right thing.
- 55:00
I, you know, here's my anchors, here's my messages, and here's the unsent message that I'm n- I'm now gonna send." And because it's a hundred percent confidence, it just goes out.
- 55:08
Right? It goes back to the text message gateway, and it just goes out. Um, that is a risk, right? You know, y- if you want it to be perfectly safe, you have a human review all of these things.
- 55:17
We don't wanna do that because we're trying to scale, right? So we are comfortable in general with things that are, are, you know, coming back with a hundred percent confidence that we just send those messages out.
- 55:25
Uh, question back there.
- 55:27
Yeah. So when you're doing that evaluation stuff-
- 55:30
Yep
- 55:30
... are you, like, bucketing those situations somehow and, like, tracking what the agent is having trouble, like, determining in any way to, like-
- 55:38
Yep
- 55:38
... reassess, like, later on?
- 55:40
Uh, yeah, so the question is just, uh, how do we determine sort of the, the, the, the situations that might have confidence issues? Um, it is very hand-tuned and geared to this evaluation team, like basically the virtual OA team that exists now as, as humans.
- 55:53
Um, we will review, you know, in sort of spot checks, you know, a bunch of situations just to kinda see, like, "Hey, is, is-- does this seem like it's okay?"
- 56:00
Um, when they-- when a patient writes back, 'cause there are cases where a patient will write back and say, "You got that wrong. Like, that's not the time I said."
- 56:06
Like, you know, "I, I'm actually taking it now." Um, the confidence system is pretty good at picking up that that happened and basically saying, "All right, even if I think I'm confident, something's wrong," right?
- 56:15
You know, "A human should take a look at this." Um, but I mean, the, the answer is it's, it's more art than science. It's not something that we are perfect at even now.
- 56:22
And because we wanna scale, we've chosen to say, "Look, the, the worst that happens is essentially something weird happens and a couple of text messages go back and forth that are just wrong."
- 56:31
Usually, the human will get involved and say, "That doesn't sound right to me," right? [chuckles] It's not, it's not a case where the patient is in danger. Um, you know, if they say, "Well, I'm having these symptoms, you're not helping me," like, a human will step in.
- 56:41
Like, that's, that's something we're pretty good at flagging.
- 56:44
Yeah.
- 56:44
Yeah.
- 56:44
Like, it's like, are you, like, tracking that somehow? So it's like if, like, the patient says, like, "Hey, I'm having this abnormal bleeding. I'm taking this medication."
- 56:52
Yeah.
- 56:53
Do you know how many times that's happened? And do you, like, go back and, like-
- 56:56
Yeah
- 56:57
... address that?
- 56:58
Yeah. I mean, so the, the short answer is, um, we can look at interactions that ultimately are scored as low confidence, and then we can trace back from there, right?
- 57:06
So a lot of what we're doing is when something gets flagged and a human is like, "Whoa, there's something weird here," um, you know, we share those things internally, right?
- 57:13
The, the Slack channel that I was talking about before where they talk to the physician's assistant, that's largely been repurposed to people saying, "Hey, this behavior is off. Like, can you go take a look?"
- 57:21
And that ends up essentially in my queue as, you know, I gotta go check my evals. I gotta see if there's something I can do to catch this, and maybe it's a matter of changing the behavior.
- 57:28
Um, but so it's usually, it, it-- when we, when we know there's an issue, we can backtrack. That's the short answer. Uh, over here.
- 57:34
I'm curious. Uh, humans also make mistakes.
- 57:37
Yep.
- 57:38
Do you, uh, do you have any data from, like, before in the system of, like, the percent-
- 57:43
Yeah
- 57:43
... error rate for human responses versus AI?
- 57:46
Yep. Uh, so the question was about human versus AI error response from prior data. So yeah, great question. And, and yes, the answer is we do have that data, and that's one of the reasons that the client is as comfortable as they are with letting an LLM kind of run amok, right, is the idea that humans do
- 57:59
make mistakes now. And when they get escalated, you know, it's something where you can look back and you can be like, "Oh yeah, that was a little bit off."
- 58:04
You, you correct it and you move on. Um, this is kind of unique in that, again, it, it needs to be, you know, precisely worded. Like, one of the biggest risks is just that you give sort of off-label medical advice.
- 58:14
But if the idea is that, like, oh, you misunderstood and you have to go back and correct yourself, that's okay, right? It, it's, it's-- that's not a fatal error, right?
- 58:21
So a lot of it is that, you know, we think that we can get better use out of our humans by reviewing these situations, you know, than we can out of just having them push the buttons because they will occasionally push buttons wrong, right?
- 58:30
Same thing happens as, as with the robots.
- 58:32
So, um, you talked about mistakes, and, uh, this has been running for a while.
- 58:37
Yeah.
- 58:37
Have you thought about, um, fine-tuning a model with de-identified, uh, messages-
- 58:43
Yeah
- 58:43
... and then, like, running it back through?
- 58:45
Uh, yeah. I mean, the, the, uh, the-- So the question was about, um, have we thought about fine-tuning? Um, we have already seen two major model releases in the time we've been working on this.
- 58:56
Um, we [chuckles] we, we generally don't think that fine-tuning is a great use of our, of our dollars. Um, it, it, it-- obviously, it could be cheaper. We-- I mean, one, one example is, um, we tried, you know, at, at one point to use Haiku.
- 59:06
Um, and you know, Haiku is not even that much cheaper. It's maybe a third the cost, right? Um, we, we got to a point where we made our blueprints better, in part because, like, we'd sort of had some shortcuts where we just didn't have to be as, as precise with Sonnet, right?
- 59:19
You know, we had to be more precise with Haiku, and then it worked. Haiku did not get the time stuff. Haiku was terrible at figuring out what times it needed to sort of put on things.
- 59:27
And so the, the kind of thing we would have to do there, like, it either just kind of requires a smarter model, and there were smarter models from multiple people.
- 59:33
Like, o4 mini really is both, you know, at-- I mean, it's co- it costs a little bit less than Haiku, I think, right? And it, it was every bit as smart as Sonnet.
- 59:40
We chose not to go with it in part because it wasn't as transparent. Um, but so in, in general, we don't believe that fine-tuning is warranted because we think the models are just gonna keep getting better and cheaper and that we, you know, we'll be able to just kind of switch wholesale as opposed to having to fine-tune
- 59:53
something.
- 59:56
Um, so when you're going through that, uh, it sort of ha- you had this, like, chattiness with the model where it was describing its actions and then calling tools.
- 1:00:01
Yep.
- 1:00:01
Is that, like, is that an intentional choice? 'Cause I feel like you could just skip that and just do standard outputs.
- 1:00:06
It-- Well, so yes, it wa- it was kind of intentional choice, right? Th-this is partly that we, we already get the, the rationales and sort of the general, you know, explanation of its actions.
- 1:00:17
Um, but there are times where you wanna be like, "Look, why did it do this?" And you know, if it's thinking out loud, it's a lot easier to catch.
- 1:00:23
Um, so yes, uh, it, it's possible that we could eradicate some of that. We don't really think the juice is worth the squeeze.
- 1:00:29
So if you've had one miscarriage, most likely you're going to suffer a second miscarriage. How does the current structure set, set up so that you have a new anchor point to see, like, this person is using this medication all over again?
- 1:00:41
Yep. Uh, great question. So, uh, question's about essentially multiple treatments or coming back again after, you know, having gone through treatment. Um- There are a couple ways to do that.
- 1:00:49
So one is that, um, you know, again, depending on how you get there, if you scan a QR code, that can start kind of a new activation, so we can know that you're coming in a second time.
- 1:00:58
Um, but people will write back after, you know, two months and say, "I have a question," right? And, and so we either can just reactivate that conversation. Uh, the other thing is different treatments would usually come from different phone numbers.
- 1:01:09
So there's a few different ways to kind of disambiguate, you know, what somebody's actually up to. But that notion of, like, you know, "The same thing happened to me again, I'm starting the regimen over again," fundamentally, you could just explain it.
- 1:01:18
You just say like, "Hey, I had a miscarriage two months ago. I had another one. Can you help me?" And it would reset itself, right? The LLM is smart enough to do that.
- 1:01:28
Can I share some skepticism on the evaluator node?
- 1:01:30
Uh, please, because there's plenty to sh- there's plenty to share. [laughs] [laughs]
- 1:01:34
I mean, intentional- obviously everyone has the intention of improving, right? And two-
- 1:01:38
Yeah
- 1:01:38
... brains is better than one. So-
- 1:01:40
Yeah
- 1:01:40
... I understand the intention behind it. I guess I'm skeptical that doubling the costs-
- 1:01:45
Yeah
- 1:01:45
... are yielding 100% better outcomes, right?
- 1:01:49
Yep.
- 1:01:50
So to another question, do you have like a funnel of how often the evaluator-
- 1:01:54
Yeah
- 1:01:54
... might be impacting? Second question: You-- They're both Claude, right?
- 1:02:00
Uh, in this case they are, yes.
- 1:02:01
Is there-- Was there an, uh, an intentional decision to stick with Claude rather than switch the model family where in theory, hypothetically, you've got-
- 1:02:09
Yeah
- 1:02:10
... a different brain looking at the other brain kind of thing?
- 1:02:12
Yep.
- 1:02:12
And last question. Sorry, I know there's a bunch-
- 1:02:14
No, no, please
- 1:02:15
... um, was the inclusion of this evaluation node like something that made the client feel better? So were there other impacts besides just like this is a performance thing that made it worth it?
- 1:02:25
Yeah. So questions are all about sort of the evaluator node and the, and the processes. So, uh, the, the shortest possible answer is, um, yes, we're also skeptical about it, and at the same time, we think that there's still value in trying, you know, uh, essentially it's get- getting a second bite at the apple, right?
- 1:02:41
We do think that just having a different system prompt in the same conversation does occasionally deliver better results. But you could have the, the virtual OA evaluating the complexity of its own situation.
- 1:02:49
I don't think you could get it to evaluate whether it was right or not. Just typically, LLMs are terrible at that anyway. And so I, I think the, I think the basic answer though is that we wanted the flexibility in part so we could do things like try a different model entirely, right?
- 1:03:01
Or, you know, have something where maybe, maybe you did fine-tune a model specifically to catch these errors, right? Like that I think wouldn't be crazy at all. Um, so yeah, we wanted kind of that optionality, and at this point, you know, it's still early enough, right?
- 1:03:12
Again, it's running, it's out there. Like, you know, we're still tuning it. Um, if we get to a point where we're like, "Look, the only issue with this is how much it costs," or like specific details about like how good it is at catching errors, um, we'd, we'd go harder at that.
- 1:03:23
But we're pretty, we're pretty happy with the balance of it usually escalates situations that need review, right? It will sometimes screw up something just because it thinks that it was easy and it wasn't.
- 1:03:33
That, that does happen. The same thing happens with humans, right? So like we, we sort of are meeting the bar that we'd set for ourselves in the first place.
- 1:03:38
That's a good distinction though, that the evaluator has a different task of sorts.
- 1:03:42
It does.
- 1:03:43
So it's not really just the same thing two times.
- 1:03:45
Correct. The evaluator is looking at it differently, and it has this explicit-- Uh, so one thing actually, though, is that the evaluator can see what the VOA is supposed to do, right?
- 1:03:53
It can see the guidelines, so it ca- it is able to basically say, "You didn't do that right because I know what you were told to do and you didn't do it."
- 1:03:59
And likewise, the eval- the, the virtual OA can see the evaluator's confidence framework, and it can say, "Well, I'm gonna be scored against these things, you know, I better get it right."
- 1:04:07
Again, this is very much more art than science. But, but I mean, you're asking the right question about like, could we just have either a more optimal or a cheaper way of doing it?
- 1:04:13
I think the answer is yes. Okay, let me keep going for a second. Please just hold your thoughts. Um, so again, th- this idea of like every interaction looks like this.
- 1:04:21
It is a starting state, a conversation, an ending state, which then goes back to the system. And so what I wanted to show you here was if I go back to a conversation, right?
- 1:04:29
In fact, let me just see what I got here. Oh yeah, in fact, this, this answered. So I said the pregnancy symptoms have decreased. The next question in the blueprint is, do you think you're done, right?
- 1:04:38
You know, do you believe that, you know, the miscarriage and sort of the, the changes that these medicines were supposed to elicit have, have completed, right? Um, and there's basically one more message after this which kind of confirms and says like, "Hey, let us know if you have any questions."
- 1:04:49
But that kind of interaction, right? Back and forth, back and forth, a- assessing the state as it currently exists, is what this is built to do. And we're compressing after every one of these interactions into only the changes that happen to the state at any given time, right?
- 1:05:01
We're not saving... You know, we're-- In LangSmith, we're saving the entire conversation, right? This data-- Sorry, that's the wrong tab. This data, you know, about like what the virtual OA and the evaluator said to each other and what tools they called, this is preserved in LangSmith.
- 1:05:13
We don't get rid of this, right? But we do not save this in the state on the blue box, right? We-- That's not part of the patient's interactions with us, and we don't reload it every time you go back with, with a new message because that would ultimately both confuse things and, and blow up the context window.
- 1:05:28
So that's the way we've, uh, we've chosen to do it. Um, so let me s- let me now show you this. Um, I have another conversation here which actually needs response.
- 1:05:37
So I'm gonna grab this and put it in the sandbox. You can see what this looks like.
- 1:05:39
Sorry, Dan.
- 1:05:40
Yeah.
- 1:05:40
Apologies if that's off, but I just got confused a little bit.
- 1:05:42
Yep.
- 1:05:42
So how do you keep persistence then if you're, if you only have it online saved?
- 1:05:47
Uh, sorry. Persistence if you only have what?
- 1:05:49
If you're, you're saying here we get rid of that a- a- after every like new event.
- 1:05:53
You, you just don't save the, the process of the model talking to itself, right? You, you have it. You can refer to it if you need to. It's a debugging tool.
- 1:06:01
Input and output?
- 1:06:01
Yep, yep. Input and output is all that we snapshot in the, in the larger system. Okay, so now let's look at this. So I think that's actually the wrong one.
- 1:06:08
Let me go... Sorry, find this again. All right. Yep. So this, this right here. Actually, you know, I can, I can-- I don't have to go to the sandbox to look at this.
- 1:06:18
This is an example of what happens when things are complicated enough that we're asking for human review. Okay? So in this case, I've just started this conversation, all right?
- 1:06:26
And I said, "Yep, I got my medicine, came from the clinic. Here's my time." Now this is a moment where in the treatment, a lot of stuff is happening.
- 1:06:33
I'm figuring out what time zone they're in, right? And I'm saving that as part of the patient data, right? So in this case, I said I was on West Coast time.
- 1:06:38
So my time zone offset is four twenty minutes before UTC. Um, I am going to a new phase of the treatment. I have my medicine. You know, now I'm not in onboarding anymore.
- 1:06:48
I'm actually taking the medicine. And I'm sending multiple messages. So in the confidence framework, and I think I can find this, but, uh, I, I, I won't dig into it until we get there.
- 1:06:57
Um, in the confidence framework, we say when you have all of these changes at once, you should deduct from your confidence scores. You see up here this confidence of seventy percent.
- 1:07:04
I have the threshold set at seventy-five. So for anything that's below seventy-five percent, I stop and I ask a human to either approve, right? So if I were to approve this, it would just say, "All right, these changes are fine," and in this case, the changes are fine.
- 1:07:18
Um, or I could give feedback, right? I could say, and I'll try this now and, uh, live demos be damned. Um, let's say, you know, I wanna say, "Please mention the patient's name in your,
- 1:07:34
in your next me- in, in your messages.
- 1:07:38
Or in your message." So I'll say submit feedback, okay? And I'm working to figure out this 'cause it's kind of an operational detail. Um, this is now going and thinking again, so I'll have to reload this in a minute and, and see what happened.
- 1:07:47
But what's actually happening here, if I go over to LangSmith again, um, which I should be able to do...
- 1:07:56
is see that what's happening now is that it is restarting a thread that I already s- ha- started in progress. So the one exception to us wiping out its brain and reloading everything is when you come back with this feedback, right?
- 1:08:07
'Cause you want it to basically be able to s- pick up right in the thread and say, "Hey, you just did that wrong, but everything else here, like, you need to be able to see how you got to that place," right?
- 1:08:14
You know, so make the right decision and, and finish it up. Um, and so I think,
- 1:08:19
let's find out here. All right, still thinking. Um, oh, there we go. So you can see the only change that happened here is that it mentioned her name, right?
- 1:08:33
Otherwise, it's the same thing. Same time zone offset, same treatment phase, same reminder. Um, you can see the rationales here. Um, and, and you can see here the rationale even includes this.
- 1:08:43
I changed it to update the name. Now, you could imagine doing a version of this where I just had a little edit box and I said, "I'm gonna change this message."
- 1:08:49
We chose not to do that, right? We want the LLM actually to drive these changes. We think that it's better for humans to speak to them as though they're talking to a person.
- 1:08:57
Um, this is a debatable choice, but it is a choice that we made [chuckles]. Um, and part of that means we can be very, very flexible about the treatment, right?
- 1:09:03
We can just give feedback on the situation rather than having to build some sort of tools that are, are, are flexible enough to deal with all different types of treatments.
- 1:09:10
Um, but so here, I'm just gonna go ahead and say approve, and now those messages go out and the changes are made, right? I have, you know, my patient local time set and I know I'm in the n- next part of the blueprint.
- 1:09:20
Okay, I'm gonna pause here. I'm about to jump over to code. I think we have something like forty-five minutes left. Um, any questions on any of this so far that are not, "I just wanna see the code"?
- 1:09:27
Because [laughs] I can do that part. Um, over there.
- 1:09:30
Have you, have you heard any feedback from their actual customers on this experience?
- 1:09:35
Yep.
- 1:09:35
It seems like it's a pretty highly emotional interaction.
- 1:09:38
Yes. Uh, the question is about feedback from patients and that it is emotional. So yes, absolutely. So remember, this is a system the patie- or the client is already running, right?
- 1:09:45
So fundamentally, they already believe that they're talking to humans even when they're n- not exactly, right? Even the humans pushing the buttons are just calling up essentially bot-generated responses.
- 1:09:56
Um, when things get emotional, um, humans can step in. You know, we, we tend to steer them towards kind of approved knowledge-based responses. Like, you don't want this to be something where it goes completely free form.
- 1:10:06
There's, there's legal and other reasons not to do that. So by stepping in and having an LLM make the decisions, it doesn't really change the kind of current context of these treatments.
- 1:10:15
They're already getting, you know, basically this sort of medically approved feedback, you know, based on a certain flowchart, and if it goes somewhere, you know, a little crazy, uh, the, the escalation point is usually to call someone, right?
- 1:10:26
It's not, you know, we keep on talking forever in text because that's messy. Um, there are a bunch of points which I'm not gonna be able to demo here which basically just say, "Yeah, I'm sorry, I can't answer that question.
- 1:10:35
Call [REDACTED:phone_number], go to your doctor," whatever it is, [chuckles] right? But that, that is usually where it goes from there.
- 1:10:41
Uh, the blueprints look a lot like graphs or state machines themselves.
- 1:10:45
Yep.
- 1:10:45
Like, why not use LangGraph for that as well?
- 1:10:48
Well, i- it's a great question. So honestly, part of it is just that we needed to have something that the patient-- or not the patients, the client was actually comfortable maintaining, right?
- 1:10:56
'Cause remember, part of it is that we do not want this in code, right? We don't want this to be something where you can only maintain it if you have a technical person.
- 1:11:03
That's the problem they had before, right? And so just to jump over for a second, I'll show you what this, this kinda looks like. So this is essentially the thing that the, the client is maintaining, and I'll blow this up a little bit.
- 1:11:13
I realize that is small. Um, but the idea here is that we're using terms and, and, you know, we'll, we'll see a bit more of this in the code.
- 1:11:20
We're using terms that are defined in the framework. A trigger is, you know, something that happens, you know, essentially after an event, right? Um, you know, we have the, the conversation of these messages.
- 1:11:30
We always tell the LLM why this is important. If we just had this, this detail and we just said, "This is the message you send," I don't think it would perform as well.
- 1:11:37
It's much more helpful to actually give the LLM justification for why it would say something because then it makes better decisions. Um, one of the many quirks. Um, go ahead.
- 1:11:45
On, on that one, I, I'm not sure if this is code or not for the next part, but how, how complicated or how simple did you build the if statements kind of logic in there in the blueprints-
- 1:11:55
Yeah
- 1:11:55
... so they don't get lost? Or if you have any-
- 1:11:57
Yeah
- 1:11:57
... any, any situations that actually got lost in the, in the, in the-
- 1:12:00
Oh, sure. Yeah. So, so the question was just about, you know, essentially why, why do we have this framework and, and why is it maybe not more declarative, right?
- 1:12:06
In terms of, like, specifically if-then and, and that sort of thing, right?
- 1:12:09
A- actually, I, I use similar with, with, uh, with, a different program, with LlamaIndex, uh, and I, like, I got Llama just get confused on some of the if-thens.
- 1:12:18
Sometimes it would, like-
- 1:12:19
Yeah
- 1:12:19
... miss the ball.
- 1:12:19
Yeah.
- 1:12:20
So I couldn't go, like, very complicated.
- 1:12:22
Yeah.
- 1:12:22
Not a lot of nested if-thens. It's gotta be, like, one, two levels max, otherwise it got lost.
- 1:12:26
Right, right.
- 1:12:26
My question would be did you, did you face that using LangChain, yes or no, and how to kind of measure that?
- 1:12:32
So I, I, uh, the, the, so the answer is just about, again, how do you def- how do you define these things as clearly but, you know, maybe not complexly as, as possible, right?
- 1:12:40
So, um, this, this framework tends to work where you're really just saying: Look, I'm giving you this approved language and I'm trying to give you, in the bold statements here, right, primarily, I'm trying to give you a sense of, you know, what, what the conditioning really is.
- 1:12:54
But part of the reason that we did it this way is because, you know, if the patient writes back after this thing and he says, "You know, yes, I have the medication, and I took the pills, and my stomach hurts, and I'm confused," right?
- 1:13:04
I mean, it could be all these things. We wouldn't want to represent something like that in a flowchart. What we really want to do is just say, "Look, this is the outline of the thing.
- 1:13:11
You can see it." You know, if you need to jump ahead, jump ahead and don't ask the patient questions they've already answered. It just turns out that this, this framework really does work pretty well for letting the LLM do that sort of thing.
- 1:13:20
Uh, it's-- I know that's kind of a magic answer, but it's pretty good at it.
- 1:13:22
I guess Claude isn't smart enough to figure it out.
- 1:13:24
Uh, yeah, no, I mean, Claude, yeah, Claude mostly nails it. Um, most of them do. Yes.
- 1:13:29
Did the... Did... Does including the instructions, like, quantify- quantifiably improve the response quality?
- 1:13:36
Uh, q- uh, sorry, what do you mean by including the instructions?
- 1:13:39
Including the, the reasoning.
- 1:13:40
Oh, the reasoning. I, I... So question was, d-do-does including the reasoning help with the response quality? I think it does, right? I mean, this is one of these things where we started also by borrowing from human documentation, right?
- 1:13:50
So th-this was a process that was originally explained to humans who were gonna push the buttons. And so we took a combination of flowcharts that existed to explain the flow of the treatment and these kinds of, you know, this is the message that you should send in these situations, and, and this was kind of the, the hybrid
- 1:14:04
output of those two things. So I wouldn't say we did aggressive testing on is it, is it really better, or is it just that, you know, this is good enough?
- 1:14:11
It's more that, like, we started with this framework based on the human materials we had.
- 1:14:14
Kind of a follow-up question to that.
- 1:14:16
Yep.
- 1:14:17
Did, did you find any sacrifices that you had to make in kind of maintaining this document as human readable versus optimizing it for LLM consumption?
- 1:14:27
Yeah. Um, so the question is, uh, maintaining this document as human readable versus LLM. Yes, there are trade-offs. I think they're still worth it. We, we may change our mind at some point, right?
- 1:14:36
So, you know, imagine the workflow here being, um, you know, this, this Google Doc is maintained essentially by our physician's assistant, right? She is the co-owner of the blueprint maybe next to me.
- 1:14:45
Um, when we w- when we make changes, we talk about them together, we recommend in this document, and then accept them, and then effectively, I, I export it to markdown and check it in, right?
- 1:14:53
That is gonna change a little bit. We're gonna build a lot of these tools into the database, and so that, that's really where you'd be doing this instead. Um, but because this is human, you know, maintained, right, because it is basically, you know, still driven by the team, um, yes, we, we are making a trade-off.
- 1:15:07
I don't think it's a trade-off that's, that's super damaging.
- 1:15:12
Um, based on your current design-
- 1:15:14
Yeah
- 1:15:14
... just now when you do the approve, disapprove thing-
- 1:15:17
Mm-hmm
- 1:15:17
... um, what does change after you do an approval? Is it, like, a one-time thing, or does it improve your answer in the future, or even-
- 1:15:25
Yeah
- 1:15:26
... in changing the blueprint?
- 1:15:28
I-- So at the moment, no, and the question was just about, uh, the approve, uh, sort of defer, um, you know, feedback mechanism. Um, so actually I'll go back and just show this really quick.
- 1:15:37
This should be done now. Um, there we go. Um, again, we are, we are saving this in the sense that I can see this in LangSmith, right? I can look at this and I can say, well, in, you know, these cases where an approval was needed, and in this case, like, just, just as a, a visual, you
- 1:15:51
know, sort of feedback, whenever you have this graph null start, right, that, that is one of these cases where, you know, there was an approve feedback defer choice. Um, I could filter by this, and I could look at all of these things, and I could say, "Well, what kinds of things were we actually trying to approve or
- 1:16:04
give feedback on?" Um, we don't learn from them, right? We, we as humans will maybe update the blueprints. We do not put this back into training data again for a bunch of reasons which are, are kind of specific to the situation.
- 1:16:15
Um, but, you know, you can see here that, like, I can go all the way down here, and I'll try to find this quickly, um, and you'll get to a point where the human says...
- 1:16:21
All right. Yeah, here it is. So we get feedback from, yeah, from the human OA. This is essentially what happens whenever I push that button and I say, "Give feedback," right?
- 1:16:31
The human OA has feedback about your unsent messages. The feedback has mentioned their name. Um, it just goes right back to business. It's like, okay, let me look at the messages that, you know, I was sending.
- 1:16:39
I'm gonna update with this one. I'm gonna probably delete, I'm not sure. Actually, no, it just updated that one, uh, in place. We rescored it. One thing we've said is that we do not change the confidence score on something that a human reviewed.
- 1:16:51
We leave it where it was, right? We let them review it again, right? And so in all these cases, this, this is also a much quicker, you know, simpler sort of operation, right?
- 1:16:59
And so you can see the evaluator here is basically like, "Yep, that message is fine, but we're not gonna do anything, you know, uh, really to change the, the overall score."
- 1:17:06
Um, so, you know, that's the kind of thing that, you know, we can look at afterwards, right? But we are not, at this point at least, you know, really trying to feed that back into the model.
- 1:17:14
It's really just for the blueprints.
- 1:17:16
Just to get a sense of your metrics, um, a lot of these are, you know, over a minute, and it says about, you know, a couple hundred thousand tokens.
- 1:17:22
How do you-
- 1:17:23
Yeah
- 1:17:23
... kind of look? Is it, like, a necessary evil, the, the time it takes and the cost, or is this-
- 1:17:27
Yeah
- 1:17:27
... just a testing environment basically?
- 1:17:28
No, no. I mean, it, it is a necessary evil. And, and actually just to point it out, um, these costs I don't believe are correct. Um, one of, one of the shortcomings of LangSmith, and I think they've admitted this in various forms, is they don't really take into account the caching.
- 1:17:40
Um, so these costs should be lower than what you see here. Um, but, but fundamentally, yeah, these are expensive operations, and, you know, we could change, we could change some of them at the potential cost of higher error rates, right?
- 1:17:51
Like, we could try to cache more and have you inherit threads in progress, and it would be faster, right? Because you've already loaded everything. It would be, you know, potentially you're not reloading any context, and so, you know, you're, you're spending maybe less on tokens, and you just might have a higher error rate.
- 1:18:04
And, and that's, you know, a, a thing we are trading off.
- 1:18:08
Is there some kind of knowledge base that your model is taking to, like, a pivot with depending on the medicine that this, your-
- 1:18:17
Yep
- 1:18:18
... that patient is taking? Because they are the-- maybe the side effects are different-
- 1:18:21
Yep
- 1:18:22
... and the questions will be different.
- 1:18:23
Uh, yeah. So the question is just, uh, uh, in terms of the, the knowledge bases. So let me, let me actually jump over and just show this really quick.
- 1:18:29
So I mentioned these blueprints. Um, I'll, I'll jump over now really into just what the, the, the implementation looks like. So you can see over here, you know, this idea of for Avila, right?
- 1:18:39
We have a handful of documents here that are, again, exported into markdown. Um, and I'll try to blow this up because I know these are small. Um, let me just shrink this down.
- 1:18:48
Okay. So the idea here is that, you know, I've got all of this, you know, uh, sort of framework data, right? The idea of defining. What do I mean by a blueprint, right?
- 1:18:56
We're doing-- We're, we're defining this every time not in, um, the, the prompt, right? We're doing this as part of the context window, in part because we do want this to be really flexible.
- 1:19:05
If you want to change the terms, um, you should be able to do that, right? We don't want the treatments to be hamstrung by, by terms we use for other treatments.
- 1:19:11
I define anchors. I talk about schedules. I talk about scheduled messages, right? So all of this stuff exists in part just to, to lay the groundwork. And then this framework, right, is now referring to specific documents, right?
- 1:19:21
And so you can see here, like, I've-- Again, these documents are all referring to each other. So I can go through here and I can look, you know, and, and click on these links and go straight to other things.
- 1:19:29
If I want to do something around the knowledge base, so the way that we do that is this triage idea. Um, so, you know, if something happens that a blueprint doesn't address, right?
- 1:19:38
So the way that we tier it is first check on the blueprint. If you have approved language, use it, right? Send it, send it back for human review, whatever you need to do, right?
- 1:19:44
But use that approved language. If you don't think you can answer that question, you go and look at this, which now again, is, is self-referential. We don't read the entire scope of medically approved knowledge all at once.
- 1:19:55
We let the LM decide, are they complaining about stomach pain or bleeding, right? You know, if I can't find anything in any one of these, I have a larger knowledge base, right?
- 1:20:02
Which is, you know, sort of just a laundry list of like random questions people ask. Um, we've chosen to do it this way in part because it is human readable.
- 1:20:09
It mirrors something the client already mostly had, right? They already had a lot of these structures. Um, and you know, we fundamentally did not believe that it made sense to over-process, you know, things like a RAG.
- 1:20:20
Now, I will say that for the thing that we have like kind of a backup, you know, sort of knowledge store, which is almost entirely a CSV, that probably is suit for- suited for RAG, right?
- 1:20:27
It would be okay to use a RAG for that. It doesn't get used that much, right? So in some sense, it's just not worth implementing that way, at least not yet.
- 1:20:36
Okay. All right. Um, yeah.
- 1:20:39
Uh, yeah, thanks. Uh, the virtual way, uh, many, many in common there was also evaluations or going back and forth.
- 1:20:45
Yep.
- 1:20:46
Is that also just above this? Is that like defined in the overview right now or is that at the instructions level?
- 1:20:52
That is prompt level. So we-- the, the virtual OA and the, and the evaluator both have relatively small prompts, which I can show. So let me just see if I can find them here.
- 1:21:00
Um, those prompts are, are... They do reference each other, right? So imagine this being, um, you know, again, built on the LangChain stuff, so the base agent class. Um, this prompt is basically aware of the other agent, right?
- 1:21:13
So in this case, it's just two. So the prompts do speak about each other. The evaluator knows about the virtual OA and vice versa, right? You know, the things that we try to do, and, you know, th-this is, I think, normal prompt engineering stuff for people who've really played with this stuff.
- 1:21:24
You have to tell it how to take turns. You have to tell it that it, you know, if it gets called on, it has to talk, right? Like, you know, one, one problem we have that we have to sort of frequently do retries on is the LM thinks that everything's done, it doesn't say anything, and, and the
- 1:21:35
whole thing dies. Um, so, you know, you have to talk, but then you can be done, right? You just have to say something. Um, we have the basic idea of you have to determine this overall confidence score, but we don't include this in the prompt because we wanna be able to show it to the virtual OA as
- 1:21:50
well, right? So the details of how you score something is, is, is factored out. Um, but the notion of here's who you are, here's who this other guy is, and here's how you work together, that is in the prompts.
- 1:22:00
Okay. Actually, on that note, let me jump over and actually show you some of the, the confidence stuff and the, the guidelines. So, um, I'll start with confidence, and I'll get into the guidelines, which are much, much longer.
- 1:22:10
This again is, you know, intended to be mostly LLM readable. This is not something the client generally maintains, right? So this is not in the same category as these blueprints.
- 1:22:18
Um, but the idea here is that, you know, I've got this confidence score, and, you know, I am trying to figure out ac-across these multiple dimensions, do I know what's going on?
- 1:22:26
You know, here are some examples. We, we are trying to be as prescriptive as possible with examples of these different situations. Um, do I know, you know, the knowledge that I need to know?
- 1:22:34
Here's an example which might speak to your question actually over here. You know, the idea of, do we want the LLM to use its world knowledge to figure out that when I'm talking about an antibiotic and I give a specific antibiotic, that it applies to the whole class of them?
- 1:22:47
Yes. That, that's a risk that we're kind of willing to take, right? We don't need to have an explicit, "This specific antibiotic is safe for this treatment," right? That would very quickly spiral out of control.
- 1:22:56
So we do have a handful of places where we ask it, use your own judgment, but refer to, you know, the, the knowledge base and the blueprints for, for your baseline.
- 1:23:03
Um, and then so after I get through these categories, then I have this idea of deductions, right? And the deductions here, um, are specifically things like, you know, you should deduct from the overall score, not the individual messages, right?
- 1:23:14
Because an individual message could be like, "Yeah, this is exactly from the blueprint," like it's the right thing to say. But overall, these situations can be complicated, right? And so what we're trying to do is explain that such that it can, again, score the overall interaction in a way that surfaces it for, for human review.
- 1:23:29
Um, okay. So I'm gonna move on to the guidelines because, again, there's just a lot more in here. [chuckles] Um, this is long enough that I'm not gonna review everything, but I'll, I'll try to get to some of the biggest parts.
- 1:23:37
Again, some of this is really simple, right? Like tool calling. Um, one issue we've certainly had over time is, you know, fabrication and, and, you know, honestly, instructions like this do help.
- 1:23:46
Um, you know, do not make up a tool call. Wait, wait your turn, right? Call the tool and step back. Um, there's a lot of stuff around time, right?
- 1:23:53
There is a lot of stuff around, hey, you need to ask about it in the right way. You don't ask about it in a way that forces someone to tell you where they are, right?
- 1:24:01
Their people are very sensitive about this. They don't want, you know, people knowing where they physically are, but you need to know what their time is so that you can schedule the messages for them, right?
- 1:24:08
Um, you want, you know, when you work with time, um, you have to use things like, you know, I-ISO timestamps. Like, that's how the rest of the system works.
- 1:24:16
Um, but calculating these things and keeping them all straight, it requires a relatively smart model. So a lot of this, you know, has, has sort of grown over time to just work with the idea that, you know, this is how you can talk to models about this and, and do a pretty good job.
- 1:24:28
Um, setting anchors, scheduling messages. Again, these are all the things that are core parts of the system. This is not treatment specific, right? This is all written to be generic enough that I don't have to rewrite this every time I add a new drug, right?
- 1:24:39
Which, which is one of the core requirements, right? We did not wanna have to do this in code. Yes.
- 1:24:43
Hey, uh, so this is a very big document.
- 1:24:45
Yep.
- 1:24:45
And every time we're setting it to a prompt, right?
- 1:24:48
Yep.
- 1:24:48
Uh, have you experimented with the prompt caching? Because this doesn't change.
- 1:24:52
Yeah. Correct. No, we have. And so right now, and I'll, I'll get to caching actually in just a second. Every-- I'm just trying to manage time here, but I do have time for that.
- 1:24:59
Like, we are, we are doing some explicit system prompt caching, and then we are caching explicitly, um, you know, the, the multiple turns of the messages such that, you know, each operation is, you know, I think the average operation with just sort of the baseline stuff is maybe ten to fifteen thousand tokens, right, per turn, all cached.
- 1:25:16
It adds up, right? So, you know, you do have maybe the average cost to generate a single message somewhere in the fifteen to twenty cent range, right? It, it's not cheap.
- 1:25:24
But we are caching as aggressively as we can. We have thought about things. All this is brand new, right? The idea of the, of the hour cache, you know, that, that Claude just introduced.
- 1:25:31
It's not clear to us that that would help because, you know, we can't guarantee that the patient's gonna get back to us within, you know, either five minutes or an hour, right?
- 1:25:38
It- it's just... it's a bit of a, a risk to take at the system level. Um, but if anybody here actually knows more about Claude caching than I do, [chuckles] ple- please talk to me because, like, we i- we-- ideally, we would like to cache a lot of these documents.
- 1:25:48
Um, it's just not clear if we can do that across sessions. It's not clear if, you know, we would get the benefits that we're looking for. So we've just tried to be as aggressive as we can within a single conversation.
- 1:25:59
That's a very long list of guidelines. How did you come up with it, and how do you optimize it, and what do you do when you optimize it?
- 1:26:05
Yeah, I mean, the... Look, the, the, the real answer is that I mentioned before that, you know, there's a few thousand lines of code and a few thousand lines of prompt.
- 1:26:12
This is most of that prompt, right? I mean, this is a lot of it. Um, it is something where, you know, we have tuned it over time. This is, you know, myself, as well as, you know, the, the, the physician's assistant.
- 1:26:21
Like, we have come up with something that we believe is fundamentally, you know, pretty good at handling these, you know, generic situations. And when we find edge cases, we, we just modify these prompts.
- 1:26:30
It-- Again, it is not perfect. We could definitely think about subdividing this. We could think about moving some of it into the prompts. Um, but we think this division is, you know, roughly correct for keeping it generic so that it's, you know, it handles a bunch of different treatments and it, it handles the situations that we see
- 1:26:44
across treatments pretty well, right? You will have cases where you're doing medicine that's all in one day. That- that's relatively unusual 'cause, you know, that's something you just send instructions home.
- 1:26:52
Um, you know, but there... I mean, I- I'll show an example around Ozempic. You know, that's weekly, monthly, right? Like, there's, there's much longer durations. We've tried to get to a point of balancing it where, you know, we, we do end up with a good result.
- 1:27:04
Okay. Uh, yeah, back there.
- 1:27:06
Yeah. So this is very long prompt. Have you ever seen, like, uh, how this model predicts? Or sometimes there's one issue, or basically, like, they have a very long prompt.
- 1:27:16
Sometimes they're adding more and the-
- 1:27:18
Yeah
- 1:27:18
... you know, the, the language model will try to stop it, right? So basically-
- 1:27:22
Yeah
- 1:27:22
... now they are trying to put more context.
- 1:27:25
Yeah. So the question is just about the prompt length. Um, so again, not, not to dismiss that out of hand. In Claude terms, this isn't actually that long, right?
- 1:27:33
I mean, this is, like I said, I think on average fifteen thousand tokens. Um, it still leaves a lot of the window, you know, behind, right? It is, it is not actually so long that we start seeing really crazy behaviors until we start doing multiple turns, like multiple conversations in one thread, right?
- 1:27:48
That's where it starts to blow up. Um, so we, we just genuinely have not gotten to a point where we're like, "My guidelines are too long." You know, the guidelines could be a little shorter, and I think we, we have optimized them in various places over time.
- 1:27:58
But, you know, we've, we've crammed this into a box where we really can sort of process one situation all in one gulp without really feeling any, any pain.
- 1:28:07
Okay. Yeah.
- 1:28:08
Any, any examples on tools that, that you have used? You mentioned tools that are not really-
- 1:28:12
Yeah, yeah. No, so I, I can, I can share a little bit of that. So let me, uh, let me go down here to the tools code itself. So, so one thing to note, this is a hybrid, and I think I mentioned this really earlier on about, um, there is some stuff coming from an MCP gateway, right?
- 1:28:25
So in this case, you know, I'm loading from files. It's just the file system MCP, again, all localized to our VPC. So there's nothing crazy going on there. Um, but I could instead load from, you know, my database, right?
- 1:28:35
I could choose to, to have that be the place where you interact. There's a bunch of other tools, right? And so, you know, I have this. This actually is where [chuckles] probably most of the code in my app actually is.
- 1:28:44
Um, the list of tools is essentially down here, and you can see it's things that enable interacting with the state, right? So all of these functions that are looking at anchors and messages and confidence and the treatment and patient data, all of that stuff is local to my graph run.
- 1:29:01
I don't have an MCP for it. I could. I, I just chose not to. Um, and you know, in this case, like, this code just lives in this Python app.
- 1:29:07
You could ob-- You could very much refactor this out. Like, the state can still live here, and the code could be somewhere else. It's, it's, you know, it's up to you.
- 1:29:13
Um, but one note actually about all of this stuff is... I'm trying to find a good example here. So I am aggressively using the command object. Um, for anybody who, who has programmed with LangGraph, the whole idea behind this is that at any given point, you're able to pass back a message, and this particular thing is just
- 1:29:27
an error message. But, like, you know, you're able to pass back a message and a place to go, right? So you can say, "Here's the re- here's the response, and by the way, I know that the evaluator asked for this, so go back to the evaluator."
- 1:29:37
You can actually get around some of the graph routing, um, this way. And so we've, we've, we've definitely, you know, tried to work this into both the MCP tools and the state tools that we have.
- 1:29:45
Yeah.
- 1:29:45
You talked about improving the prompt, right? What are your goals of doing that? Do you generate synthetic data? Or was it, you know, applied on people and using it with the past code and then-
- 1:29:55
Yeah. Uh, so question is about how to, how do we improve the prompt. So, um, it was really a fusion of, um... We were able to go from essentially the physician's assistant, who had the most experience with tricking things, right, with coming up with tricky situations, right?
- 1:30:09
So we were able to test a lot of the edges really just with her, you know, having her pretend to be the patient. We then scaled up to the full team of operations associates who, you know, then tried to trick it at a higher level, right?
- 1:30:19
And, you know, were putting out things that they'd seen from patients themselves, you know, trying to sort of, um, you know, get to these complicated cases. Um, and then, you know, with real people, you know, we're able to take that a step further.
- 1:30:29
Um, but with those first two levels, we're, we're not seeing, you know, a tremendous amount of stuff that's not expected. Again, this system exists. If we'd been doing this from absolute scratch, I think we would have a lot tougher of a time coming up with what we think the edges are.
- 1:30:41
Whereas, you know, there's a system that already exists. There's a lot of conversations to draw from. We're able to run some of those back. So we'll look at s- at conversations in the old system and replay them here and essentially just try to figure out, you know, where the edges are.
- 1:30:54
Is there some evaluation criteria you use what makes it a state versus not?
- 1:31:00
Time is one of them.
- 1:31:00
Sure. So i- is there a question about... Are the questions about, um, how to decide what goes in the state or not? Um, I mean, the... I guess the short answer is everything that comes in in that initial payload, which I'll go back over here for, um, all of this stuff.
- 1:31:14
I think this one's probably the better example.
- 1:31:17
Yeah. So all of this stuff over here, th- this is all state. Um, there's not really a distinction. Like, everything that comes in to sort of preload the conversation is state.
- 1:31:26
Some of it is editable. Um, some of it's not. I'm trying to remember examples. Like, and here examples are you can't change the source. I couldn't say, you know, in the context of the LLM operation, this is not an Avila patient anymore, right?
- 1:31:37
That's-- The LLM's not allowed to do that. Um, it also can't change past messages. So the LLM is not allowed to look at the message queue and say, "Message five that went out three days ago no longer exists."
- 1:31:46
Like, that function doesn't exist. Um, so we've sort of just calibrated it to where the only things it can do is read the entire state, modify the patient data, modify messages, modify anchors, like, uh, like unsent messages.
- 1:31:59
So, uh, it's, you know, it's just a software choice. Like, that's how we architected it.
- 1:32:05
So I guess, I guess if you have a session, you, uh, you either wanna kinda store the messages that went back and forth in, within the session. Like, nothing inside, but-
- 1:32:14
Yeah.
- 1:32:15
If the context window gets too big and you kinda wanna summarize it-
- 1:32:19
Yeah
- 1:32:19
... or go ahead-
- 1:32:20
Yeah.
- 1:32:21
In that case, would you wanna be able to deleting some messages or summarizing?
- 1:32:26
Uh, sure. I mean, I, I think the, I think the real answer though is that hasn't happened for us. Like, the way that we've structured it, there is not a single thinking operation that, that goes long enough to blow things up.
- 1:32:36
Um, but, but I will talk really quickly about, um, retries. So we do have the notion, and I'll just pop this up, maybe just an easier way to see it.
- 1:32:43
Um, so we do have the notion inside the graph of, you know, again, the virtual OA talks to the evaluator, you know, at, at a certain, a certain point in the thing.
- 1:32:51
Everybody uses tools. But then there are cases that will cause the, the graph to retry a cer- a certain operation, right? So one of them is, um, one of the agents malforms a tool call, right?
- 1:33:01
It tries to call a tool, but it uses the wrong JSON and, you know, otherwise things would've died. We can detect that. We delete the message, and we say, "You screwed up that tool call.
- 1:33:09
Try again." Right? That, you know, s- keeps it inside the graph, right? Essentially, this retry node is then able to loop back via that, um, command message. Uh, it's not shown here [chuckles] on the graph.
- 1:33:18
But via that command object, it can say, "Go back to the virtual OA. Try that again." Right? We also have some cases where, again, you know, uh, we expect the model to talk, and it doesn't, right?
- 1:33:28
There are just a bunch of cases where Claude will just end its turn prematurely. We detect that. We say, "You have to say something." [chuckles] Right? Like, literally, that's the message is like, "Don't just say nothing.
- 1:33:36
If e- even if you're gonna end your turn, just say, 'I'm done.'" Right? That's enough, you know, for us to keep the logic going. So it's things like that.
- 1:33:43
Low confidence from the evaluator will re-trigger retries?
- 1:33:46
I'm sorry, say it again.
- 1:33:46
The eva- uh, low confidence coming from the evaluator will, will retry-
- 1:33:50
No, no. So low confidence is something we wanna pass through to the human, right? So low confidence is a valid response to the graph, right? You know, like, I have a low confidence message I want a human to review.
- 1:33:59
That's fine. Uh, yeah.
- 1:34:02
You built your system prompt up over time. It's long, it's pretty complicated, right?
- 1:34:06
It's not the system prompt.
- 1:34:07
Sorry.
- 1:34:08
No, no, it's that thing. It's, that's why-- It's one reason we separated it. So yes, the guidelines have expanded over time.
- 1:34:11
The guidelines.
- 1:34:11
Yep.
- 1:34:12
Okay. Um, do you have any automated regressions against that if you wanted to change it or add to it? Because you've obviously tested it thoroughly by human.
- 1:34:21
Yeah.
- 1:34:21
How do you make sure you don't-
- 1:34:22
Yeah. Um, perfect time to talk about evals. So we're getting towards the end. Um, let me talk about evals a little bit just to sort of give you guys a, a sense of what we did here.
- 1:34:29
So, um, this was, this was a weird one because if you go to LangSmith, and I'll, I'll just... I'm not sure I can find the exact place where this happens.
- 1:34:36
Um, but there's-- Is it under-- Maybe it's under datasets. I, I forget. There is a place in here where you can basically say, "I wanna run, you know, an evaluation against, you know, uh, these, these, you know, this, this dataset that I've defined in LangSmith."
- 1:34:49
Um, I do define datasets in LangSmith. So I can see here, assuming this loads up, which hopefully it will. So I've got a Happy Path dataset here that I defined in LangSmith.
- 1:34:57
I, this isn't the whole thing. Again, some of this is redacted. But, um, if I look at these things, what I'm really doing is I'm saying, "Okay, this is an interaction that we had in the past."
- 1:35:05
Um, you know, this was one that I ran last week. Um, you know, here's, you know, the, the input, right? This is the last message from the patient. This conversation, you know, A, I've got this initial state here that I can look at, right?
- 1:35:15
So I can see, you know, what was going on in the first place. Um, I see this conversation and, you know, I get to the end, and I get a message out which says, "Great, you're ready to start."
- 1:35:24
Okay? So this is one part of my Happy Path dataset. The evals that I'm running, um, are fundamentally, it's a t- it's a custom harness. I'll blow this up a little bit.
- 1:35:33
Hopefully, it's visible to most folks. Um, but again, the idea here was that we couldn't just say, you know, when I ask for, you know, what the weather is in San Francisco, it gives me back, you know, cold and, you know, foggy.
- 1:35:47
Um, it had to be, here's examples of these input states and then, you know, let's evaluate it, you know, sort of an LLM-as-a-judge form, what the output state looks like, with the caveat that one of the big things we wanted to test was things like time operations.
- 1:36:00
So I can't put in an eval from three weeks ago and run it now and get equivalent times. I have to either give it really specific guidance on how to handle time or I have to replace all the timestamps.
- 1:36:09
Um, we ended up sort of doing a, a hybrid of both. Um, and so what this did was my eval suite is a custom Python app stapled to a Bash script.
- 1:36:18
This was all client coded. Um, I, I, I got what I wanted, but I can't vouch for much more than that. Um, I load datasets, right, from LangSmith. I call essentially the medical agent via, in this case, like I'm running this locally.
- 1:36:30
Like I could run this against my cloud LangGraph instance. But I'm literally running this against LangGraph on my laptop using that dataset from LangSmith, having preprocessed a bunch of stuff around dates, times, you know, sort of circumstances, so that when I get back the result, I can not confuse the LLM as a judge about whether it's right
- 1:36:46
or not. Um, and I'm using this LLM rubric, right, to do this. And so let me see if I can find my rubric here. So well, here's, here's a couple examples.
- 1:36:55
So I've got this, this YAML. So this is just, this is Promptfu and how it works. Um, what I do is I basically say, "Look, I'm trying to test, you know, this, this custom thing that I'm gonna call, you know, essentially with, with my, um, you know, my, my custom harness, and then I want you to evaluate
- 1:37:09
it, you know, with an LLM." And in this case, I think it's using GPT-4o. Um, you know, this, the evaluation rubric is basically, is everything basically exactly the same?
- 1:37:18
Do I see minor discrepancies, but I don't think they're a big deal? I mean, we're, we're keeping this very fuzzy. What we really wanna know is, is anything completely busted, right?
- 1:37:25
And then there are cases where it is. Um, I had to put in specific notes here like, "Hey, don't be picky about like the, you know, different wording that you might see in something like an anchor," right?
- 1:37:33
Sometimes it'll say they will take the medicine. Sometimes it says they did take the medicine. It's like, who cares? Like in this case, the spirit of it was right.
- 1:37:39
Um, and so, you know, and the times and dates, like there's some specific language here. So all of this turns into- Basically this guy right here. So, um, when I look at this, and again, I realize there's a lot of text here, um, it's not worth looking at all of it.
- 1:37:54
I ran these three examples from my dataset, and I got, you know, basically a, a passing grade, right? I'll, I'll go into the fail, one fail, one pass in a second.
- 1:38:01
But in each of these cases, right, I can see what the LLM-as-a-judge said, right? So if I go down here, this is GPT-4o going, you know, opining on, "Well, this is the source data you gave me and the sample output.
- 1:38:13
Here's the run I just did. Close enough," right? Um, you know, it does point out some things, right? So if I ran these evals and I was like, "Well, it's actually a big problem that, you know, the reminder, you know, unsent message didn't show up," right?
- 1:38:24
I, I could make a choice to do that. I could st- I could strengthen my rubric. But the way that we did this was just to say, "Look, at any given point, we do need to be able to test the current state of the system.
- 1:38:33
We want to do it against, you know, first the happy path, and then we can certainly do it against edge cases." Um, we're actively maintaining this. I think this might change once we actually hand this over, you know, more fully to the client, right?
- 1:38:42
We, we want them to have all the protection they might need. But it's this style, right? This style of eval. Um, does that answer your question more or less?
- 1:38:48
Yeah.
- 1:38:48
Okay. Okay, um, yes. Go ahead.
- 1:38:52
Or did you check to see if it, uh, caught specific cases? You know, like, "I know this is bad in certain ways. Make sure the judge catches that."
- 1:39:00
Yeah. So I mean, the, the short answer is yes. Some of that I'm redacting for, for a couple reasons. But yeah, more or less, we, we started with a happy path.
- 1:39:06
We do have a handful of, of specific like, "Hey, this is busted, and it's frequently busted. Let's make sure it's not," um, you know? The... But it, it's just a different dataset.
- 1:39:14
Yeah. Uh, yeah.
- 1:39:16
Just curious, this is a fantastic graph, but I, I think you mentioned earlier about blueprints and all that. So how much Python code or Node.js code have you written for this particular-
- 1:39:26
Well, so again, the, the part that I'm showing here is entirely Python in a LangGraph container. Um, I would guess it's, you know, maybe four thousand lines of code.
- 1:39:35
Um, most of it is honestly the tool calls [laughs] like that, that just... And I, I'm sure I could refactor that to be shorter also. Um, it's not much. It's really just enough to run essentially this graph, right?
- 1:39:45
You know, I have to have the routing between all of it. I have to have the tools that it can call. Um, everything else is in the prompts and the guidelines, right?
- 1:39:51
And so, you know, it really is more English than it is code. On the other side, right, on the other side of the box, right, this thing, um, that blue box is, I don't know exactly how much code, but it's entirely, um, Node and React and, and Mongo.
- 1:40:05
Um, and, and frankly, I haven't been that involved in it. Sorry. Yep.
- 1:40:09
Um, I also work in a highly regulated industry, so this has a lot of overlap for me.
- 1:40:14
Yeah.
- 1:40:14
Um, how do you think about taking this, like this is one medical procedure-
- 1:40:20
Yeah
- 1:40:21
... among hundreds, thousands.
- 1:40:22
Yep.
- 1:40:22
How do you think about taking that and scaling it across an industry or another client who comes to you wanting something similar?
- 1:40:28
Yeah. Great question. So, uh, the question's about scale. So let me actually jump over and, and show you one thing I should have probably already shown, but I, I forgot.
- 1:40:35
I mentioned before the idea of us doing different treatments. So, um, what I did in, in, in this particular case, so, you know, again, we were focused on, on the, the, you know, early pregnancy loss.
- 1:40:45
Um, I did this. In fact, I think I have this sitting here somewhere. So let me zoom out and find it, and then I'll zoom back in. So I took this to Cline, and I basically said to Cline, "Hey, I've got..."
- 1:40:56
Yeah, this should be it right here. Um, I said, "I'm looking to make a new treatment," right? "I defined one for Avila. Here's the structure," right? "Have a look.
- 1:41:04
Um, here's the link to, you know, Novo Nordisk's suggestions on how to dose Ozempic. Um, make me something new," [laughs] right? Um, and it basically went through this process, and I'll, I'll shrink this down, um, and created, you know, a basic treatment, you know, for Ozempic.
- 1:41:19
It read these files. It then decided, "All right. I got it. Here's the thing I'm gonna do." Um, I just said, "Cool. Go for it, and here's a couple of tweaks based on the thing that you said."
- 1:41:27
So I, again, this is my client process. I use this all the time. Um, and I ended up with what I think is a pretty serviceable Ozempic treatment, which I will very quickly show you here.
- 1:41:36
Uh, if I go back to these conversations, I'm pretty sure I named these people all O. So there you go. Here's Oliver with Ozempic. Um, and you can see it's still Ava.
- 1:41:45
I didn't change that, right? I, I, I... You could obviously get to the point of having it be a different personality. But in this case, this is Ava with a new treatment, asking all about Ozempic pens and helping me figure out their time, same as we were before.
- 1:41:56
It's a weekly injection. I didn't change any code for this. I literally threw this through Cline, got a new set of treatments out, and it just kinda works. Yes.
- 1:42:05
Yeah. On your graph, what's with the retry node?
- 1:42:09
What is that for?
- 1:42:10
It, it's for catching, uh, the retry node. The question was, um, it's for catching those errors around, um, malforming tool calls is one easy way for the graph to terminate, right?
- 1:42:18
So if it, you know, forgets a bracket and, you know, sends back something that's invalid JSON, um, Claude does this a lot less now. Like, Sonnet 4 is pretty good at it, but, um, Sonnet 3.5 was not nearly as good.
- 1:42:28
So we can detect that, and we can just say, "Well, you're trying to make a tool call because I see certain things in here." I either see tool call ID as a parameter.
- 1:42:35
I see weird brackets. You know, we, we can parse that in code. Um, and then, again, we wipe that message out, and we go back to the guy that called it and said, "Hey, you fucked up."
- 1:42:42
Pardon my French. Um, you know, "Try again." And so that idea of, you know, the retry node, there's a handful of those situations where we aren't gonna take over and do it in some deterministic way.
- 1:42:52
We're just telling the LLM, "You made a mistake. Here's the character of your mistake. Try again." And we do have to wipe out its memory of that mistake because it can get very confused.
- 1:43:01
Like, one of the reason-- one of the ways this happened a lot earlier in our testing was it would malform tool calls and then hallucinate the results, right? And so if you left the message in there, it would think that it understood the blueprint even though it had made the entire blueprint up, right?
- 1:43:12
So I don't wanna sugarcoat where, like, there are some weird cases that if you don't very carefully control for it, you can end up with some very bad behaviors.
- 1:43:19
But we were able to catch the main ones and essentially just give it another shot.
- 1:43:22
I was looking for, like, a loop.
- 1:43:24
Oh. Oh, right. Sorry. The reason there's no loop, this, this is just totally, uh, this is actually one thing where I think LangSmith and LangGraph could be better at this.
- 1:43:30
The retry node is capable of calling back to the other ones using the command object, but it doesn't show up on the graph. So it turns out, like, they were very invested, LangSmith was or LangGraph, in graph flows, and then they introduced this idea of, well, you don't really have to define it in the graph.
- 1:43:43
You can send anything anywhere. And so that's what we're using. Yeah.
- 1:43:53
Yeah.
- 1:43:55
Um, so there is this 0.25 gap. How do you bridge that gap to delivering real-world human recommendations? So there should be some, like, human intervention.
- 1:44:04
No, sure. But remember... Sorry, the question was about confidence scoring and how do we get it to, to sort of to be higher. Um, we want the score to be lower when there is a complicated situation, not because we think it's wrong, but because we want a human to review it.
- 1:44:16
I mean for the one that, that passed at 0.75.
- 1:44:21
Uh, well, so, w-what... Sorry, what do you mean?
- 1:44:23
So human can review the trigger if it's below 0.75.
- 1:44:27
Correct. Yes.
- 1:44:28
So if the one above 0.75 is triggered lower than this.
- 1:44:32
So if, if the confidence score is above the threshold, meaning higher than it, right? So let's say I have a confidence score of 0.9. What that could mean, right, is that I am sending multiple messages at a time, but there's no other reason for me to think that those messages are wrong.
- 1:44:45
Um, we have chosen, with the client, to set the threshold lower than that because they don't want their humans having to get involved every single time something is a little complicated.
- 1:44:53
But we've agreed that if it's below 0.75, they should. It, it's purely just a calibration, right? It's you, you decide. If you wanted your humans involved in everything, set the confidence score to 100, right?
- 1:45:02
You could see every single thing that happens. You're just clicking approve all day. Um, you're George Jetson. But, um, that's hopefully not what people actually want, right? You want a bunch of these things if they are high confidence and, you know, fundamentally recoverable, let's say.
- 1:45:15
Like, you know, there might be certain circumstances where you don't want a message automatically sent out. Hopefully, we can control for that, but in general, like, we wanted to set it in a place that felt like we were gonna get some scale out of our humans, right?
- 1:45:25
Which meant messages going out automatically. Right. Uh, yes.
- 1:45:32
So it seems like you're kind of merging confidence and case complexity into one metric. Um, do you find that that kind of has any side effects with it instead of-
- 1:45:45
Yeah
- 1:45:45
... kind of trying to evaluate those two things separately and then making a determination as to whether to get a human involved-
- 1:45:52
Yeah
- 1:45:52
... um, using those two signals?
- 1:45:55
Yeah. So the question is about conflation of, of confidence scoring and complexity. Uh, yes, 100% agree. And again, it's one of the things where I think we're happy with how it sort of operates now, but I'm not sure that we're doing the confidence piece right.
- 1:46:08
Um, a-and at the same time, like, the concept of it, I think, is, is good, right? You know, do you have the information you need? You know, is, is there anything...
- 1:46:15
Like, do you understand the user's intent or, you know, did they say something ambiguous? And we do see cases where it triggers, right? It's not that we never see it, you know, correctly rate itself as being like, "Well, I'm not totally sure what they said."
- 1:46:25
But it's not as frequent as we would like. So we combine it with complexity in part because, you know, we just wanna get it below that threshold so that we can have a human review it.
- 1:46:32
It, it's not a perfect system, but it is a system that we think works.
- 1:46:35
So it's more of a confidence that this falls within the happy path or defined blueprint scenarios-
- 1:46:43
Correct
- 1:46:43
... than it is confidence that we know what the response is.
- 1:46:46
Uh, correct. So y- uh, to, to that point, it's, it's not confidence that the specific response is exact and perfect and whatever it is. It's confidence that we don't think that there's a blend of uncertainty on, you know, what the response is and complexity of the situation, right?
- 1:46:59
Either of those things can push it below the threshold.
- 1:47:03
How are you hosting the, the graphs in the US?
- 1:47:08
Uh, so this is all-- Uh, the question was about hosting. So this is all, um... LangGraph has a pre-built containerization that you can use. Um, we are using it with a couple of modifications.
- 1:47:18
Um, we are essentially deploying both halves of this, right? So to go back to this. We're deploying both halves of this with Terraform. Uh, you know, everything's hooked up with GitHub actions.
- 1:47:26
I mean, like, you know, we're, we're doing as much of it as automatically as we can, but we are using most of the built-in LangGraph containerization.
- 1:47:33
I thought there was a... Like, you have to pay for the platform. You have to have an API key to get a certain amount of runs.
- 1:47:39
Yep.
- 1:47:39
So are you guys not doing that?
- 1:47:40
Uh, you do. No, I mean, we're-- we, we, we have a conversation with, with LangChain actively about that. So the question actually is also, um, the, the specific nature of how we need to be deployed.
- 1:47:49
Right.
- 1:47:49
Um, LangGraph platform doesn't currently support that, right? So, like, there... It's, it's all, it's all evolving. But, um, we, we have had those conversations. Um, let me pause for a second.
- 1:47:57
We still have ten minutes, um, and so... Or roughly ten minutes. Nine minutes. Um, let me just double-check my own list of things that I wanted to talk about just to make sure I didn't miss anything major.
- 1:48:06
Um, talked about that, talked about that. Dun, dun, dun, dun, dun. All right. I, I think we more or less hit everything. Um, and, and so I'm happy just to open it up.
- 1:48:15
Um, I have a couple of final thoughts that I'll leave sitting up here, um, if anyone's curious. Um, any other questions?
- 1:48:22
Uh, just a bit of a practical question.
- 1:48:24
Yep.
- 1:48:25
Um, it does feel like maybe you have spent a lot of your resources maybe for each project. Uh, how do you... Are you, like, allocating two to three months, like you and your other people who work in a project for you to adapt full time?
- 1:48:38
Like-
- 1:48:39
Yeah. Uh, so the question was just about, uh, resource allocation to build these things. So, I mean, I think the best way to say it is just that we had built a handful of things like this.
- 1:48:49
We had not done it for healthcare, right? But we were able to bring some things in in terms of, you know, I had an open source LangGraph project that I was comfortable using as a base for, for this client code, right?
- 1:48:58
Again, we, we forked it, we brought it in private. Um, you know, but it got us, you know, part of the way there in, in part because this isn't a totally special snowflake.
- 1:49:05
It's, it's a workflow with tools. So, you know, we have, I think, spent less time on that upfront scaffolding each time we've done it. Um, and you know, we're looking at another project right now, which would be that much quicker.
- 1:49:16
So a, a lot of it is just once you do this and you understand the mechanism, getting to, you know, good enough or getting to a starting point is just a lot quicker.
- 1:49:23
Um, it is something though where like, you know, a-again, we, we are, we are still working on this and we've been at it for a few months, but I think we probably built what would've been a year's worth of conventional software, maybe more, right, you know, in that short period of time.
- 1:49:38
Anything else? Yeah.
- 1:49:40
So you mentioned you did a lot of vibe coding on this one.
- 1:49:42
Yeah.
- 1:49:42
Can you explain that a little bit? I'm, I'm curious of how much of specific Lang- LangGraph-
- 1:49:48
Yeah. Yeah
- 1:49:48
... the tools, like, know in terms of-
- 1:49:50
Well, okay, so, so the, uh, the question was just about, uh, vibe coding and how much the tools know of these frameworks. So the single biggest issue with vibe coding, right, is when you're dealing with something new enough that the models don't really understand it.
- 1:50:01
Um, you can always just send them the API docs, and they usually do pretty well. Um, I ran into plenty of places with LangGraph where nobody had tried to do this before and no one had this exact bug, and we just had to actively debug it back and forth.
- 1:50:13
Um, I can definitely endorse, uh, o3 as a better debugger than Claude. Um, so, you know, we, we did have cases where we had to do that. Um, but for the most part, like, I mean, as long as you can at the docs, you know, you, you can get to a reasonable place.
- 1:50:26
And again, like, I am a former engineer. I, I may eventually call myself a current engineer, but I, I kinda wouldn't right now. Like, I don't have all the same practices and sort of all the same hygiene, you know, that our, our professional engineers do.
- 1:50:35
But I do know how to smell kind of bad behavior, and I'm pretty good at prompting to get what I want. So that, that's kind of how it, how it worked out.
- 1:50:46
Um, any... Oh yeah, please.
- 1:50:47
In this example, uh, you, you did not use MCP, right? The-
- 1:50:51
No, th- there is MCP. MCP is, is in a couple places, right? So if you look at the connection, there, there's actually two, and one of them I just didn't totally label.
- 1:50:59
The blueprint knowledge base at the top, that yellow box, that is an MCP connection in the LangGraph context. And then vertically, there's another set of connections back to the database.
- 1:51:08
So there's a couple places where it happens. But a- again, it's mostly for the inter- the exchange of state, and it's for this, you know, reading of documents. That's primarily where it happens.
- 1:51:18
I had a question about, like, um, the types of people that like to send a whole sentence but in like five different messages-
- 1:51:27
Yep. [laughs]
- 1:51:28
... like they'll respond to one after the other.
- 1:51:28
Yep, absolutely.
- 1:51:30
How do you handle that? Are you queuing it up-
- 1:51:31
Yeah
- 1:51:31
... or some type of time-based?
- 1:51:33
Yep. So the question is about rapid-fire text messages. So yes, the, the way we handle that is, is twofold. So one is that, um, we do have a configurable...
- 1:51:41
'Cause I'm not totally sure we have this turned on, but we have the idea of a configurable delay before we actually send it for processing. So if someone is gonna send five text messages in a row, we wait five seconds before doing anything, right?
- 1:51:51
Um, that's when we can catch it. But the other way is, um, we will invalidate previous running threads if someone texts in afterwards, right? We don't want an in-process response.
- 1:52:01
We want to take whatever the complete context of the conversation was, send all of that in, and then we respond to five messages at once, right? So it's a combination of smart retries and smart invalidation.
- 1:52:14
So in your case, um, if a user decides to act rogue, rogue basically-
- 1:52:19
Yeah
- 1:52:20
... and giving some irrelevant [laughs]-
- 1:52:23
Yep
- 1:52:23
... absolutely wrong responses-
- 1:52:26
Yeah
- 1:52:26
... because you are not using a RAG, the system is not getting influenced by it, right?
- 1:52:31
W- well, uh, so the question was about rogue responses.
- 1:52:34
How are you, uh, how are you detecting that it's a rogue?
- 1:52:37
Well, in general, we are, we are giving some pretty basic guidance around you're only here to answer treatment-related questions. If someone wants to talk to you about the weather, you just say, "I'm sorry, I can't help with that," um, you know, or other worse things.
- 1:52:49
Um, so generally speaking, that works pretty well. Um, we have it pretty well tuned to escalate, right? If someone just is essentially going off the rails and you can't think of a response, you just say, "I'm sorry, I'm gonna get somebody to help you."
- 1:52:59
And it sets the confidence low, and then a human can get involved. Um, it doesn't in practice happen that much. I mean, again, if you're involved in this, like, you know, if you're involved in this and you take the time to actually engage with the system, you want the result, and you probably wanna get back to your
- 1:53:11
life. So we don't see a tremendous amount of it, but, but that's the way that we would deal with it.
- 1:53:17
At any point in your development process, did you get frustrated enough with LangGraph where you thought like, "This is certainly more than just a framework"?
- 1:53:24
Uh, yeah. Did, did I get frustrated enough with LangGraph, um, at any point? Honestly, so the, the one thing I'll say, and I mean, this, this is just totally, uh, a personal project.
- 1:53:32
Um, I was doing a side thing, um, around college counseling, um, and I just have this sitting here 'cause I show it off sometimes. Um, and the point was, I, I went into this project e- explicitly trying to avoid it.
- 1:53:44
I said, "I'm gonna do s- a very similar thing where I have an AI in the middle of a workflow, and I want it to ask questions and I want it to think, and I do not wanna use LangGraph or Crew or any of these other frameworks because I don't wanna be dependent on them."
- 1:53:54
Um, and so I asked, you know, uh, I asked Cline to write me a layer that was pretty good at, you know, uh, talking to these models and structuring a thinking process, and it, it did fine.
- 1:54:03
I mean, the reason to have the framework is in part because, you know, again, we won't be with this client forever. We want them to have something they can operate.
- 1:54:10
We want them to have something explainable and easy to use. Like, this is all, you know, logs in Google Cloud. Like, it's not the most fun experience. So, you know, it is very much choose your tools wisely, you know, good and bad.
- 1:54:20
Like, you get a better experience with the other ones.
- 1:54:22
Speaking of that-
- 1:54:23
Yeah
- 1:54:24
... why Cline over Cursor?
- 1:54:26
Uh, uh, thank you so much for asking me that question. Um, I, I... Uh, so Cline, Cline over Cursor. Um, I, I will admit that Cursor and, and Windsurf is great, by the way.
- 1:54:35
Like, I mean, there's a handful of them that I really do like. Um, Cursor and Windsurf a- as, as two examples are, are, are just trying to hit this sort of narrow, um, you know, thing about it has to cost $20 a month, and therefore it has to be heavily optimized about how it sends tokens in different
- 1:54:49
places, otherwise their economics blow up. Um, Cline doesn't do that. It's very simple. It's real-- I mean, it, it's very smart, but, like, it's just giving tools to a smart model, and that smart model can cost whatever it costs.
- 1:54:59
Cline is the spiritual cousin to Claude Code and Codex, right? I mean, it's that style of thing as opposed to an IDE that has to hit this very narrow target and has to do a lot of pre-optimization on how the, the tokens flow.
- 1:55:09
Thank you.
- 1:55:09
Yeah.
- 1:55:10
How do you balance what you put into tests that you run before you deploy versus, like, relying on eval and production data to guide the following model improvements?
- 1:55:19
Yeah. So, I, I mean, honestly, uh, th- this is also another place where I will raise my hand and say, "This is why I'm not a real software engineer."
- 1:55:26
I, I don't have a robust testing framework on my Python code. I, I haven't really needed it, right? Or at the very least, like, I'm using evals and sort of overall performance of the system as the better benchmark, right?
- 1:55:36
So evals are important. We have to have those. Um, uh, you know, the other side of the code, like, you know, Stride is a TDD shop. Like, our, our, our software engineers are very, very good at test-driven development, and so the blue box is very well tested.
- 1:55:48
Um, but just my code, you know, very much is, is eval sort of tested instead. Um, it's not a great answer, but, like, that, that is kinda how I thought about it.
- 1:55:58
All right, um, I think we are at time. Thank you guys so much. This was awesome. [audience applauding] Um, I'll be here if you wanna stick around. [upbeat music]