AI Engineer World's Fair 2025
The Build-Operate Divide: Bridging Product Vision and AI Operational Reality
About this talk
Freeplay product leader Jeremy Silva and Chime speech-analytics leader Chris Hernandez explain why promising generative AI prototypes encounter operational and reliability problems in production. They advocate iterative evaluations, human-in-the-loop review, hallucination monitoring, and feedback loops, while positioning customer-experience and quality-assurance teams as prompt testers, model reviewers, and ongoing AI performance monitors.
Chapters
- 0:00The build-operate divide and speaker introductions
- 1:12From traditional machine learning to the GenAI quality chasm
- 3:34Hallucinations, human oversight, and review feedback loops
- 6:06Turning QA and customer-experience teams into AI operators
- 11:05Continuous monitoring and operational responsibility after launch
Talk transcript
- 0:00
[upbeat music] Um, we titled this talk the build operate divide, 'cause we've ob-observed this troubling trend where, like, good AI product concepts fail to reach their full potential due to operational
- 0:25
challenges. So today, Chris and I are going to be talking about some of the learnings we've gleaned from, like, our experiences on the front lines operationalizing AI products at scale.
- 0:34
By the end of this talk, we hope you're going to leave with an understanding of how you can kind of bridge that gap between product concept, um, and operational reality by understanding how to deliver quality through evals, human review, and how you might build your teams around that.
- 0:48
So a little bit about ourselves. My name is Jeremy. I lead product at a company called Freeplay, um, which exists to help solve a lot of these operational problems and, and help companies ship great AI products.
- 0:59
Hi, everyone. My name is Chris Hernandez. I lead our speech analytics team at Chime. Uh, I've been in the industry for about ten years, uh, from a CX perspective and, uh, about nine years on the ML space.
- 1:09
So excited to be here.
- 1:12
So both Chris and I kind of came from the traditional ML world into the GenAI world, and I thought it was, like, helpful to maybe take a note and talk about, like, what we feel like has changed in this transition.
- 1:24
Um, so I think the biggest difference in my mind is this decreased barrier to entry, right? In the traditional ML world, like, you need tons of data to even get started.
- 1:32
There are these long model training cycle times, and that barrier to entry just kind of like goes down significantly, um, in a GenAI world, right? Like, you can now use the intelligence of the base models to be able to, like, make use of and leverage, uh, smaller data assets within your organization.
- 1:49
And then what comes along with that is this, like, increased iteration speed, which is, like, as the barrier to entry goes down, the speed at which you can iterate starts to go up.
- 1:59
And this increased iteration speed starts to accentuate the need for, like, a high-quality ops function. Um, so we want to take a look at why.
- 2:09
Working across, like, dozens of different enterprise teams at Freeplay, we've noticed this common trend emerge, which is companies will build an initial prototype, maybe they'll even ship a V1 of that thing into production, but they inevitably hit this sort of quality chasm as they try and go from a V1 to a V2 that really drives true value
- 2:27
for the customer. And there's this reliability problem in there. And what we've seen is, like, the only way that you really cross that quality chasm that, and get to a reliable V2 is through iteration.
- 2:41
And so we talk about this iteration loop as you move from monitoring to experimentation to testing, evaluation, and I, I use that word broadly speaking, human review, auto-evaluation, like all these things all kind of come into play.
- 2:55
Um, but what happens is your product quality becomes a direct function of your ability to move through this loop, right? The faster you can iterate, the more times you can move through this loop, the better your product quality becomes.
- 3:09
So it, ops kind of sits at the foundation of this, and especially as you start to scale, right? Like, your ability to move through this loops is, like, ends up being an ops function.
- 3:19
And the, like, not so subtle irony here is that to deliver high-quality AI products, you actually need a ton of human elbow grease. And I'll pass it to Chris here to talk a little bit more about the importance of human experts in this process.
- 3:34
Cool. Awesome. Um, you know, I made the joke that one day I was like, you know, the first time I saw LLMs, we started playing around with it. You know, I typed in like, uh, "I wonder if it actually knows like basic things like, uh, who invented Wi-Fi?"
- 3:45
And then it spit out Abraham Lincoln. I was like, "Oh my goodness- [laughs] ... we're in some trouble." But, uh, all jokes aside, and maybe not that extreme, it is a good reminder that LLMs, and even the smartest ones, can make mistakes, uh, and they often do it with a lot of confidence.
- 3:59
Um, so, uh, that's of course what we call hallucination. And, uh, in the GenAI space, um, it, it, at times, like you got to make sure that, you know, we don't leave it unchecked and that these hallucinations can actually be tricky and risky.
- 4:11
Uh, they could also mislead customers, and they could misinform decisions. And when you think of like companies or industries like healthcare, uh, just a small hallucination could have a lot of risk in it.
- 4:21
Um, so that's where human-in-the-loop comes in. That's why it's important to talk about this today. And human-in-the-loop kind of ensures while GenAI does the heav-heavy lifting, uh, humans are still there to steer the ship and, and drive, drive in the right direction.
- 4:34
Um, so today we're going to talk a little bit more about like human-in-the-loop, why it's important. I don't think this is like groundbreaking, like everyone understands human-in-the-loop. It's been around forever, but just reiterating on the importance of it.
- 4:45
So LLMs right now are just, you know, we know that they're good at, uh, generating content, and they do it really, really quickly. But without humans, they often fall short of real-world reliability, especially when nuance, empathy, and context is required.
- 4:58
Without human-in-the-loop, you're also not scaling productivity, you're scaling risk. And so a [REDACTED:marital_status] hallucination in isolation might not seem like a big idea or big deal, but at scale it could be dangerous.
- 5:09
Um, so, uh, the last point I want to make on that is also each time a, a human flags or corrects an output, it's a signal that we need to be using to help retrain, reinforce the models that we see.
- 5:21
And that's of course how, uh, you know, models are evolving over time. So I want everyone to think of human-in-the-loop not as like a safeguard, it, but as a feedback mechanism or feedback engine, if you will.
- 5:32
So every model output that goes through a human review, that feedback then gets fed for refinement and, and of course, uh, improving the model. And over time, this loop brings AI closer to real human expectations and behaviors.
- 5:45
But there is a challenge and, and speaking with many different teams, the challenge is that we don't have enough people to do all these reviews. Like model-graded evals and evals in general are great, but you still need to have the human element in there.
- 5:56
So most teams don't have enough people to review thousands of the different outputs that you have. And that makes it really, really hard to improve on the models and even measure, uh, the current state of the models that you have.
- 6:06
But there is good news, and so this might not seem like a, a clear path in some of the organizations that you all might be, be part of. Um, but we already have teams that are kind of trained to do this type of work of human-in-the-loop, and those teams are like your quality folks or CX folks within
- 6:21
the operations side of the business. And so especially when you look at call centers, they're already experts in doing the jobs that you need them to do, which is evaluate interactions at scale, spotting a, spotting edge cases and defining what really, what good looks like.
- 6:35
So in this age of A, uh, GenAI, their roles from my perspective is not shrinking, it's continuing to evolve. Um, and they're not just, right, they're not just measuring outputs anymore, uh, or outcomes.
- 6:47
Uh, I think they're really shaping, like, the future and what's next.
- 6:52
So as GenAI becomes, like, more embedded in the operations, quality is also evolving. It's no longer about auditing what's already happened, uh, but it's also about shaping what's gonna happen next.
- 7:02
So the shift change, uh, in the QA teams who are in this ops space, they're not just scorekeepers anymore. They're becoming model shapers, prompt testers, and AI performance monitors.
- 7:12
And if there's any group that's already built for this kind of work, as I mentioned before, you might want to look at your ops team and see if there's anyone in the QA contact centers that could also help expand your projects, uh, even further.
- 7:23
Um, they al- already know how to evaluate the nuanced conversations. They also know how to surface these edge cases like I mentioned before.
- 7:30
So talk a little bit about, like, the evolution of just quality in general or CX. Um, they've been for the most part focusing on auditing, coaching, compliance, uh, but as automation continues to scale, these roles aren't disappearing, but they're actually transforming.
- 7:46
When you look back, you know, 25 years ago, um, right, you know, the, the main roles were just QA professionals were listening to phone calls and evaluating interactions. But we see that as, like, continu- like, a continued evolution where automation is now coming, uh, into play, and these folks are, like, transforming their skill sets to help solve
- 8:05
for larger problems. [sniffs] The diagram that we have here kind of just shows, like, where the QA scope is today and how it's expanding to the GenAI space. At the core, we have what QA has always done, uh, which is auditing interactions, capturing quality behaviors.
- 8:20
But now, as you can see, we scope... The, the scope itself is expanding, and QA professionals are testing prompts now, and they're tagging outputs, and they're really helping shape the model, uh, behavior that we're expecting.
- 8:31
Uh, and there is a beauty in the GenAI space, uh, which is it opens the doors to non-technical folks. Uh, traditionally, uh, only the ML teams, the engineers are the ones that are heavily involved in the outputs and deciding what's gonna, you know, uh, come from it.
- 8:44
Uh, but now, um, uh, you don't need to know how to build, uh, the model pipeline to know what, you know, a good output looks like, uh, just like you don't need to know how to, um, make wine to be a good wine connoisseur.
- 8:54
Uh, the, the outputs or the expertise still matters and just as valuable. Uh, and then the contact center CX teams are, again, already equipped with this. And so, um, it might be something to, to look into and just see how you could expand that reach even further.
- 9:09
So one of the things Chris is talking about here is, like, something we have observed, like, working across a number of custom- customers as well, which is this stor- sort of emerging role of, like, the AI quality lead.
- 9:20
And importantly, like, I've almost never seen it actually called this, but it seems to exist at companies who are having, like, a lot of success in the GenAI space.
- 9:29
And I expect this role to actually become more formalized and gain more traction, right? Like, Chris and his team are an example of this role. Um, and importantly, like, people in this role can come from a variety of different backgrounds, product, ops, engineering.
- 9:43
But, like, the key attributes of what makes someone a good AI quality lead is someone who first and foremost has a deep, deep understanding of the customer need in the domain, and then importantly is a systems thinker and is able to, like, systematically think about how to diagnose and solve these quality problems.
- 10:01
What that looks like day to day is this person is doing a lot of these just kind of, like, this new skill set, labeling data, writing evaluation criteria, running experiments and tests and things like this, and prompt engineering.
- 10:15
And you'll notice that, like, there is something notably missing from that, which is, like, these are often not the people who are writing production code. But I think what has changed so significantly is writing production code is not the only way now that you can contribute in, like, a really hands-on way.
- 10:33
All of these things, like, given the right tool set and the right structure of your team, you can contribute to, like, this iteration loop, the prompt engineering, the evaluation, all this kind of stuff, without necessarily being the one, like, writing the production code day to day.
- 10:47
Chris is talking about, like, how you do this at scale, but we've also seen a lot of success where companies will have one person or two people in this role, um, especially when you have a smaller footprint, and that goes a long way, you know?
- 10:58
So I think what, what Chris is painting is, like, in larger enterprises and as you start to scale this stuff up, you really need, like, a meaningful quality team.
- 11:05
But you can do this by, like, kind of empowering a [REDACTED:marital_status] or couple individuals in this role.
- 11:13
Cool. So I'm just gonna leave you all with a, a, a few last points. You know, I think while human-in-the-loop is important, and, like, we've just spent the last few minutes talking about it, uh, if you don't have the resources, I think, you know, looking at high-risk, high-trust areas is abs- like, an absolute must.
- 11:28
And so insert human-in-the-loop at decision points and not just for show. The next piece is, like, continue to bring your co- like, your ops teams and CX teams into the life cycle early to help what, you know, define what good looks like, uh, so they can help build out golden sets and, uh, tests against real world, real
- 11:45
world edge cases. Um, and then this is another point I wanted to make too, is just, like, the, the fact of, like, launch is not the finish line. Uh, track performance, flag hallucination, me- measure impact, and iterate.
- 11:55
Uh, time and time again, I see teams, like, celebrating, which is great to see celebration of, like, a product actually launching or solution being actually launched. Uh, but the important part there that it's not the end.
- 12:04
It's the beginning, where you need to set up your teams to make sure that the outputs are performing how you expect them to perform. Um, and last but not least, scale is not just about tech anymore.
- 12:13
I think it's about people. And so leveraging QA, ops, uh, and support, and frontline teams, I think as a strategic partner in the GenAI space is gonna be what makes teams successful.
- 12:23
And if there's just one key takeaway today, it's that scaling GenAI isn't just a technical challenge anymore. It's an operational, uh, liability and responsibility. Um, when you embed quality and then human feedback into that loop, the right people, uh, into your GenAI systems, you're not just building faster, but you're also building better.
- 12:41
Um, so anyways, thank you, guys. Appreciate it. [audience applauding] [upbeat music]