← All AI Engineer talks

AI Engineer World's Fair 2024

What It Actually Takes to Deploy GenAI Applications to Enterprises

About this talk

Arjun Bansal of Log10 and Trey Doig of Echo AI explain how enterprise conversation-intelligence systems ingest and normalize customer interactions, apply configurable LLM pipelines, and surface insights across entire conversation datasets. They discuss earning customer trust through accurate evaluations, using Log10 AutoFeedback to address LLM-as-a-judge weaknesses and track hallucinations, and supporting production-scale throughput with self-hosted, domain-adapted models.

Chapters

  1. 0:00Customer conversations and the opportunity for comprehensive AI analysis
  2. 5:05Enterprise data ingestion, LLM pipelines, and customer trust
  3. 7:45Log10 accuracy infrastructure and automated evaluation
  4. 14:51Product demonstration, self-hosted models, and hallucination tracking
  5. 21:07Closing remarks

Talk transcript

  1. 0:00

    [on-hold music] So Echo AI. So Echo AI is a, uh, uh, bringing this, like, fundamental new technology into the world of customer support and, and customer-facing teams.

  2. 0:24

    Uh, you've probably all seen this metaphor once before. You know just the tip of the iceberg. There's so much that lies underneath, and this is especially true with, with companies that are dealing with exceptionally high volumes of customer interactions.

  3. 0:36

    So whether that be customer support or sales or s- customer success teams, any kind of conversation that you're having with your customers is an opportunity to learn more about your business and what your customers need from it.

  4. 0:47

    So, you know, at the surface, you probably get the signals around the things that are going wrong or the things that are just routine to, to kinda handle. And, and most enterprises have a good idea of the, the approximate kinda set of categories that they, they generally have to deal with on a normal basis.

  5. 1:04

    But there's just so much that lies underneath and it's, it's been virtually unlocked by just pure manpower that's required to look at these conversations. Um, and there's this sort of cycle that happens with these types of companies and, you know, everyone's kinda heard the startup mantra of just be obsessed about your customer, right?

  6. 1:25

    Just do everything you can to listen to the customer, understand them to the, the best possible degree you, um, you're capable of at, at, at your scale. But as you grow and you get bigger and you get more customers, a lot of that becomes this, um, this cycle that, uh, where you just, like, you kinda lose touch

  7. 1:43

    with, with what they're saying, what they're telling you every day. Um, and that only really comes with being successful, right? Uh,

  8. 1:50

    at like a, you know, a, a meaningful or, like, a manageable size of customers, you can typically use, uh, either, you know, staff or your sales team, your customer support team to sort of derive those insights and better, um, you know, navigate your, your company towards, uh, greater revenue growth.

  9. 2:05

    But when you get too big to the point where that's, like, virtually impossible just due to the scale of conversations that you have to deal with, uh, there, there becomes this issue where you're just-- there's so much that is just happening right underneath you without your, your, your knowing.

  10. 2:18

    So the way we think about it, um, and this-- the, the first three columns here are, are effectively what every enterprise, every, every company of scale is trying to do, where they do manual reviews.

  11. 2:31

    So they do these, like, small sample sizes. They try to collect, I don't know, five percent of the conversations. They run through an evaluation of it to check for, you know, maybe certain compliance things, like how well did the agent handle it or perhaps are there any sort of common subjects or, or, um, themes in these conversations.

  12. 2:49

    At the end of the day, you're just sampling and, and it's-- leaves everyone unhappy at the end of the day. They understand that it's not a very, uh, kinda accurate process.

  13. 2:57

    So then you, like, involve engineers and you start building these scripts. You start doing retroactive analysis. You pull this data from all the different systems that your, your customers are interacting with you on and you're, you know, doing these, like, kinda long-form analysis.

  14. 3:10

    And y- you're, you're ultimately, um, that, that, that can typically, like, derive insights that you know that you wanna, uh, track in a, in a forward-moving time. So then you, like, move into this situation where you start building software to, to look for very specific things in every single conversation and everything's just so retroactive.

  15. 3:29

    Like, everything is after the fact. It's already after fires have formed. You, you have no sense of where the smoke is and that's where generative AI-- sorry, generative AI can come in and really transform this process because generative AI unlocks this, this amazing capability of one hundred percent coverage.

  16. 3:47

    So now rather than doing that sampling or just, like, looking for the things that you know to look for, you can use generative AI to surface all the things that you didn't know to look for and look at everything all at once.

  17. 3:58

    So here, here's a great example. This is just a small interaction between some random, uh, you know, agent that's handling a, a chat from one of their customers. And in this one message that comes from the customer, you're able to find things like why is this customer talking to me?

  18. 4:12

    That'd be your intent. Uh, you know, what are the aspects of our business that are, that are really kinda at the root of this? So, you know, routers and it was broken on delivery, so maybe you have a supply chain issue.

  19. 4:22

    You've got the basics, things like sentiment, understanding, like, not only the sentiment of your customers, but also the sentiment of your, your, your representatives, which is, you know, maybe more important.

  20. 4:31

    Um, and then at the end of the day,

  21. 4:34

    with each of these messages, you can effectively extract so much more and so much depth that's ever been available to, to pull from just a single conversation and that's what our, our, our, our platform seeks to provide.

  22. 4:47

    Here's a great example from one of our customers. Uh, uh, we don't have to read this, but, uh, wine enthusiasts, they, they, they shipped a new unit. It was a brand-new, uh, uh, version of their, their...

  23. 4:56

    They, they, they, they sell, like, really high-end wine refrigerators, so, um, their customers are, are spending a lot of money. It's a, it's a very special relationship that they have with their customers.

  24. 5:05

    They want those customers to buy more fridges. A lot of these customers are retail locations. So, uh, one of the, the, the insights we were able to surface for them was a, you know, a, a, a defect in their manufacturing process.

  25. 5:18

    Uh, something that, uh, you know, could have gone on for weeks and weeks and weeks and become a much bigger problem versus what our, our, our platform was able to surface in real time.

  26. 5:29

    So yeah. How, how does this all work? Uh, it all really starts with gathering all of those conversations. This is kinda like non-AI boring stuff. We're, like, connecting to a bunch of different, like, contact systems and ticket systems and so forth.

  27. 5:41

    Uh, we pull all that in and we normalize it. We, we make it super clean, uh, ready to go and compressed to, to pass it into LLM prompts. And then we have dozens of these pipelines that are assessing these conversations in, in, uh, you know, a variety of different ways, all of which are configurable by the user.

  28. 5:56

    So the user, the customer can come in, they can tell us, like, what they care about, what they, what they're looking for, and they'll actually work with us to write these prompts and eventually they write the prompts themselves and manage it over time.

  29. 6:08

    And why this is like ultimately most important is when you are dealing at the accuracy-- uh, sorry, when you're dealing at the like enterprise scale, they're ultimately most concerned around accuracy.

  30. 6:20

    Uh, there's a, there's a, I think, a huge, uh, hesitation right now in the market around accuracy. I can deploy generative AI to try to understand these conversations, but do I really trust the insights?

  31. 6:31

    Is it gonna be better than what my, my business analysts are doing, or my, my CX leaders, my, my, my VP of sales? Does this system, does this technology, is it really capable of giving me insights that I ultimately trust?

  32. 6:43

    So for us, it's important that we establish trust from the very beginning. So when we bring on a customer, within seven days, we try to introduce them to like, "Okay, here are the insights."

  33. 6:52

    And then from there, we work towards a place where we're, you know, sort of a, a-- We like to say ninety-five percent accurate, but, you know, it's a lot of sampling, a lot of figuring out.

  34. 7:02

    But virtually, we wanna create that trust with the customer 'cause that's ultimately what's gonna get them to keep renewing and, and be a customer for a longer period of time.

  35. 7:12

    So Log10 plays a huge part in our ability to do this. So not only have they created a, a huge amount of, of features and capabilities that allow for our engineering team to build faster and to understand, you know, the quality of the code that they're writing, but maybe more importantly, our solutions engineers, our implementation engineers, who

  36. 7:32

    are actually working hand-in-hand with the customer, who is telling us that, like, "This isn't s- this isn't exactly right." Um, and Log10 is, is, is a kind of a go-to tool for p- managing that process, so...

  37. 7:45

    Jump in the back there. Thanks, Trey. Um, so really excited to share with you what we've built at Log10. Uh, so we're basically an infra layer to improve LLM accuracy for your AI applications.

  38. 7:58

    We started with this vision of building self-improving systems, so having LLM applications that can improve prompts and models themselves to ultimately drive accuracy improvements. Obviously, we're not there as a field yet, but that's the vision we're driving towards, and I'm excited to share some of the work we've done along that path.

  39. 8:21

    So today, measuring and improving LLM accuracy is hard. Uh, you've probably had this experience if you've tried to deploy a prompt and kind of try to YOLO accuracy in prod.

  40. 8:34

    Um, and you've probably heard in the news of these, um, instances where there was the Air Canada chatbot, which hallucinated a refund policy, and a judge in Canada f- um, forced them to honor the policy that was made up, so obviously causing some financial problems as well as brand image problems.

  41. 8:55

    Uh, there was the case of the Chevy Tahoe dealership chatbot, which was convinced to sell a truck for a dollar.

  42. 9:01

    Uh, we've run into similar issues where we were having some, uh, issues with a product, and a chatbot told us to, uh, play a game while we wait, uh, kind of missing our emotional state in that moment.

  43. 9:14

    And, uh, there's also issues with semantic search engines where, uh, because they're doing this kind of RAG-based lookup, uh, they sometimes just miss out on common sense, even though it may not be explicitly present in the source documents that they're pulling up.

  44. 9:28

    And so, uh, people have tried to use human review as an alternative to, um, a-as a way to sort of look at the output of the LLMs, and that's ultimately become the gold standard in terms of getting LLM apps into deployment, but it's time-consuming and expensive.

  45. 9:47

    And as an alternative, uh, we've tried to use, uh, AI-based review, so, um, LLMs-as-a-judge, for example. But obviously, this also has a lot of issues with accuracy. Uh, people have found that models tend to prefer their own output.

  46. 10:03

    Uh, they exhibit positional bias, so just the ordering of, um, if you present, uh, an option first, it might just prefer that over the second one, even if the second one's better.

  47. 10:14

    Uh, they have verbosity bias, a bias towards diversity of tokens, and so forth, so, uh, kind of failing in these trivial ways. And so with Log10, we did a bunch of algorithms research work to try to address this issue and ask this question of, could we get the accuracy of human review with the speed and cost advantages

  48. 10:34

    of model-based review? And that's exactly what we've solved for with our auto feedback system. So I think even in the previous talk, you saw a graph kinda like this, where when you use LLMs-as-a-judge, uh, the predicted feedback can often be, uh, just set at one score, regardless of what the ground truth is.

  49. 10:56

    But with our auto feedback system, you get that much nicer, much better correlation between the, uh, predicted feedback and the actual feedback.

  50. 11:06

    And just to kinda motivate what do you do once you have this measure of accuracy, so some of the downstream ways in which you can use the auto feedback is for things like, uh, monitoring, so you get like an ongoing quality signal on how well your LLM application's doing.

  51. 11:23

    It can be used for triaging, so you make the most optimal use of the limited human resources you might have. And it can also be used for curating datasets, high-quality datasets, for things like automated prompt improvement and fine-tuning, which can ultimately improve the accuracy of your LLM application.

  52. 11:45

    Uh, so next I'll say a little bit more about what the system is. Um, so for some of the, uh, experiments we did as part of the research, we came up with these three different ways of building auto feedback models.

  53. 11:58

    Uh, so on the left, we have the ground truth datasets, which consist of the input and output, some kinda grading rubric, which you might give to a human for review, and, uh, their feedback.

  54. 12:09

    And we had three variations where we could build these models with few-shot learning or with some kind of fine-tuning, um, and then finally, uh, creating these models with some bootstrap synthetic data and then fine-tuning the auto feedback model.

  55. 12:25

    And I won't have time to go into all the details. We published, uh, some of this work in a blog post, which is available on our Substack, and there's a QR code there.

  56. 12:35

    But, uh, in summary, uh, for a summary grading task where we use the TLDR dataset, we were able to get a forty-five percent improvement in evaluation accuracy by going from aggregate to, uh, annotator-specific models, uh, going from GPT-3.5 to GPT-4 as the base model, uh, going from few-shot learning to fine-tune models and, uh, with the use of

  57. 12:59

    the bootstrap synthetic data. Uh, our approach was also very sample efficient, so by using the bootstrapping approach, we were able to achieve the accuracy of almost as if we had one thousand, uh, ground truth labeled examples with just using fifty ground truth examples.

  58. 13:19

    So much faster to get started and not needing as much data to get to that high level of accuracy with the evaluation model.

  59. 13:29

    Uh, we also extended this in a follow-up blog post to open source models. So, uh, we were able to match the accuracy of GPT-4 and GPT-3.5 fine-tuned evaluation using Mistral 7B and Llama 70B Chat, and that's in a follow-up blog post, which is accessible on this link.

  60. 13:52

    And so once we were able to show that we're able to get high confidence in the eval models that we were building in this way, uh, we set it up for deployment within this auto feedback module on our platform.

  61. 14:04

    And, uh, just kind of zooming out, we have, as part of the Log10 platform, a fundamental LLM Ops offering, which includes things like logging, debugging, and evaluation, as well as auto-tuning features to do prompt optimizations and manage your fine-tunes.

  62. 14:20

    And we have a seamless one-line integration which sits in between your LLM application, uh, and your, um, LLM SDK. Uh, we have integrations into many of the common LLM SDKs, including OpenAI, Anthropic, uh, Gemini, and a few of the open source ones, and we integrate with frameworks as well.

  63. 14:40

    So next, I'll hand it back to Trey for a demo.

  64. 14:51

    All righty. Let me, uh, just make this slightly bigger.

  65. 14:58

    Well, okay. So, uh, here's, uh, uh, Echo AI. So Echo AI is, uh, like I said earlier, we, we, we connect into all of the different, um, channels that your customer conversations are coming in.

  66. 15:08

    We transcribe them, we clean them, and then we allow you to, uh, basically, like, codify all the different things that you're looking to, uh, kinda analyze these conversation against, uh, as well as offering a product that is purely generative and is, uh, kinda tasked with surfacing those insights in ways that you, um, you, you basically know the

  67. 15:24

    question, like, "Why are my customers canceling?" And then we, uh, generate, uh, like a, you know, uh, an ontology of different reasons for why that's happening. So let me give you an example of how, uh, we are, um, making use of this feedback tool.

  68. 15:36

    So if I, like, jump into one of these... This is a demo account, so, uh, all these conversations are, are generated, so they're kinda silly. But nonetheless, uh, you can see here we've got this, uh, example, uh, phone conversation that comes from a customer.

  69. 15:48

    Their, their TV was broken. Uh, one of the most, you know, simple, uh, I think on the surface, uh, features would be summarization of the transcript. Uh, we actually rely on summarization for, uh, a variety of different downstream, uh, sort of analysis, so it's really important to us to, uh, kinda ensure high levels of accuracy.

  70. 16:06

    Not only that, uh, this is also, um, uh, you know... Because of that, uh, kinda technical reason, we have, like, an, an immense number of prompts and throughput that has to get through LLMs, so we do quite a bit of self-hosting and are constantly training, uh, new models to better handle different, uh, domains of, of our customer

  71. 16:24

    base. So here's an example. Like, you can see that this, this summary came in. It's pretty decent. Uh, all of these other sort of insights that you see here are all being generated by, um, various different models and pipelines, all of which are being created from LLMs.

  72. 16:39

    So, uh, each of these things are, are questions and specific, um, sort of, uh, requirements that the customer is providing us. So it's really hard to stay on top of accuracy and quality as a result.

  73. 16:51

    Basically, every customer is different. Um, so what we've done is we've leveraged Log10 not only for the ability to, um, very-- Sorry. Very quickly, uh, uh, allow our, our engineers and solution engineers to go in and actually understand, like, what was the, the generated prompt for all of this pretty standard stuff, as many of you know.

  74. 17:10

    But m- which maybe, um, kind of more interesting is the ability to automatically generate feedback for these things. So we've, we've created a criteria that analyzes each of the, the summaries that we create against very specific, uh, user-defined, like, uh, were defined by us, uh, criteria for how q- how good of a quality, um, the summarization was.

  75. 17:29

    So in this case, uh, the-- It actually looked pretty good, and there's, there's one, uh, point deducted. You can actually read down here why it deducted that point. But let's say that, like, I wanna provide a kinda human override here.

  76. 17:42

    I could just come in here and change the point value and, uh, accept it. Uh, and why that's useful for us is, like I said, uh, it's, it's critical that we continue, um, a process that allows our solution engineers to...

  77. 17:55

    This is an example, one against Mistral, um, to kinda give us, like, really high fidelity, uh, human-provided feedback in a way that, uh, virtually is, uh, effortless for them to do so, because what we're ultimately trying to do is collect as much of that as possible.

  78. 18:09

    Uh, and Log10 has been, uh, I think, a great tool in not only, uh, making that possible for the solution engineers and our engineers to do, but also automatically doing so behind the scenes.

  79. 18:18

    It's really ch- changed our, um, our processes towards our fine-tuning datasets.

  80. 18:25

    Thanks. Actually, I'll show one more thing, uh, to give an example of, like, how it's, uh, maybe more useful to an engineer. Here's another, uh, conversation where, uh, the summarization failed.

  81. 18:36

    Uh, we, we have just a, uh, reiteration of the instructions. That looks, that looks like the system prompt. Um, no idea why. Let's... We can, we can actually g- kinda go in here and look why.

  82. 18:47

    Okay, well, there's, there's a good reason why. Um, and you can see that it's been graded accordingly, uh, which is, you know, uh, exactly what we would expect. So we, we've been able to track, uh, hallucinations via this process.

  83. 19:00

    We've been able to see model drift, uh, in a meaningful way, in a data-driven way that we previously were unable to do. You know, a lot of it has been, uh, sort of, um, requiring humans to sample these things and give feedback on that.

  84. 19:11

    So this has been a, a huge, uh, tool in our, like, maintenance of, of, of achieving the, the utmost trust that we can, uh, kinda retain with our customers.

  85. 19:22

    Awesome. Thanks. Great. So maybe just in the interest of time, I'll skip forward. So, uh, one of the big, uh, achievements we were able to get was, uh, using this Auto Feedback approach, get a 20 F1 point improvement in accuracy in one of the use cases with Echo AI, and we publish all of the details in a

  86. 19:44

    case study that's accessible there. And, uh, just in terms of getting started, obviously, we cannot share, um, you know, customer data, uh, here, so we created this new summarization app, which kinda, uh, shows an example of a summary grading, uh, application, as well as, um, has a live version of the website that you could check out.

  87. 20:04

    Um, if you wanna, uh, take a picture of those QR codes, um, you can adapt this for your use cases, for your tasks, and try out Auto Feedback yourself.

  88. 20:15

    Uh, we also have an SDK, so everything that was shown in the UI is available programmatically from, uh, Python and SDKs in other languages. Uh, the feedback type is pretty flexible.

  89. 20:27

    Uh, we have a bunch of recipes that you can get from our GitHub, as well as notebooks to get started as well. Um, I'll skip over some of the mechanics, but basically goes over how you create the task, uh, create feedback, run the Auto Feedback locally o- on a simpler model, and then fetch feedback from a more

  90. 20:49

    complex model that runs on our cloud. And maybe just in the interest of time, we'll, uh, I think, Trey, you already covered this, but, um, yeah, invite you to use our platform, and, uh, for a limited time, you can use even the more advanced models on our platform for free.

  91. 21:07

    Thanks a lot.

  92. 21:08

    Thank you. [audience applauding] [upbeat music]