← All AI Engineer talks

AI Engineer World's Fair 2024

Cohere for VPs of AI

About this talk

Cohere engineering director Vivek Muppalla briefs AI leaders on the company's enterprise AI offering, contrasting Command R and Command R+ and describing embeddings, reranking, retrieval-augmented generation with built-in citations, multilingual evaluation, and agentic tool use. He also discusses enterprise data privacy, intellectual-property indemnification, and deployment across clouds, private infrastructure, and on-premises environments before audience questions on text classification and model evolution.

Chapters

  1. 0:00Cohere's enterprise AI mission and research team
  2. 1:36Command R, Command R+, reranking, and enterprise evaluations
  3. 4:14Privacy, deployment flexibility, and enterprise benchmarks
  4. 5:46RAG citations, multilingual capabilities, and embeddings
  5. 13:27Audience questions on classification and model evolution

Talk transcript

  1. 0:00

    [on-hold music] Um, hey, folks. Uh, this is Vivek. Uh, super excited to chat with all of y'all.

  2. 0:16

    Uh, we'll give a, a quick talk about, uh, what Cohere is all about, and, uh, we'll make sure we have enough time to chat about, uh, your production challenges with these models or anything else you wanna chat about.

  3. 0:29

    Um, cool. Uh, quick intro. Um, so we are a leading data security-focused, uh, enterprise AI company. Uh, our focus is building trustworthy, uh, enterprise AI models, uh, with our partners for real-world business, uh, use cases.

  4. 0:45

    Um, and, uh, we work with a lot of, like, strategics a-across, uh, various clouds, uh, and are, uh, have a bit of a Switzerland play when it comes to where we can, uh, ship and deploy, and I'll talk to that, uh, a little later.

  5. 0:59

    Um, so we have a crack team of, uh, ML, uh, researchers and seasoned enterprise, uh, operators. Uh, Aidan, uh, who's our co-founder and CEO, was one of the, uh, authors on the, uh, seminal transformer paper.

  6. 1:12

    Uh, we have, uh, Phil, who's our chief scientist. Uh, he was an NLP lead at DeepMind and a professor at Oxford. Uh, Nils, uh, who leads a lot of our retrieval efforts, was a BERT expert.

  7. 1:25

    Uh, and then Patrick, uh, was, uh, the co-author on the RAG paper, which is a big, uh, focus for us at, uh, Cohere. So, um, fantastic team, uh, that's, uh, helping us build lots of amazing things.

  8. 1:36

    Um, so here's a quick overview of our product line. Um, so we have, uh, two main arcs, so the generative side and then the advanced retrieval models. Um, on the generative side, uh, we have, uh, two flagship models, which is Command R and R+.

  9. 1:52

    R is our, uh, workhorse, uh, super low-cost model, uh, that's great, great for most enterprise use cases. Uh, and then we have R+, which is, uh, your more powerful, larger model, uh, for more complex, like reasoning, tool use, RAG, uh, use cases.

  10. 2:09

    Um, and on the retrieval side, uh, uh, most people have used some form of an embedding model, um, and I'll get into that a little later in the talk.

  11. 2:18

    Uh, but, uh, the Reranker is something special, uh, that I haven't seen quite often in the market, uh, but we think, uh, it adds a lot of value, especially to your, uh, RAG pipelines.

  12. 2:28

    Um, so we'll get into that too. Um, so when we build at Cohere, uh, so we have, uh, five guiding, uh, ethos as to how, uh, we want to go about things.

  13. 2:37

    Um, the first is, uh, obviously everybody's testing their models in all sorts of academic benchmarks. Uh, but for us, what is really important is the performance on, uh, enterprise use cases.

  14. 2:48

    Um, so we've worked, uh, with a lot of our partners to ensure that we have an eval suite, uh, that is highly customized to enterprise, uh, use cases across, like let's say health, HR, finance, uh, and that's, uh, a bit of our goalpost as we ship, uh, e-each of these like model versions and, uh, we constantly benchmark,

  15. 3:08

    uh, on how we are performing at each of these industries, uh, and use cases that our customers care about. Um, and then when-- the next thing is ef-- all about efficiency and scalability, right?

  16. 3:18

    We're, uh, not particularly chasing the race for having the largest model out there, but what we really care about is the practical use of these models, right? How, how do these models get used, and how cheap is it, uh, uh, and how easy is it for you to run it as a customer?

  17. 3:35

    Um, the next big thing is obviously customization. Um, you know, as, as much as we'd like for all of these models to work out of the box, there's always a, a certain niche that, uh, customers want to customize this for.

  18. 3:49

    Uh, and we offer a variety of things, uh, some of which are pretty intrusive. Uh, we've helped our customers, uh, with taking our base model, uh, and retraining that with their enterprise-specific data, uh, for domain adaptation.

  19. 4:03

    We can do a full retraining of the model for you with your data. Uh, and then obviously the, uh, pretty typical last few layers, uh, retraining, which is self-serve on our platform.

  20. 4:14

    Um, data prominence and privacy is another, uh, big focus, uh, for us. Uh, we've, uh, worked quite a bit to ensure that all of the data that we've, uh, collected for building our models, uh, is, uh, meets up to the enterprise, uh, standards, uh, and we offer indemnification for any IP claims, uh, that you might run into

  21. 4:33

    as a customer. Uh, and then obviously we don't ever use any of your data to train our models. So, um, so that, that's a, uh, guarantee from us. Uh, and, uh, deployment flexibility.

  22. 4:44

    Uh, as I mentioned, we're available on pretty much every major cloud provider. Uh, and then we also allow you to deploy, uh, on-prem or in your own VPC, um, uh, wherever, uh, your compute and your data is, that's where we'll meet you at.

  23. 4:58

    Um, cool. Um, so just a quick look at, uh, uh, you know, your typical performance metrics. Uh, as I mentioned, uh, something that enterprises like repeatedly tell us is, uh, they care about like multilingual, they care about like RAG, um, and tool use for upcoming like agentic use cases.

  24. 5:16

    Uh, so a, a lot of our focus has been, uh, in these areas. Uh, these are some benchmarks from HotpotQA, Bamboodle, uh, Berkeley Function Calling, uh, that our models, uh, are quite, uh, good at.

  25. 5:30

    Uh, and, um, and another example of like how we try to innovate is, uh, we try to make sure that as we are building, uh, w- these stacks, we're incorporating all of the features that people care about out of the box, and they don't have to do extra work.

  26. 5:46

    Citations on RAG is a great example of this. Uh, for most people, you have to do a lot of work to actually build this functionality with like other APIs, uh, but this really comes out of the box with like Cohere's, uh, models and APIs.

  27. 5:58

    You don't have to do anything additional as a developer, uh, to build, uh, get citations and which is, uh, very important for any RAG-based like application. Um, uh, on the multilingual front, uh, we have one of the best performance when it comes to, uh, the FLORES multilingual, uh, evaluation.

  28. 6:19

    Uh, and we also have a bit of a secret sauce with our, uh, tokenizer, uh, which, uh, helps keep costs really low, right? Uh, and, uh, that's again, uh, TCO is again a very big thing for enterprises.

  29. 6:30

    Uh, and, uh, it allows our customers to take that same model and de- deploy across the globe on, uh, with their customers, uh, which is very important. Um, switching gears towards the embeddings models, uh, again, uh, given, uh, Nelson and his team, uh, have been innovators in this space for a while now, uh, and, uh, we've, uh,

  30. 6:53

    done quite a bit of work to make our-- make sure our performance is great on, like, noisy data and at a super low cost, uh, uh, in this particular space.

  31. 7:02

    Um, so we're, we're actually pretty excited about what our embeddings models, uh, can do, and this is al- almost always, like, one of the top things that, uh, our customers are, uh, excited about.

  32. 7:14

    Um, but embeddings, uh, is a pretty complex space and not without its, uh, challenge. Uh, so we, we try to build this, like, fun demo where we took all of the archive papers, uh, and we asked it a question, uh, "When was the attention paper, um, built by, uh-- paper published by Aidan, uh, Gomez," who's our founder?

  33. 7:33

    And we tried this across, like, a bunch of, like, embeddings model. So some common patterns that we see is, um, archive is a great example of, like, where you have different kinds of, like, data, right?

  34. 7:44

    You have, like, the title, when was the paper published, the dates. Uh, you have the various authors. You have, like, the actual paper itself. And in many ways, this represents the kind of data you might see in enterprises, right?

  35. 7:55

    Um, so when you, uh, actually build the embeddings for this, um, you get a fairly complex, like, vector space, uh, and your search queries might not actually, like, map neatly to this, right?

  36. 8:07

    Uh, and this is sort of like where our reranker comes in. Uh, so what our reranker does is once you have, uh, all of these, uh, retrieved documents, uh, it hel- it's a cross encoder that helps you, uh, rerank the, uh, retrieved set and make sure that that's the one that you send into your context with the

  37. 8:25

    generative model. Um, and, uh, here's the reranker in, uh, action. Um, so what this demo is showing you is, uh, we search, uh, for the transformer paper by Aidan.

  38. 8:38

    So you have three different types of, like, search patterns over here: a lexical search, then an embeddings-based search, uh, and, uh, the Cohere rerank-based search, right? Uh, and the, uh, various forms of these, like, retrievals obviously give you the responses, but they are stack ranked at different places in the retrieval set, which means the overall accuracy of

  39. 8:58

    your RAG system might be low. Uh, and rerank is what's helping you to make sure, uh, that isn't the case. Uh, this builds on top of, like, other things that people care about, like chunking strategies, uh, but making those more optimal.

  40. 9:12

    Uh, another impact of this reranker is, again, total cost of operation, because most of the expense for your models is coming in from the input tokens, right? Uh, and if you were able to, like, narrow in, uh, to the right, uh, uh, context and do that quickly, uh, you could pass in very minimal amount of context to

  41. 9:32

    your large language model, which drives on-- drives down your overall cost of, like, operation. Uh, and that, uh, is again, very important in the enterprise, uh, setting. Um, yeah, and when it comes to deployment options, like I said, we have our SaaS API, uh, that we can help, uh, you manage run your wo- uh, workloads.

  42. 9:52

    Uh, but then we're also on all of the major cloud AI services, Sa-Sage, uh, SageMaker, Bedrock, um, OCI, uh, and private deployment across all of these cloud providers and pri-- um, also on-premise deployments if, if that's, uh, something you care about.

  43. 10:08

    Um, security and privacy obviously is a pretty, uh, top of mind for us, so we make sure we're compliant with, uh, uh, the standards that our customers are often asking us, uh,

  44. 10:20

    for. Um, and then the last bit is just enabling, like, developers. Um, so we, uh, obviously have a, a pretty tight integration with things like La-LangChain, LlamaIndex. Uh, but we also have an open source, like, toolkit, uh, that comes out of the box, uh, with, uh, various forms of, like, connectors, um, and that lets you, uh, ingest,

  45. 10:40

    uh, data pretty easily into your systems and lets you have full control over the things you're, uh, building, um, and don't have to really, uh, look, uh, for a ton of different, like, options, uh, as you're developing your, uh, enterprise applications.

  46. 10:55

    So, um, that's it from me, and I'd love to take any questions or chat. And we also have Sandra here, uh, from Cohere, so she'll, she'll be happy to help.

  47. 11:06

    Thank, thank you very much.

  48. 11:07

    Yeah.

  49. 11:08

    May- maybe two questions. So on the first or second slide, you show your, uh, investors and the selected partners, right?

  50. 11:13

    Yeah.

  51. 11:13

    I saw Accenture, I saw McKinsey.

  52. 11:15

    Yeah.

  53. 11:16

    Can you explain a little bit, like, how that partnership, uh, work, right?

  54. 11:20

    Yeah.

  55. 11:21

    And then, and then the second question, like, you can skip if someone else, like-

  56. 11:23

    Yeah

  57. 11:23

    ... has another one, is can, can you, like, maybe without disclosing, uh, customers and all, give us a, like a few samples of where you- ... clients pick to Cohere versus other solutions and, and kind of why?

  58. 11:37

    So explain where you win, right, on the, on the enterprise world.

  59. 11:40

    Yeah, absolutely. Uh, happy to chat about that. Um, so, uh, I think the typical challenge with all of these enter- um, enterprise, like, generative AI models is the last mile challenge, right?

  60. 11:50

    Like, uh, there's, uh, so many different, like, arcs of, like, customization that's needed, uh, with the enterprises. Uh, and, uh, there's a lot of, like, traditional, uh, like, players who've been around for a while and have great relationships with the enterprises, uh, have a deep understanding of, like, the various, like, uh, business domains.

  61. 12:10

    Uh, that's where, uh, the McKinsey and Accenture and all of these companies come in. Um, so they've been, uh, really, uh, helpful for us to co-develop, like, the product.

  62. 12:20

    Like, make sure that we're able to effectively, uh, bridge that, like, last mile gap with them. Um, and yeah, that... hopefully that, that helps. Uh, and then onto your second question, I would say, um, uh, in terms of, like, winnability, I think, like, the main, uh, aspects that has been, uh, resonating a lot with our customers is

  63. 12:39

    this control over the data and control over the compute, right? Uh, given we're available pretty much everywhere, uh, a lot of, like, the customers care about that private cloud deployment.

  64. 12:49

    The ability to fine-tune in that, uh, private cloud environment, uh, which is pretty big, uh, for a lot of people, and making sure that their enterprise data does not leave their own ecosystem.

  65. 13:00

    Um, and specific examples for that have been, uh, companies in, uh, let's say HR or healthcare or even, uh, folks who are trying to take their in-house, like, code and, uh, build a custom model that's, uh, working with, uh, their code base.

  66. 13:16

    Uh, those are the styles of applications, uh, that, uh, we've seen a lot of like, uh, impact and success with.

  67. 13:21

    Thank you.

  68. 13:22

    Yeah.

  69. 13:27

    So I'm curious, um, uh, for text classification, well, what is the kind latest best practice? I see on your website you have a classify endpoint, right, build a classifier.

  70. 13:38

    Uh, is that still the recommendation?

  71. 13:40

    Yeah. Uh, that's a, that's a great question. I, I... The, the way I like to think about, uh, these things is there's always the arc of like, uh, um...

  72. 13:49

    A- and I think, like, Jerry in his earlier talk did this, right? Which was what phase of, like, development you're in, uh, if you're trying to, like, prototype or you're trying to productionize, and what's your scale of, like, a production setting, uh, which, uh, is important to consider for these things.

  73. 14:04

    Uh, for, uh, if you're just trying to, like, get off the ground, like, quickly, I would say just using the generative model off the shelf is obviously always great.

  74. 14:13

    It, it gets you off the ground really quickly. Uh, when you're trying to, like, productionize something, that's when I would start thinking, "Hey, do I need a bespoke model?

  75. 14:22

    Uh, or, like, the, the general model is good enough," and what are sort of, like, the cost of operation, like, differences and, like, the cost of, like, maintenance dif- differences and also the scale, right?

  76. 14:32

    Like, I mean, if, if you're going to try to do something that's, uh, you know, tens of thousands, like TPS, uh, then having a, a purpose-built, like, model for that is the route I'd go, uh, versus, you know, uh, a more heavy general model, uh, which might serve other, other needs.

  77. 14:48

    Um, so yeah, both of those are good options depending on what you're trying to accomplish.

  78. 14:53

    Just a quick-

  79. 14:53

    Yeah. Yeah

  80. 14:53

    ... uh, quick follow-up. If we actually train a classifier with you, is that also a transformer model, just more specialized?

  81. 15:00

    Yeah, exactly.

  82. 15:01

    Okay.

  83. 15:01

    Yeah.

  84. 15:02

    Got it. Thank you.

  85. 15:05

    Yeah.

  86. 15:06

    No, I think we're done. Oh, no, we've got one more question. Here we go.

  87. 15:10

    Just a quick question. When we were looking at Cohere, you know, probably early last or middle of last year-

  88. 15:18

    Mm-hmm

  89. 15:19

    ... um, one of the challenges with the models that we found were the input context size limit-

  90. 15:24

    Yeah

  91. 15:24

    ... were quite small. H- how has that evolved a- as you guys have sort of created the next, you know, sets of models on your side?

  92. 15:31

    Yeah, that's a great question. So our latest generation models are, uh, fairly competitive, 128K, uh, context input, uh, windows. Uh, and we're constantly looking to, uh, figure out how to, like, up them.

  93. 15:42

    Um, so context window, I would say, should be not a problem for, like, most applications, uh, at the moment.

  94. 15:50

    Okay.

  95. 15:50

    Yeah.

  96. 15:53

    Cool. Well, we're actually slightly early, but yeah, I'd like to thank you for giving the talk.

  97. 15:58

    Yeah.

  98. 15:58

    Thank you, Vivek.

  99. 15:58

    Absolutely. Thank you so much. Yeah.

  100. 16:00

    Thank you.

  101. 16:00

    Yeah.

  102. 16:01

    Thank you. [audience cheering] [upbeat music]