AI Engineer World's Fair 2024
Cohere for VPs of AI
Read the talk
Cohere for enterprise AI: retrieval, control and the cost of production
Enterprise AI depends on more than model quality: retrieval, customization, deployment control and operating cost determine whether a useful prototype can become a production system.
From a talk by Vivek Muppalla
Before you start: Basic familiarity with language models and document search will help; retrieval-augmented generation and reranking are explained as they appear.
What does an enterprise need from a model?
How do you put an AI model to work on business data while keeping control of where it runs? That is the starting problem for Vivek Muppalla’s introduction to Cohere: trustworthy enterprise models, developed with partners, that can deploy across cloud providers rather than requiring customers to commit to one infrastructure ecosystem.
The research background helps explain the product emphasis. Muppalla points to cofounder Aidan Gomez’s coauthorship of Attention Is All You Need, Phil’s background at DeepMind and Oxford, Nils’s expertise in BERT, and Patrick’s coauthorship of Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. Generation and retrieval are both central to the resulting lineup.
The generative family has two main choices: Command R, presented as the inexpensive enterprise workhorse, and Command R+, a larger model for more demanding reasoning, tool use and retrieval-augmented generation, or RAG. The retrieval family supplies embeddings and Cohere Rerank, which adds a relevance-ranking step to a retrieval pipeline.
| Component | Role in the application |
|---|---|
| Command R | General enterprise generation at lower cost |
| Command R+ | More complex reasoning, tool use and RAG |
| Embeddings | Retrieve candidates by semantic similarity |
| Rerank | Reorder retrieved candidates for the query |
This separates two decisions that are easy to conflate: which model should generate the answer, and which information should reach that model.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Evaluate the workload, then design for its constraints
Cohere’s development principles begin with enterprise-specific evaluation. Academic benchmarks provide one view of model performance; the stated release goal is an evaluation suite built with partners around health, HR, finance and other customer workloads. Repeated evaluation across model versions is meant to reveal whether a release improves the tasks customers actually need.
The next principle is operating efficiency. The objective is not simply to produce the largest model, but to make the model useful, affordable and straightforward to run. Customization follows from the same practical concern: a broadly capable model may still lack the behavior or domain knowledge a particular enterprise needs. Muppalla describes domain adaptation by retraining a base model on enterprise data, including full retraining, and contrasts that with the then-available self-service option he describes as retraining the last few layers.
Data provenance and privacy form the fourth principle. Muppalla says Cohere applies enterprise standards to its training data, offers indemnification for IP claims, and does not use customer data to train its models. That assurance concerns Cohere’s own model training, distinct from customer-requested adaptation. It is also a historical assurance: the later Enterprise Data Commitments distinguish SaaS training opt-out controls from private and third-party cloud deployments where prompts and generations are not sent to Cohere.
The fifth principle is deployment flexibility: major cloud providers, a customer’s own VPC, or on-premises infrastructure. Together, the five principles turn model selection into a system decision—task performance, operating efficiency, customization, data handling and the location of compute all have to fit the application.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
RAG citations and multilingual operating cost
The capability priorities follow the enterprise workloads: multilingual support, RAG and tool use for agentic applications. Muppalla names HotpotQA, Bamboogle and Berkeley Function Calling as evaluations where Cohere performs well. These are qualitative performance claims in the presentation, not a basis here for a numerical comparison between models.
A more concrete integration feature is built-in RAG citations. Cohere’s models and APIs can return citations with an answer, reducing the work required to connect generated statements to supplied sources. That does not remove the application’s responsibility to retrieve documents and construct the request; it moves citation generation into the model/API path instead of requiring developers to build that behavior themselves.
For multilingual applications, Muppalla cites strong FLORES performance and attributes lower operating costs partly to the tokenizer. Tokenization matters economically because the amount of text represented by a token affects the token volume processed for a workload. His practical objective is to let a customer use the same model across geographies while controlling total cost of ownership.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Finding the right attention paper
The retrieval demonstration makes the pipeline problem tangible. After describing embedding quality on noisy data and low cost as product strengths, Muppalla introduces a search over arXiv papers: when was the attention paper associated with Aidan Gomez published? The query sounds simple, but the records combine titles, dates, author lists and full paper text. That mixture resembles enterprise data much more closely than a collection of uniformly formatted passages.
An embedding represents that mixed content in a vector space. A query asking for a particular author’s paper and its publication date may not align cleanly with the representation of the whole record. A semantically related result is not necessarily the precise document needed to answer the question.
Reranking separates candidate discovery from final relevance ordering. Cohere’s reranker is a cross encoder: it evaluates a query together with each candidate document and reorders the already retrieved set before the generator receives its context. The procedure is:
- Retrieve candidate documents using the existing search system.
- Score the candidates against the user’s query with the reranker.
- Select the most relevant candidates for the generative model’s context.
The important boundary is the retrieved set. Reranking can improve its ordering; it cannot promote a document that the initial search never returned.
The demo compares three columns: lexical search using OpenSearch, embedding search, and Cohere Rerank. All can return results for the transformer-paper query, but the relevant document appears at different positions. In the displayed rerank column, Attention Is All You Need appears first. That ordering matters because a RAG application may pass only a small part of its retrieved set to the generator.
Reranking complements chunking rather than replacing it. Chunking determines the units available for retrieval; reranking helps choose among the resulting candidates. A small Python function makes the handoff explicit: given the arXiv candidate records and scored candidate indices, preserve the records while selecting them in relevance order.
python
from collections.abc import Mapping, Sequence
from typing import Any
def select_context(
candidates: Sequence[Mapping[str, Any]],
scored_indices: Sequence[tuple[int, float]],
limit: int,
) -> list[dict[str, Any]]:
if limit < 0:
raise ValueError("limit must be nonnegative")
ranked = sorted(scored_indices, key=lambda item: item[1], reverse=True)
return [dict(candidates[index]) for index, _ in ranked[:limit]]
Here, scored_indices contains the reranker’s candidate indices and relevance scores. Keeping each selected record intact preserves its title, authors, date and text for the answer-generation step. The function performs selection; the cross encoder supplies the relevance judgment.
Better selection also changes the cost calculation. If a small set of relevant documents is sufficient, the generator needs fewer input tokens. Muppalla emphasizes input-token expense as a major cost driver, but whether it dominates—and whether reranking lowers the total bill—depends on the workload and the reranker’s own cost. The mechanism is narrower and more useful than a universal savings claim: send less irrelevant context to the generator.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Putting the pipeline inside the enterprise
Deployment options extend from Cohere’s managed SaaS API to cloud AI services including SageMaker, Bedrock and OCI, alongside private cloud and on-premises deployments. Muppalla also describes compliance with standards customers request, without naming particular standards in this discussion. The architectural choice is where the model runs relative to the organization’s data, compute and security requirements.
Developer integration is the other half of that deployment story. Muppalla names LangChain and LlamaIndex integrations, then describes an open-source toolkit with connectors for ingesting data. The intended benefit is a shorter path from enterprise data sources to a working application while retaining control over the system being built.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Partners, private fine-tuning and customer choice
The first audience question pairs two concerns: how do partnerships with Accenture and McKinsey work, and why do customers choose Cohere over alternatives? The first answer concerns the last mile between a general model and a deployed business application. Enterprises need customization, and established partners bring existing customer relationships and knowledge of the relevant business domains. Muppalla describes their role as co-developing products and helping close those implementation gaps.
For customer choice, his emphasis is control over data and compute. Private deployment is useful not only for inference but also for fine-tuning inside the customer’s own cloud environment, so enterprise data can remain within that ecosystem. He gives HR, healthcare and models adapted to proprietary codebases as examples of application categories where that control matters. These are descriptions of customer needs, rather than named case studies or quantified competitive wins.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
When should classification use a specialized model?
The next question asks whether the Classify endpoint is still the recommended way to build a text classifier. Muppalla answers through development stage and production scale rather than treating an endpoint as a universal recommendation. An off-the-shelf generative model offers a fast starting point; moving into production creates a reason to reassess whether a general model is adequate and economical.
| Situation | Decision to examine |
|---|---|
| Early prototype | Use a general model to get started quickly |
| Production workload | Compare quality, operating cost and maintenance |
| Very high-volume classification | Consider a purpose-built model |
Muppalla uses a hypothetical workload of tens of thousands of TPS to motivate a purpose-built classifier, not to report measured throughput. A heavier general model may serve several needs, while a specialized model can be a better fit for a narrow, repeated task. Both remain options; the decision depends on the workload and the cost of maintaining the solution.
A follow-up asks whether training such a classifier still produces a transformer model, just a more specialized one. Muppalla confirms that it does. Specialization here changes the model’s task and operating role; it does not imply abandoning the transformer architecture.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A larger context window, with retrieval still doing useful work
The final question returns to a concrete limitation: an audience member had evaluated Cohere in the previous year and found the input context limits small. Muppalla says the then-current generation offers 128K-token input context windows, with further expansion under investigation.
His assessment is that this capacity should suffice for most applications at that point. Capacity, however, is not a guarantee of answer quality across the entire window. The earlier retrieval demonstration remains relevant: having room for more documents does not make every document useful, and selecting the right context still affects what the model can answer and what the application costs to run.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
From the talk
The original Transformer paper, coauthored by Aidan N. Gomez and first submitted to arXiv on June 12, 2017.
Patrick Lewis and colleagues describe generation that combines a pretrained model with retrieved Wikipedia passages.
Further reading
- Command R's historical release noteDocumentation
The 2024 release note describes Command R's RAG and tool-use positioning and 128k context capacity.
- Cohere ToolkitRepository
An archived RAG application repository with configurable retrieval, model providers and data integrations. Archived May 14, 2026.
Updates since the talk
A current Python walkthrough of retrieval, reranking, document-grounded generation and returned citations.
- Cohere model and API deprecationsDocumentation
Lifecycle notices covering historical Command models, Classify, connectors and fine-tuning capabilities.
- Enterprise Data CommitmentsDocumentation
Cohere's deployment-specific data handling and SaaS training controls, updated December 5, 2025.
Read the complete timestamped transcript
- 0:00
[on-hold music] Um, hey, folks. Uh, this is Vivek. Uh, super excited to chat with all of y'all.
- 0:16
Uh, we'll give a, a quick talk about, uh, what Cohere is all about, and, uh, we'll make sure we have enough time to chat about, uh, your production challenges with these models or anything else you wanna chat about.
- 0:29
Um, cool. Uh, quick intro. Um, so we are a leading data security-focused, uh, enterprise AI company. Uh, our focus is building trustworthy, uh, enterprise AI models, uh, with our partners for real-world business, uh, use cases.
- 0:45
Um, and, uh, we work with a lot of, like, strategics a-across, uh, various clouds, uh, and are, uh, have a bit of a Switzerland play when it comes to where we can, uh, ship and deploy, and I'll talk to that, uh, a little later.
- 0:59
Um, so we have a crack team of, uh, ML, uh, researchers and seasoned enterprise, uh, operators. Uh, Aidan, uh, who's our co-founder and CEO, was one of the, uh, authors on the, uh, seminal transformer paper.
- 1:12
Uh, we have, uh, Phil, who's our chief scientist. Uh, he was an NLP lead at DeepMind and a professor at Oxford. Uh, Nils, uh, who leads a lot of our retrieval efforts, was a BERT expert.
- 1:25
Uh, and then Patrick, uh, was, uh, the co-author on the RAG paper, which is a big, uh, focus for us at, uh, Cohere. So, um, fantastic team, uh, that's, uh, helping us build lots of amazing things.
- 1:36
Um, so here's a quick overview of our product line. Um, so we have, uh, two main arcs, so the generative side and then the advanced retrieval models. Um, on the generative side, uh, we have, uh, two flagship models, which is Command R and R+.
- 1:52
R is our, uh, workhorse, uh, super low-cost model, uh, that's great, great for most enterprise use cases. Uh, and then we have R+, which is, uh, your more powerful, larger model, uh, for more complex, like reasoning, tool use, RAG, uh, use cases.
- 2:09
Um, and on the retrieval side, uh, uh, most people have used some form of an embedding model, um, and I'll get into that a little later in the talk.
- 2:18
Uh, but, uh, the Reranker is something special, uh, that I haven't seen quite often in the market, uh, but we think, uh, it adds a lot of value, especially to your, uh, RAG pipelines.
- 2:28
Um, so we'll get into that too. Um, so when we build at Cohere, uh, so we have, uh, five guiding, uh, ethos as to how, uh, we want to go about things.
- 2:37
Um, the first is, uh, obviously everybody's testing their models in all sorts of academic benchmarks. Uh, but for us, what is really important is the performance on, uh, enterprise use cases.
- 2:48
Um, so we've worked, uh, with a lot of our partners to ensure that we have an eval suite, uh, that is highly customized to enterprise, uh, use cases across, like let's say health, HR, finance, uh, and that's, uh, a bit of our goalpost as we ship, uh, e-each of these like model versions and, uh, we constantly benchmark,
- 3:08
uh, on how we are performing at each of these industries, uh, and use cases that our customers care about. Um, and then when-- the next thing is ef-- all about efficiency and scalability, right?
- 3:18
We're, uh, not particularly chasing the race for having the largest model out there, but what we really care about is the practical use of these models, right? How, how do these models get used, and how cheap is it, uh, uh, and how easy is it for you to run it as a customer?
- 3:35
Um, the next big thing is obviously customization. Um, you know, as, as much as we'd like for all of these models to work out of the box, there's always a, a certain niche that, uh, customers want to customize this for.
- 3:49
Uh, and we offer a variety of things, uh, some of which are pretty intrusive. Uh, we've helped our customers, uh, with taking our base model, uh, and retraining that with their enterprise-specific data, uh, for domain adaptation.
- 4:03
We can do a full retraining of the model for you with your data. Uh, and then obviously the, uh, pretty typical last few layers, uh, retraining, which is self-serve on our platform.
- 4:14
Um, data prominence and privacy is another, uh, big focus, uh, for us. Uh, we've, uh, worked quite a bit to ensure that all of the data that we've, uh, collected for building our models, uh, is, uh, meets up to the enterprise, uh, standards, uh, and we offer indemnification for any IP claims, uh, that you might run into
- 4:33
as a customer. Uh, and then obviously we don't ever use any of your data to train our models. So, um, so that, that's a, uh, guarantee from us. Uh, and, uh, deployment flexibility.
- 4:44
Uh, as I mentioned, we're available on pretty much every major cloud provider. Uh, and then we also allow you to deploy, uh, on-prem or in your own VPC, um, uh, wherever, uh, your compute and your data is, that's where we'll meet you at.
- 4:58
Um, cool. Um, so just a quick look at, uh, uh, you know, your typical performance metrics. Uh, as I mentioned, uh, something that enterprises like repeatedly tell us is, uh, they care about like multilingual, they care about like RAG, um, and tool use for upcoming like agentic use cases.
- 5:16
Uh, so a, a lot of our focus has been, uh, in these areas. Uh, these are some benchmarks from HotpotQA, Bamboodle, uh, Berkeley Function Calling, uh, that our models, uh, are quite, uh, good at.
- 5:30
Uh, and, um, and another example of like how we try to innovate is, uh, we try to make sure that as we are building, uh, w- these stacks, we're incorporating all of the features that people care about out of the box, and they don't have to do extra work.
- 5:46
Citations on RAG is a great example of this. Uh, for most people, you have to do a lot of work to actually build this functionality with like other APIs, uh, but this really comes out of the box with like Cohere's, uh, models and APIs.
- 5:58
You don't have to do anything additional as a developer, uh, to build, uh, get citations and which is, uh, very important for any RAG-based like application. Um, uh, on the multilingual front, uh, we have one of the best performance when it comes to, uh, the FLORES multilingual, uh, evaluation.
- 6:19
Uh, and we also have a bit of a secret sauce with our, uh, tokenizer, uh, which, uh, helps keep costs really low, right? Uh, and, uh, that's again, uh, TCO is again a very big thing for enterprises.
- 6:30
Uh, and, uh, it allows our customers to take that same model and de- deploy across the globe on, uh, with their customers, uh, which is very important. Um, switching gears towards the embeddings models, uh, again, uh, given, uh, Nelson and his team, uh, have been innovators in this space for a while now, uh, and, uh, we've, uh,
- 6:53
done quite a bit of work to make our-- make sure our performance is great on, like, noisy data and at a super low cost, uh, uh, in this particular space.
- 7:02
Um, so we're, we're actually pretty excited about what our embeddings models, uh, can do, and this is al- almost always, like, one of the top things that, uh, our customers are, uh, excited about.
- 7:14
Um, but embeddings, uh, is a pretty complex space and not without its, uh, challenge. Uh, so we, we try to build this, like, fun demo where we took all of the archive papers, uh, and we asked it a question, uh, "When was the attention paper, um, built by, uh-- paper published by Aidan, uh, Gomez," who's our founder?
- 7:33
And we tried this across, like, a bunch of, like, embeddings model. So some common patterns that we see is, um, archive is a great example of, like, where you have different kinds of, like, data, right?
- 7:44
You have, like, the title, when was the paper published, the dates. Uh, you have the various authors. You have, like, the actual paper itself. And in many ways, this represents the kind of data you might see in enterprises, right?
- 7:55
Um, so when you, uh, actually build the embeddings for this, um, you get a fairly complex, like, vector space, uh, and your search queries might not actually, like, map neatly to this, right?
- 8:07
Uh, and this is sort of like where our reranker comes in. Uh, so what our reranker does is once you have, uh, all of these, uh, retrieved documents, uh, it hel- it's a cross encoder that helps you, uh, rerank the, uh, retrieved set and make sure that that's the one that you send into your context with the
- 8:25
generative model. Um, and, uh, here's the reranker in, uh, action. Um, so what this demo is showing you is, uh, we search, uh, for the transformer paper by Aidan.
- 8:38
So you have three different types of, like, search patterns over here: a lexical search, then an embeddings-based search, uh, and, uh, the Cohere rerank-based search, right? Uh, and the, uh, various forms of these, like, retrievals obviously give you the responses, but they are stack ranked at different places in the retrieval set, which means the overall accuracy of
- 8:58
your RAG system might be low. Uh, and rerank is what's helping you to make sure, uh, that isn't the case. Uh, this builds on top of, like, other things that people care about, like chunking strategies, uh, but making those more optimal.
- 9:12
Uh, another impact of this reranker is, again, total cost of operation, because most of the expense for your models is coming in from the input tokens, right? Uh, and if you were able to, like, narrow in, uh, to the right, uh, uh, context and do that quickly, uh, you could pass in very minimal amount of context to
- 9:32
your large language model, which drives on-- drives down your overall cost of, like, operation. Uh, and that, uh, is again, very important in the enterprise, uh, setting. Um, yeah, and when it comes to deployment options, like I said, we have our SaaS API, uh, that we can help, uh, you manage run your wo- uh, workloads.
- 9:52
Uh, but then we're also on all of the major cloud AI services, Sa-Sage, uh, SageMaker, Bedrock, um, OCI, uh, and private deployment across all of these cloud providers and pri-- um, also on-premise deployments if, if that's, uh, something you care about.
- 10:08
Um, security and privacy obviously is a pretty, uh, top of mind for us, so we make sure we're compliant with, uh, uh, the standards that our customers are often asking us, uh,
- 10:20
for. Um, and then the last bit is just enabling, like, developers. Um, so we, uh, obviously have a, a pretty tight integration with things like La-LangChain, LlamaIndex. Uh, but we also have an open source, like, toolkit, uh, that comes out of the box, uh, with, uh, various forms of, like, connectors, um, and that lets you, uh, ingest,
- 10:40
uh, data pretty easily into your systems and lets you have full control over the things you're, uh, building, um, and don't have to really, uh, look, uh, for a ton of different, like, options, uh, as you're developing your, uh, enterprise applications.
- 10:55
So, um, that's it from me, and I'd love to take any questions or chat. And we also have Sandra here, uh, from Cohere, so she'll, she'll be happy to help.
- 11:06
Thank, thank you very much.
- 11:07
Yeah.
- 11:08
May- maybe two questions. So on the first or second slide, you show your, uh, investors and the selected partners, right?
- 11:13
Yeah.
- 11:13
I saw Accenture, I saw McKinsey.
- 11:15
Yeah.
- 11:16
Can you explain a little bit, like, how that partnership, uh, work, right?
- 11:20
Yeah.
- 11:21
And then, and then the second question, like, you can skip if someone else, like-
- 11:23
Yeah
- 11:23
... has another one, is can, can you, like, maybe without disclosing, uh, customers and all, give us a, like a few samples of where you- ... clients pick to Cohere versus other solutions and, and kind of why?
- 11:37
So explain where you win, right, on the, on the enterprise world.
- 11:40
Yeah, absolutely. Uh, happy to chat about that. Um, so, uh, I think the typical challenge with all of these enter- um, enterprise, like, generative AI models is the last mile challenge, right?
- 11:50
Like, uh, there's, uh, so many different, like, arcs of, like, customization that's needed, uh, with the enterprises. Uh, and, uh, there's a lot of, like, traditional, uh, like, players who've been around for a while and have great relationships with the enterprises, uh, have a deep understanding of, like, the various, like, uh, business domains.
- 12:10
Uh, that's where, uh, the McKinsey and Accenture and all of these companies come in. Um, so they've been, uh, really, uh, helpful for us to co-develop, like, the product.
- 12:20
Like, make sure that we're able to effectively, uh, bridge that, like, last mile gap with them. Um, and yeah, that... hopefully that, that helps. Uh, and then onto your second question, I would say, um, uh, in terms of, like, winnability, I think, like, the main, uh, aspects that has been, uh, resonating a lot with our customers is
- 12:39
this control over the data and control over the compute, right? Uh, given we're available pretty much everywhere, uh, a lot of, like, the customers care about that private cloud deployment.
- 12:49
The ability to fine-tune in that, uh, private cloud environment, uh, which is pretty big, uh, for a lot of people, and making sure that their enterprise data does not leave their own ecosystem.
- 13:00
Um, and specific examples for that have been, uh, companies in, uh, let's say HR or healthcare or even, uh, folks who are trying to take their in-house, like, code and, uh, build a custom model that's, uh, working with, uh, their code base.
- 13:16
Uh, those are the styles of applications, uh, that, uh, we've seen a lot of like, uh, impact and success with.
- 13:21
Thank you.
- 13:22
Yeah.
- 13:27
So I'm curious, um, uh, for text classification, well, what is the kind latest best practice? I see on your website you have a classify endpoint, right, build a classifier.
- 13:38
Uh, is that still the recommendation?
- 13:40
Yeah. Uh, that's a, that's a great question. I, I... The, the way I like to think about, uh, these things is there's always the arc of like, uh, um...
- 13:49
A- and I think, like, Jerry in his earlier talk did this, right? Which was what phase of, like, development you're in, uh, if you're trying to, like, prototype or you're trying to productionize, and what's your scale of, like, a production setting, uh, which, uh, is important to consider for these things.
- 14:04
Uh, for, uh, if you're just trying to, like, get off the ground, like, quickly, I would say just using the generative model off the shelf is obviously always great.
- 14:13
It, it gets you off the ground really quickly. Uh, when you're trying to, like, productionize something, that's when I would start thinking, "Hey, do I need a bespoke model?
- 14:22
Uh, or, like, the, the general model is good enough," and what are sort of, like, the cost of operation, like, differences and, like, the cost of, like, maintenance dif- differences and also the scale, right?
- 14:32
Like, I mean, if, if you're going to try to do something that's, uh, you know, tens of thousands, like TPS, uh, then having a, a purpose-built, like, model for that is the route I'd go, uh, versus, you know, uh, a more heavy general model, uh, which might serve other, other needs.
- 14:48
Um, so yeah, both of those are good options depending on what you're trying to accomplish.
- 14:53
Just a quick-
- 14:53
Yeah. Yeah
- 14:53
... uh, quick follow-up. If we actually train a classifier with you, is that also a transformer model, just more specialized?
- 15:00
Yeah, exactly.
- 15:01
Okay.
- 15:01
Yeah.
- 15:02
Got it. Thank you.
- 15:05
Yeah.
- 15:06
No, I think we're done. Oh, no, we've got one more question. Here we go.
- 15:10
Just a quick question. When we were looking at Cohere, you know, probably early last or middle of last year-
- 15:18
Mm-hmm
- 15:19
... um, one of the challenges with the models that we found were the input context size limit-
- 15:24
Yeah
- 15:24
... were quite small. H- how has that evolved a- as you guys have sort of created the next, you know, sets of models on your side?
- 15:31
Yeah, that's a great question. So our latest generation models are, uh, fairly competitive, 128K, uh, context input, uh, windows. Uh, and we're constantly looking to, uh, figure out how to, like, up them.
- 15:42
Um, so context window, I would say, should be not a problem for, like, most applications, uh, at the moment.
- 15:50
Okay.
- 15:50
Yeah.
- 15:53
Cool. Well, we're actually slightly early, but yeah, I'd like to thank you for giving the talk.
- 15:58
Yeah.
- 15:58
Thank you, Vivek.
- 15:58
Absolutely. Thank you so much. Yeah.
- 16:00
Thank you.
- 16:00
Yeah.
- 16:01
Thank you. [audience cheering] [upbeat music]