AI Engineer Summit 2025
Anthropic for VPs of AI
About this talk
Anthropic’s Alexander Bricken and Joe Bayley outline practical enterprise AI adoption, introducing Claude 3.5 Sonnet and coding evaluations before explaining interpretability through semantic feature activations and the Golden Gate Claude steering demonstration. They discuss Model Context Protocol, Intercom’s Fin customer-service deployment, and implementation practices centered on LLMOps, use-case-specific metrics, rigorous evaluation, and solving concrete business problems.
Chapters
- 0:00Introductions and enterprise AI implementation agenda
- 1:52Claude 3.5 Sonnet, SWE-bench, and interpretability
- 4:08Feature activations, Golden Gate Claude, and customer problems
- 9:38Model Context Protocol, deployment support, and Intercom Fin
- 13:08Evaluation practices and closing remarks
Talk transcript
- 0:00
[on-hold music] I'm Alexander Bricken.
- 0:18
I'm on the Applied AI team at Anthropic, so I work very closely with customers to do technical implementation work, and I also bring that advice back to product research and model research.
- 0:29
Um, I'm gonna pass it over to Joe.
- 0:31
Hey, everyone. It's great to be here. My name is Joe Bayley. I work on the go-to-market team at Anthropic. I joined Anthropic, uh, over a year ago now, so I've seen our models evolve from, uh, two point one to today's capabilities.
- 0:42
And I think day to day what's really exciting is we're working with AI leaders who are solving real business problems, um, that just seemed impossible a year ago. So really excited about how quickly everything is moving.
- 0:55
Okay. For today, we will do, um, a quick overview, uh, you know, who we are, o-our mission, and then we'll focus a lot on implementing AI and best practices and common mistakes.
- 1:08
Uh, Alex and I actually didn't just take this from our own experience, but we, um, talked to a number of our colleagues, so this is all based on hundreds and hundreds of customer interactions.
- 1:18
Um, so we hope there's some actionable insights to take out of this.
- 1:23
Awesome. So what is Anthropic? So we are an AI safety and research company building the world's best and safest, uh, large language models. We were founded a few years ago by some of the leading experts in AI.
- 1:37
And since our inception, we've not only released, uh, multiple iterations of our frontier models, we've done so while being at the bleeding edge of safety techniques, of research and policy.
- 1:48
I'm gonna pass it over to Alex to talk a little about our marquee model.
- 1:52
Awesome. And so some of you are probably familiar, but the most recent model we launched was Sonnet 3.5 New, uh, in late October of last year. Um, you might be familiar with it because if you're a developer, uh, Sonnet is actually one of the leading models in the code space.
- 2:08
So if you're familiar with evaluations like SWE-bench, which is an, an agentic coding eval, uh, Sonnet is still at the top of the leaderboard for that. Um, I won't go too much into the details on the eval side, so let's keep moving.
- 2:21
Um, so yeah, in addition to what Joe mentioned, we have a lot of different research directions that we're focused on. Um, and these are really distributed but have overlap between, you know, model capabilities, product research, and AI safety.
- 2:35
The one that differentiates us, I would say, is the interpretability, and this realistically is reverse engineering the models and trying to figure out actually how they're thinking and maybe why they're thinking, and then a-an additional capability in terms of steering them in the right direction, depending on a use case.
- 2:50
So let's dive into that a little bit more. We're still very early in interpretability research, it's worth mentioning. As you can see, there's kind of like a longer timeline, and we're really only at the, the first half of that, maybe even the first twenty-five percent.
- 3:02
Um, but we're, we're really approaching it in these stages that build upon each other. So these in-include things like understanding, so grasping AI decision-making. Detection, so actually being able to understand specific behaviors and put labels on those.
- 3:16
Steering, so influencing the AI input in some, some way, shape, or form, and I'll get to an example of that in a second. And then finally, explainability, and that's really where you unlock business value associated with interpretability methods.
- 3:28
And so while we see interpretability in the long term providing, you know, a lot of significant improvements in AI safety, reliability, and usability, specifically, our interpretability team uses methods to understand feature activations at the model level, and then has published research on these, uh, in-- towards model semanticity and scaling model semanticity, which are two papers I highly
- 3:50
recommend. Um, and then as the technology improves into kinda detection landscapes, for example, you can imagine having a much better grasp at, uh, at the actual thinking and behavior of the model or even discovering sleeper agents for safety reasons that might be buried deep within, uh, model capabilities.
- 4:08
So a good example of that is imagining you ask the model, "What were the scores of the NBA matches today?" Right? And let's say it knows the answer, and it says how Steph Curry, you know, scored thirty points.
- 4:20
This would lead to a feature activation of, for example, feature number [REDACTED:generic_id], famous NBA players. Realistically, that's a group of neurons activating in a recognizable pattern that we've identified across all mentions of famous basketball players when a model's answering a question, not just Steph Curry.
- 4:40
Um, and you also might have heard of Golden Gate Claude. That was an example of us steering the model, uh, basically amping up the activation, uh, in the Golden Gate direction.
- 4:49
And thus, whenever you'd ask a question like, "What should I paint my bedroom?" Claude would respond, "Oh, you should paint it red like the Golden Gate Bridge, and maybe it should have some like, you know, pillars in it or something."
- 5:00
I'm gonna pass it over to Joe to talk a little bit about some of the customers we work with.
- 5:03
Yeah. So I'm gonna frame this in two ways. One is, uh, sort of early on discussions, and the other would be just examples of customers that are doing really cool things.
- 5:12
So in conversations, there's obviously a lot of noise and buzz and everything, and that's fantastic, but we often encourage our customers, uh, to sort of get back to the basics and how can they use AI to solve, uh, the core problem that your product is trying to solve.
- 5:28
We also get to work with a ton of, uh, sort of AI native, like, AI startups, and this is how they're thinking about their product. And I think you wanna move beyond, uh, chatbots and summarization.
- 5:39
These can be great options, but I'd be thinking more like, where do you wanna place bigger bets? And to give an example, um, if you just click one more time.
- 5:47
Fancy slide. Uh, imagine you're an onboarding and upskilling platform. The problem that you solve for customers is you help them get-- you ramp really quickly, and then you help them get to the next phase o-of their career by equipping them with skills.
- 6:01
So for instance, it might be public speaking you wanna get good at, or you might wanna become a manager. And so it would be easy to say, "Okay, let's summarize course content."
- 6:09
Or let's, um, [lip smack] uh, let's, uh, have a Q&A chatbot that answers questions along the way, and they could be helpful. But I'd actually think about it differently. So what about if you could hyper-personalize, uh, course content based on each indiv- individual employee's context?
- 6:24
Or if someone is, like, breezing through all the course content, could you adapt it dynamically to make it more challenging, uh, so they're actually getting more value out of it?
- 6:32
And then the last one that I particularly like would be, what if you could, uh, uh, dynamically update, uh, course material based on people learning about the customer? So if someone was a visual learner, great, let's make visual content for them and do-- having the AI, having the ML, uh, sorry, the, the large language model just do
- 6:50
that automatically. And you have to think, does that solve the problem more than summarization or a, or a Q&A chatbot? Um, so really good food for thought. And to sort of talk about some of the customers, uh, that we see achieving really industry-leading results, uh, by combining, uh, their own domain expertise and our, um, our model.
- 7:11
So I won't read off each, but just a couple of call-outs. One is, uh, AI impacting different industries. We have, uh, taxes, we have legal, we have, uh, project management.
- 7:21
They're using AI to, uh, drastically enhance their customer experience. They make it more, uh, like, easier to use. They make it more trustworthy. Um, and so it's really improving the experience versus just being like a nice to have.
- 7:34
And then they're achieving a real h-- a real high quality, um, of output, right? You can't be giving-- you can't be hallucinating when you're doing your taxes. Uh, it could just, you know, that could lead to all sorts of things.
- 7:44
So we're thrilled that they're seeing these, these sort of like business-critical workflows powered by AI, driving really positive outcomes for them and also their customers.
- 7:55
Awesome.
- 7:56
I can do this one. Yeah. So, [chuckles] um, getting started. I just quickly, uh, there's two, two key points here. So if you go on the next slide, um, what are our products?
- 8:06
We have our API, we have Claude for Work. Our API, uh, is for businesses that wanna embed AI in their product and services, and then Claude for Work empowers your entire organization to take advantage of AI in their day-to-day work.
- 8:19
We also have, uh, next one. We also have a partnership with AWS and GCP, and you can kinda get the best of both worlds here. You can access our frontier models on Bedrock or on Vertex.
- 8:32
You can deploy these applications in your existing environment, um, and you, but you-- and so you don't have to manage any new infrastructure. So it really sort of breaks down like any barriers to entry.
- 8:43
So you're getting the best of both worlds here. We talk a little bit about support throughout this talk. It-- to us, it doesn't matter if you're accessing us through a third party or a first party, so I just wanted to call that out.
- 8:56
Awesome. So now that we've talked a little bit about some of the customers, how do we actually set customers up for success when working with them at Anthropic? So [clears throat] just to preface on kind of what my team does, as I mentioned, it's at this intersection of product research, uh, customer-facing interaction, and then also just actual research within
- 9:13
the org. Um, and we support, support the technical aspects of the use cases, so helping to design architectures, evals, tweak Claude prompts to get the best out of our models, et cetera.
- 9:23
And then we also bring whatever we see back into Anthropic, and we try to build some of the best products we can for our customers. So some examples of projects we've worked on or things that we've published include the Building effective, uh, research, uh, paper that, uh, my colleague Barry published.
- 9:38
He's gonna be speaking tomorrow. And then as well as that, we've launched Model Context Protocol, which is a open source protocol for language models to interact with data sources.
- 9:47
And, um, Mahesh is gonna be leading a workshop on that, uh, on Saturday, I believe.
- 9:54
So Anthropic as a whole, we try to effectively support our customers, but where we really start to embed, at least my team in particular, is, um, we work closely with customers that are using Claude a lot, and they're facing really niche challenges in specific use case domains, and they need support from our team to try to apply
- 10:10
some of the newest, kind of latest and greatest research or get the most out of the models from a prompting standpoint, et cetera. And so this approach is pretty additive.
- 10:17
We often kick off a pr-- a sprint once the customer is facing those tricky challenges. Uh, and that could be LL- LLMOps, architectures, or evals. We help to define certain metrics that they, they deem to be important when they're evaluating the model against, uh, the use case.
- 10:32
And then finally, um, we help them deploy that kind of iterative loop, the result of that into an A- AB test environment, um, and then hopefully into production. And so a part of that is the importance of evals, and I'll get onto that in a second.
- 10:45
Um, but I'm gonna pass it over first to Joe to talk about some stuff that we did for Intercom.
- 10:52
Yeah. So sort of the-- I think this is a good segue on what Alex was describing. So for those of you don't-- who don't know, Intercom is an AI customer, uh, service platform.
- 11:01
They have an AI agent called Fin. By many measures, it's the best in the market, and it's a pretty competitive market. So they had their product, uh, out for, I think, about a year or so.
- 11:12
And when we spoke to them, they shared wh- where they wanted to go, where they saw the future as-- of like customer support and agents. And based on some of the capabilities of our model, we felt that we could have a pretty good impact on these metrics.
- 11:25
And so what we started with was the Applied AI, uh, lead met with their data science team, and we ran a quick two-week sprint. We took their hardest prompt for Fin, and we compared it, uh, against a prompt that we helped them sort of figure out with, uh, with Claude.
- 11:40
And they saw really good results after the first two weeks. So much so we went on this sort of sprint of about two months where we were basically, um, fine-tuning and optimi- optimizing all of their, uh, prompts to get the best performance out of Claude.
- 11:55
At the end of this, they're able to look at all their benchmarks and see that Anthropic was outperforming the current LLM. It's also worth noting that they do a resolution-based pricing model, so there's an incentive for everyone for the model to be really helpful and help customers solve problems and not be like a deflection machine where it's
- 12:11
like, you know, we've probably all experienced them before. And so, uh, at the end of this two months, they decided to move forward with Anthropic. They launched it. You can read about it.
- 12:19
It's called Fin 2. And I think just some of the metrics are really, like, mind-blowing. Like, can solve up to eighty-six percent of customer support volume, fifty-one percent out of the box.
- 12:28
Our support team
- 12:29
Thought about lots of different options, and they actually adopted, uh, Fin as well, and they saw very similar resolution rates. But also making it more human, so they-- I think with our model, we can-- there's a much more of a human element to it, so they could do, like, uh, adjustment of the tone, uh, answer length, and
- 12:45
then it was also really good at doing policy awareness, so like refund policy, for instance. So unlocking some new capabilities. And we're thrilled to be partnering with them as they sort of, I think, march forward as a leader in, in this space.
- 12:57
Yeah. On, on a kinda separate note, one of the things I've seen recently is, uh, Claude on Twitter acting as some sort of therapist for a lot of people, and I always find that an entertaining example- [laughs] ...of, like, its character being expressed.
- 13:08
Yeah.
- 13:08
Um, cool. So let's get onto some best practices and mistakes that we see in the field, uh, on the go-to-market team. So firstly, testing and evaluation. I'm sure this-- those two words have been mentioned a lot today and probably tomorrow too.
- 13:22
Um, there are some typical common mis- common mistakes that we see, uh, customers struggling with. So the first one is they build a really robust workflow. They've spent, like, a bunch of time building some architecture out, and then they're like, "Okay, now we need to evaluate it.
- 13:34
Like, let's build some evals." That's not really how it should work in practice because your evals are actually the thing that directs you towards a perfect outcome, right? You can't build a whole workflow without evals probably from the get-go or very shortly after.
- 13:48
And so, you know, sometimes customers, as a result of struggling with data problems, might not be able to design their evals. You could use Claude to clean that up, do data reconciliation.
- 13:58
Um, or they're just, you know, trusting the vibes too much. Maybe they run a couple queries, they're like, "Hey, it looks good," right? Are they really testing that on a representative sample though?
- 14:07
Like, do you have enough s-samples to say that the thing that you're looking at is statistically significant? And like, or are you gonna, you know, run a hundred things when it actually goes into broad and then there's gonna be, like, loads of outliers, uh, because you didn't actually predict correctly what, uh, the customer is gonna ask of
- 14:22
the model, for example. So I challenge you to think about your use cases as this sort of latent space, right? Let's take this kinda chart here on the left-hand side of the slide, right?
- 14:34
As you explore the latent space with different functions that you can apply to the model, let's say prompt engineering, prompt caching, stuff like that, you're, you're basically moving your kind of position in that latent space around between l- attractor states, you know.
- 14:50
And eventually you wanna find an optimized point, but you don't really know where that is, right? Like, if you're changing an instruction, you don't know how the attention mechanism of the transformer is gonna eventually result in some different outcome that might not be performant.
- 15:03
And so the only way you can truly know that is empirically, and that's through evaluations. And so I think that's why evaluations are so, so important, and a lot of people just don't understand that soon enough.
- 15:13
In many ways, I actually tell customers, "Evals are your intellectual-- intellectual property." Like, if you wanna be competitive in a space, you need to be able to outcompete people by navigating that la- latent space and finding the attractor sta-state faster than anyone else.
- 15:29
Um, and so h-- part of, you know, how you do that is, well, firstly, setting up some sort of telemetry, uh, to back test. Ideally, that architecture's set up in advance, but, you know, you should invest in it.
- 15:41
Um, designing representative test cases. So let's say you're working on that customer support agent eval. You know, you might have a kid come on your website, let's say you're building it for, and might ask some crazy question like, "How do I, you know, kill a zombie in Minecraft?"
- 15:56
Like, totally unrelated to your product. That's, you know, still probable, and so you should probably include silly examples like that in your eval set to make sure that your model's actually approaching the response in an appropriate way or rerouting the question, et cetera.
- 16:10
Cool. Moving on to the next one. Um, identifying metrics. So a lot of the time, you know, there's this intelligence, cost, latency triangle of trade-offs that people are trying to move in between.
- 16:22
And most or-organizations can optimize for one or two of those things, but it's very difficult to meet three, at least right now. But realistically, that balance should be defined in advance, and you should know that for your specific use case, you're going to make a trade-off between those things.
- 16:39
So let's say a customer support use case again. You care about your customer getting a response within ten seconds, right? If more than ten seconds, I think there's been research done on this, uh, the customer's likely gonna just log off the page, and then they won't get the response, and then they'll probably complain about your product to
- 16:53
their friends, right? Whereas if you're looking at a financial research analyst agent, you probably don't care that it works for ten minutes to come up with the actual response to your question because the decision being made after that is very important.
- 17:05
It's an allocation of capital, for example. And so the stakes and time sensity-- sensitivity of the decision should really drive your optimization choices and, you know, maybe more instruction sets lead to longer latency but higher performance, et cetera.
- 17:20
The other thing is UX could be important, right? So again, on that customer support agent, because we spoke about Intercom, you could have different ways of circumventing that ten second to fifteen seconds.
- 17:31
Specifically, you could add like a little thinking box that bounces around. You could send the customer to another webpage in the meantime, have them read something, right? Like, there's loads of ways to distract and kind of push on those boundaries, but you still need to know what that important indicator is, and you need to optimize accordingly.
- 17:49
Finally, fine-tuning. So a lot of people, you know, I go into these calls and they're like, "Oh, we wanna do fine-tuning." I'm like, "Oh, here we go again." [laughs] Um, fine-tuning's not a silver bullet.
- 18:00
So it comes at a cost, and most people aren't aware of that cost. Um, the cost is generally you're doing brain surgery on the model, and thus there can be kind of limitations to its reasoning in other fields outside of the thing you're fine-tuning towards.
- 18:14
Um, so my encouragement is try other approaches first, right? Most people, they don't even have their eval set when they're trying to do fine-tuning, right? They need to have a cri-- clear success criteria in advance.
- 18:25
And it's like, only if we can't get that in our specific intelligence domain, do we then do fine-tuning. Don't try to boil the ocean in advance. The difference in fine-tuned capabilities and the wide variance of which, you know, failure versus success looks like in fine-tuning land means that you should be able to justify the cost of fine-tuning
- 18:46
and the effort of doing it, right, getting a team, fine-tuning, working with us, for example, uh, you should be able to justify that difference. And so in, in terms of best practices, you, you know, don't want let, to let fine-tuning slow you down, right?
- 19:00
You don't wanna say, "Oh, I'm only going to convert this language model use case if we can actually finally fine-tune our model." It's like, no, no, pursue it and then realize that you need to do fine-tuning, and then you can just sub in the fine-tuned model.
- 19:12
And then explore other methods first, and there are loads of different methods that, you know, are Anthropic as well as other companies are working on these days, and I just wanted to, like, flash up a few of them as we wrap things up here.
- 19:25
And so I'm not gonna go through all of these, but alongside just base prompt engineering, which granted is very important, there are loads of different features or architectures that will change the success of your use case drastically.
- 19:38
So for example, you might not need to sacrifice on intelligence of your model if in order to speed it up by, like, removing instructions if you can just leverage prompt caching and have a 90% recut- reduction in cost and a 50% increase in speed, right?
- 19:53
Or contextual retrieval will drastically improve the performance of your retrieval, retrieval mechanisms, and thus you feed the information to the model more effectively, and thus it has less of a time processing all the instruction set that you've given it.
- 20:06
So there are quite a few things that you can apply here, and some of them are even out of the box, like citations. And then there are also architectural decisions like agentic architectures.
- 20:15
You know, Barry, my colleague who's speaking tomorrow, will have a lot to say on that. Um, but that pretty much does it. Um, thank you so much for, uh, for your time.
- 20:25
We'll be in the theater level lounge after this chat for follow-up questions. Um, anything else from you, Joe?
- 20:32
No. Thank you so much.
- 20:34
Cool. Cheers. [outro music]