← All AI Engineer talks

AI Engineer World's Fair 2025

Building Agents at Cloud Scale — Antje Barth, AWS

Antje Barth· AWS19:00

About this talk

Antje Barth explains how AWS approaches production-scale AI agents, using Alexa Plus and its specialized expert systems to illustrate coordination at scale. She demonstrates Amazon Q Developer CLI integrating MCP servers and grounding responses in AWS documentation, highlights the open-source awslabs/mcp repository, and discusses agent-tool authorization, a D&D-themed demonstration, and related Strands sessions.

Chapters

  1. 0:00Building AI agents at cloud scale
  2. 1:40Alexa Plus and specialized agent systems
  3. 4:43Amazon Q Developer and MCP documentation demo
  4. 12:27Open-source AWS MCP servers
  5. 13:41D&D demonstration and agent authorization
  6. 18:28Strands sessions and closing

Talk transcript

  1. 0:00

    [upbeat music] [laughs] [audience applauding] Hi, everyone.

  2. 0:24

    I'm thrilled to be back on stage here again at the AI Engineer World's Fair, and it's amazing to see this community grow. So today, I'm gonna speak about how we can build agents at cloud scale.

  3. 0:39

    Now, at Amazon and AWS, we truly believe that virtually every customer experience we know of will be reinvented with AI. And not just the existing experiences, but there will also be brand-new experiences we are now able to build with the help of AI agents.

  4. 1:01

    And we're not just theorizing about this, right? We're all here together to actually build the future.

  5. 1:11

    Now, I wanna start just with a little bit of what that means internally across Amazon as a business.

  6. 1:19

    At Amazon, we have over one thousand generative AI applications that are either built or in development, transforming everything from how we forecast inventory, to how we optimize delivery routes, to how customers shop, and how they interact with their homes.

  7. 1:40

    And one of the most ambitious deployments of AI agents is the complete reimagining of Alexa. And I know many of us have been waiting for this for a long time. [laughs]

  8. 1:53

    So what you're about to see here represents the largest integration of services, agentic capabilities, and LLMs that we know of anywhere. So let's have a brief look. [upbeat music]

  9. 2:10

    Wow, wow. Look at my style. I know you ain't seen it like this in a while.

  10. 2:15

    Oh, hey there. [upbeat music] So we can just, like, talk now? I'm all ears, figuratively speaking. Cool. Do you know how to manage my kids' schedules? I noticed a birthday party conflicts with picking up Grandma at the airport.

  11. 2:28

    Want me to book her a ride? Billie Eilish is in town soon. No way. I can share when tickets are available in your city. Yes, please. Got any spring break ideas?

  12. 2:38

    Somewhere not too far.

  13. 2:39

    Only if there's a beach.

  14. 2:40

    And nice weather.

  15. 2:41

    Santa Barbara is great for everyone. I found a restaurant downtown I think you'd like. What is Santa Barbara known for? It has great upscale shops and oceanfront dining. Can you go whale watching?

  16. 2:51

    Absolutely. Want me to book a catamaran tour?

  17. 2:54

    Wow.

  18. 2:55

    What's the next step? Remove the nut holding the cartridge.

  19. 2:58

    Should I get bangs?

  20. 2:59

    You might only love them for a little while.

  21. 3:01

    You're probably right.

  22. 3:02

    Make a slideshow of baby Tina.

  23. 3:03

    Mom.

  24. 3:05

    What part am I looking for again? Two-inch washers. Your Uber is two minutes away. For real? [record scratching]

  25. 3:11

    Wait, did someone let the dog out today? I checked the cameras, and yes, in fact, Mozart was just out.

  26. 3:16

    Wow, wow. Look at my style. I know you ain't seen it like this in a while. Wow. [audience applauding] [laughs]

  27. 3:26

    I love sharing this video because it shows really the power of agents at scale. And just to have a quick look what that means in terms of numbers,

  28. 3:38

    we have over six hundred million Alexa devices now out in the world, and with the help of the latest advancements in AI, we were able to really reimagine this experience.

  29. 3:52

    Alexa Plus works through hundreds of specialized expert systems. That's what the Alexa team calls groups of capabilities, APIs, and instructions to accomplish a specific task for you.

  30. 4:07

    And all of these experts also orchestrate across tens of thousands of partner services and devices to get the things done, which you just seen a glimpse of this here in this video.

  31. 4:21

    And we truly believe that the future will be full of those specialized agents, each with their own unique capabilities and working together seamlessly with other AI agents.

  32. 4:35

    Now, this example shows what's possible at this massive scale, but how do we get there?

  33. 4:43

    How do we operate at this scale? Or said differently, how do we move from web services that we've built for many years now into developing those agentic services? And luckily, many of the underlying principles remain the same, whether you're building for millions of devices, whether you're reimagining and integrating AI experiences into your enterprise

  34. 5:08

    applications, or you're a startup and you're really just looking to kind of scale your idea to the next level.

  35. 5:15

    Now, another example I wanna show you is an agentic service that we built at AWS.

  36. 5:23

    You might have heard about Amazon Q Developer, which is our code assistant that helps you really kind of across the software development life cycle. And just a few month ago, we released an Q Developer agent for your CLI.

  37. 5:39

    So it brings the agentic chat experience into the terminal. It helps you to debug issues. You can ask it natural questions. It can read and write files and really kind of help to make your day-to-day in the terminal more productive.

  38. 5:52

    So let's have a quick look how this looks.

  39. 5:56

    Here is Amazon Q in the CLI, and I'll just ask a good question here. In this case, "Hey, what do you know about Amazon Bedrock?" CLI is integrated with MCP.

  40. 6:06

    So what it does, it actually figures out there is a tool. Our AWS documentation team has released an MCP server. It's connecting to it. You see the toolie is happening, and it's asking for permission, so I give it the permissions, and then it comes back with a response that is grounded in the official AWS documentation.

  41. 6:27

    Now, I don't wanna talk much more about Q, but I do wanna ask for you just to quickly think about how long did it take for the AWS internal teams to build and ship this agentic service.

  42. 6:42

    And let's just do it with a quick raise of hands. Who think it took two months to develop and ship this?

  43. 6:49

    It's a few hands. Who thinks three weeks?

  44. 6:52

    All right, it's a bunch of more hands. Who do you think it took half a year?

  45. 6:58

    Almost none. Wow, you folks are great. We built and shipped this within three weeks, and to me, this is just almost insane, right? Like, the speed, and we heard it earlier, like, the mode of, of AI, um, one of the keynote speakers called it out, is execution, right?

  46. 7:16

    And I think three weeks is super impressive. Now,

  47. 7:21

    how do we enable teams, and not just internally at AWS, but in general, to build and ship production-ready AI agents this quickly? What we did internally, our teams, we needed to fundamentally rethink how to build agents.

  48. 7:38

    And what we did is we developed a model-driven approach that really kind of taps into the power of LLMs these days and models that are so much more capable in deciding, planning, reasoning, taking actions, and let the developers focus on what their agent should do rather than telling it exactly how to do it.

  49. 8:00

    And the great news is we made it available for all of you to use as well. So just a few weeks ago, we released Strands Agents. It's an open source Python SDK, which you can check out and start building and running AI agents in just a few lines of code.

  50. 8:21

    So let me show you quickly how this looks like.

  51. 8:24

    And before I go in here, just a fun fact if you wonder, why did they call it Strands Agents? Well, this is what happens if you let AI pick its own name.

  52. 8:36

    All right. So the reasoning behind, because again, the AI agent is, is capable of reasoning, it came up with, like, think about the two strands of DNA. And just like the two strands of DNA, Strands Agents connects the two core pieces of an agent together, the model and the tools.

  53. 8:57

    And it helps you building agents. It simplifies it by y- really relying on those state-of-the-art models to reason, to plan, and take action. You can simply start with defining a prompt and your tools in code and then test it out locally and then once you're ready, deploy it, for example, in the cloud.

  54. 9:19

    And this is how simple it is. Again, just a couple of lines. Should look pretty familiar. You install Strands Agents, you import it, and then it comes with pre-built tools, which I talk about a little bit more in detail.

  55. 9:32

    And basically, you just add the tools to your agent, and then you can start asking questions or building more complex workflows with it.

  56. 9:40

    Now, by default, Strands Agents integrates with Amazon Bedrock as the model provider, so you can check the model config here using Claude three point seven Sonnet. But of course, it's not just limited to AWS.

  57. 9:56

    You can use Strands Agents across multiple providers. For example, we have integrations with Ollama, so you can start developing locally, testing it out. We have integrations, Entropic added integrations, Meta added integrations to the Llama API.

  58. 10:11

    You can use OpenAI models and any other providers available through the integration with LightLLM. And of course, you can also develop your own custom model provider.

  59. 10:23

    Now, quickly on the tools. As I said, Strands Agents comes with over twenty pre-built tools. So anything from simple tasks like, hey, I just wanna do some file manipulation, some API calls, obviously integrate with AWS services, but then also more complex use cases, and I just wanna call out a, a couple of them.

  60. 10:44

    So there's a whole group of integrated tools from memory and RAG. One tool specifically called Retrieve, which lets you do semantic search over a knowledge base. And just to show you the power of this, we have an internal agent at AWS that manages over six thousand tools.

  61. 11:05

    Now, six thousand is a hard number of tools to put into a single context window and give, um, one model to decide. So what we did is we put the descriptions of those tools in a knowledge base and used the Retrieve tool here so the agent can find the most relevant tools for the task at hand and

  62. 11:24

    only pull those tools back into the model context for the model to decide which one to take. So that's just one use case how we're leveraging that. Also, there is support for multimodality across images, video, and a- audio with Strands.

  63. 11:41

    There is a tool to kind of prompt for more thinking and deep reasoning, and it also comes with pre-built tools to implement multi-agent workflows, whether it's graph-based workflows or a swarm of sub-agents working together.

  64. 12:00

    Now, you cannot talk about tools without mentioning MCP, right? [laughs]

  65. 12:05

    So obviously we integrated MCP here natively within Strands. So you can just use this also to connect to thousands of available MCP servers and make them available as tools for your agent.

  66. 12:18

    Support for A2A is also coming soon. But let's start and talk a little bit about MCP first.

  67. 12:27

    If you're building on AWS already, make sure to bookmark this GitHub repo. It's aws-labs/mcp, and here you can find a very long list, much longer than you would see here on this slide, of a growing number of the MCP server implementation, specifically if you're working and building on AWS.

  68. 12:48

    Now, one of the challenges stems from the fact that once we all started building MCP servers, what we had was standard IO, right? So it started out to help locally connect your systems, your clients to respective tools.

  69. 13:05

    And here's just a quick example, which is important for a demo I'll show in a little bit [laughs]. This is just a standard IO implementation of an MCP server.

  70. 13:14

    Should look familiar to most of you working with MCP using the Python SDK, using FastMCP. All I'm doing here is set up my server and using the decorator to define a tool.

  71. 13:25

    In this case, my tool is to roll a dice. And you might see in the code here it has an input to define the number of sides. And I had to put a picture here because I have to admit, um, I just learned this myself.

  72. 13:41

    Do we have D&D fans in the room? Woo-hoo. All right, a few of them. So you all know what I'm talking about [laughs]. For the rest of us, I just learned, um, there are dices, and I have one here, not sure if the camera can catch this.

  73. 13:56

    Um, it's just one of them here on the slide. A dice that has, for example, this one has 20 sides. Something very normal in the D&D world to start, I think, your game.

  74. 14:07

    Um, don't ask me questions about D&D. My colleague Mike Chambers, who is either here or in the expo right now, he built the demo, so kudos to him, and he can answer all of the D&D questions [laughs].

  75. 14:17

    All right, just keep that in mind. Um, I'll come back to this in just a second.

  76. 14:22

    Now, what we wanna do here is to decouple and kind of connect to remote MCP servers because the topic is to scale, right? And the way to do this is, in the AWS world, as easy as just deploying it as a Lambda function.

  77. 14:39

    So we can do this now with Streamable HTTP, and the same concepts apply. You put your Lambda functions as you would have before behind an MCP gateway and then connect.

  78. 14:50

    And because we care about security and authorization, in the quick demo I'm gonna show you, I'm using an authorizer. Um, you can also plug in a Cognito framework for this part, and I'm also gonna store session data in a DynamoDB table.

  79. 15:04

    So let's roll this quick demo here. So what you see here is an MCP Lambda handler that we developed, it's available on the GitHub repo, which makes it really easy to kind of set up your MCP server in Lambda.

  80. 15:16

    Here's a very simple "Hello, world" example. The tool is just, um, again, defined with a tool decorator in here, and then in the Lambda handler function you can reference, um, the input here, the invoke function, and pass it to that MCP server.

  81. 15:30

    Now, if we're looking at the server implementation, and here we're doing a little bit more, you can see how we're adding session table support, which is a DynamoDB table.

  82. 15:39

    We're defining the tool. This is the rolling dice tool that I just pointed out, but this time it's hosted as a Lambda function. You can write all the code you wanna have there as well.

  83. 15:50

    And then at the very end, it's the same single line that basically when you call the Lambda function passes this on to the MCP server.

  84. 16:00

    Let's deploy this. And again, we're using the existing tools to deploy Lambda functions as we have before. So this one is using AWS Sam to just deploy that to the cloud, and then we will receive the API gateway URL as well.

  85. 16:14

    Now, from the client side here, I'm using Strands Agents, as you can see, and then I am using the MCP integration. I'm passing here my API gateway URL to connect.

  86. 16:27

    For authoriz-authorization, I have a bearer token. Again, this is a simple concept demo, but you can build more robust integrations here as well. I'm calling the list tool, and then I'm passing those tools to my agent, as we've seen before.

  87. 16:41

    This time it's the MCP available tools. And then if we run this here, we can quickly see

  88. 16:49

    this in action and basically gonna ask it here to roll a dice.

  89. 16:55

    And we're asking it to roll a D20, so again, 20 sides. And it's coming back. What did we roll? You can see the tool here is kicking in here.

  90. 17:04

    We rolled a seven. Great [laughs]. So this is just really a quick example. The good news is, once you're in the AWS world and you're working on Lambda, everything you can build with Lambda, you can integrate there.

  91. 17:16

    So basically you have access again to all of the great features, capabilities, applications you might have already built on AWS. Now, the next step here is how do we make agents talk to each other, right?

  92. 17:27

    That's kind of the, the next frontier. And we're super excited about the, all the open protocols that are emerging right now. With MCP, for example, we joined the steering committee.

  93. 17:38

    We're active part of the community contributing code and helping to further evolve MCP. If you wanna learn more about this, here is the QR code. We have a whole blog series started on our open source blog.

  94. 17:50

    Feel free to check that out as we continue to help evolve those protocols.

  95. 17:56

    Now, what's next? We all are aware that this is just the beginning, right? There will be so much more coming. And if you had a chance to check out my colleague Danielle's talk yesterday on useful general intelligence, I just wanna quote her a little bit.

  96. 18:11

    She said, "The atomic unit of all digital interactions will be an agent call." So we can imagine a future here where you might just have a personal agent like shown like this connecting to an agent store and really kind of having agents together accomplishing tasks for you.

  97. 18:28

    And some of you here in the room might already be building this, right?

  98. 18:32

    So let's go and build this future together. Thanks so much. Check out the additional sessions we have. My colleague Mike is going much more into the rolling dice demo, everything MCP and Strands, and my colleague Suman tomorrow will also have a deep dive on Strands.

  99. 18:47

    And with that, thank you very much. Check us out in the expo hall and grab your own D20. [applause] [upbeat music]