AI Engineer World's Fair 2025
AI Red Teaming Agent: Azure AI Foundry — Nagkumar Arkalgud & Keiji Kanazawa, Microsoft
About this talk
Microsoft presenters Keiji Kanazawa and Nagkumar Arkalgud explain how adversarial prompts can circumvent model safeguards and why AI agents require systematic security and safety evaluation. They introduce Microsoft's AI Red Team, the open-source PyRIT framework, and its integration into Azure AI Foundry through an evaluation SDK and hosted reporting. Demonstrations include testing a locally running Ollama model and scanning an Azure OpenAI configuration, followed by an audience question about guardrails.
Chapters
- 0:00Introductions and AI security challenges
- 2:06Adversarial prompts, agent risks, and trustworthy engineering
- 4:17Microsoft AI Red Team, PyRIT, and Foundry integration
- 5:26Live demonstration with Ollama and Azure OpenAI
- 17:14Guardrails question and closing resources
Talk transcript
- 0:00
[upbeat music] Hey, thank you and welcome everybody.
- 0:17
Uh, I'm Keiji Kanazawa from Microsoft. I work in Product and AI Foundry. If you were at the keynote session from Asha, she's my CBB boss, so I work in that org.
- 0:26
And Nakumar is, uh, a engineer in the team. He's gonna be showing you some code in a little bit, which I'm sure most of you are excited to see.
- 0:35
So here we are at the AI Engineer Conference. I'm sure you're learning about, you know, all kinds of stuff, reinforcement learning, agents, V agents, evals, all kinds of stuff.
- 0:44
And so we're all really excited, you know, at least I am, to get AI into the hands of people, right? And help people, whether it's for, for, you know, your end users and your internal users.
- 0:57
And so AI is, you know, obviously where it's at, or I think it's where it's at, but it comes with a bunch of headlines that you've probably seen. And some of them, again, if you were at the keynote this morning with Simon, he showed some examples.
- 1:10
So you know, it's very easy, uh, it can be easy to trick chatbots into saying stuff that you don't necessarily want them to say. You can actually trick them into giving you information, uh, potentially that you don't want, you know, leaking out.
- 1:23
Um, and you know, AI is built. Like AI engineering, it's built on a whole ecosystem of different things, including Python packages, MPM packages, you know, other s- services that you may be using that, that are hosted, right?
- 1:36
And of course, so this San Francisco is the home of self-driving cars. Uh, you know, there, there... This, this picture is showing like a, you know, it's a, it's a frame from a video clip where a self-driving car is, is, you know, driving happily right past a school bus with a stop sign on it, right?
- 1:54
And if you think, you know, hey, is, you know, what does it have to do with me? You know, well, you, you... That this is one test that you can kind of think of in terms of whether something you built or thinking you're thi- something you're thinking of building.
- 2:06
Again, it, it, it's really easy to, you know, kind of get around some of the defenses of, of the AI models. So for example, the prompt on the left, if you say, "How to loot a bank," um, a lot of the models will actually, you know, refuse to answer, right?
- 2:20
They'll say like, "Oh, no, I can't help you with that." Uh, but if you... And, and some of the examples, if... So how many of you were at the Sanders workshop yesterday on prompt engineering red teaming?
- 2:30
Yeah, so like, it's really, like if you preface that question with a whole, but like maybe a bit of your life story, you can maybe convince the AI to tell you something, you know, that, that, that it, it's not supposed to.
- 2:42
And there are also other tricks. Like, um, on, on the right-hand side, it says, "Nabatuotwo," which is how to loot a bank ba- spelled backwards, right? Right to left.
- 2:51
And actually it turns out that's one of the patterns that an AI, uh, AI model can be tricked into giving you, uh, the answer. And so, and so you know when...
- 3:00
And then especially like, um, we talk started this morning also, but like of course it's all agents, agents, agents, 2025 is the year of agents. Um, and there are a lot of concerns, like if you talk to businesses about how in this world of agents, AI can be, you know, tricked into, into different kind of, uh, risks
- 3:19
and stuff and different malfunctions. So, but we're here at the AI Engineer Conference and, you know, like what, what I wanna [laughs] kinda like convey is that we as engineers know how to do this stuff, right?
- 3:33
So like engineers build bridges and dams that people trust. We build trucks and trains. And so AI engineering is early, so we've got a lot of work to do to get to the point where people trust AI as much as they trust, you know, bridges.
- 3:48
But, you know, this is something we know how to do as engineers. We, we build something, we iterate, we check it, we test it, you know, we continue to iterate.
- 3:57
Um, and that's what we're here to show you, uh, kind of how to do. And as engineers,
- 4:03
you know, we also rely on not just ourselves, but other people, right? So, uh, what we like to say is trust is a team sport. So when we, when we're looking to build trustworthy AI systems, we depend on other people.
- 4:17
So the engineers need to depend on people who have a lot of expertise in these areas like security and AI risk. So at Microsoft, we have a team called the Microsoft AI Red Team, which I've been working with for a few years actually.
- 4:30
They were one of the pioneers in identifying risks, you know, kind of, of, of AI in general as well as LLMs. Like two, three years ago they were already talking about, hey, you know what?
- 4:41
Like these GPT-3, GPT-4 models, you can get them to do things that, you know, you really want, don't want them to do. And so we partnered, uh, in Azure AI Foundry with the AI Red Team to offer a solution that makes it easy for you, AI engineers, to basically have a teammate that can help you with the
- 5:01
AI, uh, red teaming. And so they, so the AI Red Team, they built a, a Python package called PyRIT, PyP-Y-R-I-T. Uh, and what we offer is a hosted version, uh, and wrapped it around a easy to use SDK, and also hosted dashboard to show you the evals, you know, that come from this.
- 5:21
And so here, uh, to show you how it works is Nakumar.
- 5:26
Awesome. Thank you. Thank you, Keiji. Uh, here we go. So
- 5:36
hello. And so this is the sample project that I'm gonna run for you all. It's a simple RAG on PostgreSQL, uh, with an Azure samples.
- 5:54
I'll have this QR code up again. Uh, so running locally right now. You can ask simple questions like this, and it's talking to a locally running model via Ollama.
- 6:07
Uh, well, tool call didn't work. Live demos, right? [laughs]
- 6:11
So in this piece here, logs for everything. But, um, what we are trying to showcase is something called the semantic kernel, uh, agent, which here's some code for it.
- 6:26
Uh, it exposes-- It takes in the Azure chat completions, and our Red Team plugin, uh, is something that our SDK exposes. It has all the functions that are needed, uh, for, for an agent to tool call, call into our Red Team agent to help someone with their Red Team process.
- 6:45
And then it's simple chat completions agent afterwards. Uh, so for now, I, I'm gonna start running this. Um,
- 6:55
and when I run this, y- it will go through a few user inputs that I have, uh, hard-coded, and then we can jump into it in, in interactive mode.
- 7:04
But the target for this semantic kernel is going to be the same RAG app which, uh, uh, is running locally. So once this loads in... Yep. The first question...
- 7:15
Oops. Live demos. [laughs] Looks like tool calling isn't working today. Um, but anyways, so this call RAG app can be switched into a call to any other application which takes in a query as an input and then responds with a string as an output.
- 7:38
Uh, so internally, we ask you to call your application, which then you can run evaluations on. Uh, so in this agent mode, what would usually happen, I can scroll up to like a previous, previous output, um,
- 7:53
which ran earlier. Not, not lying to y'all. Um,
- 7:58
so these were the strategies which were available, and then, uh, use one of the... Uh, I asked it to like, "Hey, figure out, uh, get me a harmful prompt in the violence category."
- 8:10
And then it gives you, gives me some sort of prompt, and then I'm like, "Hey, send it to my target." And then this is what the target responds with.
- 8:17
And then there is some details about some sort of ski goggles and products that were supposed to be answered, uh, from our database. Uh, and then I tried to be like, "Hey, convert the prompt using Base64."
- 8:30
And the agent converts it, and then be like, "Hey, now send it to my target." And then the target responds with something else. So this is an easy Copilot-style way for anyone to get started to red teaming an application.
- 8:45
Now, we have-- we can take it a step further and run, uh, uh, run the whole scan end to end, and this is how you would run the scan.
- 8:54
You saw a little bit of code that Keiji showed earlier. Uh, so you usually set up your AI project, throw in the URL, and then you can set up, um, initialize it with your, uh, the, the URL and then your credentials.
- 9:07
You can select risk categories. We have four of these risk categories right now. These map to our evaluators. So, uh, this is how you set them up. If you don't include any, we include all by default.
- 9:18
And then the number of objectives is the number of questions that, uh, will be sent out to your application. Uh, and then the scan method looks like this. So you give it a scan name, you can give an optional output path which stores all the results there.
- 9:33
And then, uh, your attack strategies will include a list of different attack strategies. Um, I'll pull up a docs page later on, which has all the information about different risk stra-strategies that you can use.
- 9:44
Uh, there are combination strategies like Easy, which, uh, does like flip, the one that, you know, reverses the string and things like that. So there's also Mars. These are our simple converters which are, which live within PyRIT, but then expose via our SDK, so our SDK can offer it to people to use it easily.
- 10:04
Uh, you can also compose an attack with two different strategies. So con-- get a tense converted, uh, strategy and then do a URL-style conversion on it, so it does both and then sends it out.
- 10:17
And then you pass in a target, a target which supposedly decided not to work today. Again, [laughs] call to the same application. Um, and, uh, once you run this, you usually see an output which looks like this.
- 10:35
So... Oops. I'm gonna unplug the Ethernet cable.
- 10:47
Which... Okay. There we go. And so yeah.
- 11:03
So this is when I had it running yesterday with, uh, GPT-4.0 as my model, and GPT-4.0 comes with a lot of security inbuilt within our Azure AI Foundry. So once you have all those guardrails up, it kind of was pretty good.
- 11:19
Uh, it was a very small sample size, a hundred and sixty. I think I selected ten different harm types or like five harm types with s- a few categories.
- 11:28
So none of them was able to break into our, into our application. But then I switched it on and I used Phi-3. And with Phi-3, you can see the results show, uh, five out of forty in hate and fairness, uh, was successful.
- 11:42
So we can take a look at it, filter data based on what was successful. Um, and then you can, you know, look at like what was the response that was determined as, you know, harmful in our, from our evaluators.
- 11:56
Um, and finally, we have one more, uh, way of doing this. So initially, a lot of people might just be building models. You don't even have an application. You can directly call, uh, the scan against an Open-- Azure OpenAI config.
- 12:13
So if you have models running on, uh, our, uh, on Azure, you can set it up, uh, as a target which is just- You know, these three things, and then once you have these three things, you can run the whole scan.
- 12:25
This scan runs against, uh, the model directly and gives you a, an output. I guess this should be able to work. [laughs]
- 12:35
Let's see. Um, I can probably run it here.
- 12:48
There we go. So yep, scan model, runs a direct model scan. Um, and here I have some results from a pre-run. Was prepared for this. So [laughs] this is when four dot one, if you take off all the guardrails, here's the result.
- 13:03
Uh, it says that twenty-five percent of, uh, violence category was successful and, you know, twenty percent of all the difficult complexity attacks were successful. And again, you can filter out on the data, see which was successful.
- 13:18
Um, and yep, there we go. Lots of violence. And then this was the flip where, like, some sort of we, we can see what strategy it was. I think it was a Caesar encoding strategy.
- 13:33
And you can see that the assistant kind of decoded it. Uh, so we did not want that to happen, so we [laughs] wanted a successful attack. So that's one of the things.
- 13:43
And then here's a response when you set up all the guardrails, uh, with four GPT-4-1 Nano, and you can see that the difference is that we reduce our attack success by a little bit.
- 13:55
So that's an overview of how things go, uh, and how this scan is running. So it usually gives you an ETA, six minutes. We'll probably be running out of time by then, but yeah, as soon as this is done, it shoots you an UR- a URL which will directly take you to this page.
- 14:12
So that's safety evals in AI red teaming. Uh, I will be at the Microsoft booth, uh, towards the end for questions. So back to you, Keiji.
- 14:22
Yeah, thanks. [audience applauds] So basically, well, you know, the rest of the talk is, is, uh, is talking about essentially how this fits into an overall strategy, right? So that's... AI red teaming is, uh, really important and, you know, uh, part of kind of your defenses and, and kind of in your toolbox to be able to develop and
- 14:47
deploy trustworthy AI systems. But it really, what you wanna do is incorporate this within a whole, you know, kind of a, again, from the engineering mindset, a framework, you know, and kind of a process to, to get these things out.
- 15:01
So right, so first, what you wanna do is before, you know, you develop, de-develop like a production application that goes to customers, you wanna, uh, you know, you wanna kinda map out what are the kinda risks that, you know, we're, we're anticipating here?
- 15:13
Is this an agent? Is it, you know, using kinda like external data or internal customer data? You wanna think about what are the, you know, kinda the risks that, you know, your app is gonna have, plan for it, start to implement the guardrails in the first place, and then do the evaluations of which, you know, red teaming
- 15:30
is one of those, uh, possibilities, right? So within Azure AI Foundry, we have a suite of evaluators both for quality check, quality evals. I think there's a lot of talks today, you know, at AI engineer about quality, you know, kind of quality evals, right?
- 15:46
With something you could do in a, uh, AI Foundry. And then there's a wh- a, a whole set of risk and safety, uh, evaluators that we have, of which AI red, you know, red teaming agent is one of them.
- 15:57
But we, we also have a lot of different class, you know, kinda classifiers, uh, in terms of, of, you know, kind of both input as well as output, uh, because you wanna check for both.
- 16:08
Um, and then there's a specific set of, of evaluators we just, uh, created for agentic applications as well, like is the agent, uh, following your instructions well? Is it, you know, um, things like that.
- 16:21
And then you can add your own custom evaluators. And, uh, Nakumar showed you kinda like some of the mitigation strategies that you wanna apply, right? So for example, there are guardrails and controls that you can have in your application.
- 16:34
So once you've, you know, ran the A- AI red teaming agent, you figured out, well, actually twenty percent of, you know, like, stuff gets through. What do you do?
- 16:41
Well, that's the point at which you can apply the guardrails, um, you know, to, to content filters and other, other capabilities that, again, we have in Azure AI Foundry that makes it easy for you to add those guardrails.
- 16:54
And, uh, among other ones, there's ones called like Prompt Shields, which are for-- to guard against prompt-based attacks, especially the kind that, you know, are incorp, you know, involved with AI red teaming.
- 17:05
Um, and we have time for maybe, uh, one or two questions.
- 17:10
Yeah.
- 17:11
I think we have like a minute maybe.
- 17:12
Yeah, we have two more. So.
- 17:13
Yeah.
- 17:14
International coffee. But how does, um, guardrails work under the hood?
- 17:21
Uh-huh.
- 17:22
Like, is it, is it like filtering after it gets the answer, or is it, like, before the LLM even sees the prompt? Just, just-
- 17:32
There, so there, there's both kinds. Yeah, you can apply both, right? So you-
- 17:36
Yeah, but the guardrail feature.
- 17:37
Yes.
- 17:37
How, how does it work under the hood?
- 17:40
Maybe like-
- 17:41
Um, do you know? I mean, I think there, there are both... I think we have, uh, both filters, you know, for both for the input end as well as filters at the output end.
- 17:51
So it, it'll... So there are content filters, for example, you know, ba- basically if people are typing in, you know, uh, like, like, a CBR or like, you know, like, how to, you know, how to build a bomb kind of thing, right?
- 18:03
So there's the input guardrails. But then there's also like, uh, guardrails in terms of, well, actually, I wanna make sure I'm not outputting like sexual content or something, right?
- 18:12
So like there's also guardrails in terms of the output of the model as well.
- 18:16
And that, that's what's happening under the hood with the guardrail feature or-
- 18:20
Yeah, yeah, yeah. With the guardrails
- 18:21
... you have to implement several features to get that.
- 18:24
Oh, no, no. You, you... So there are guardrails that are, are... that AI Foundry offers directly and yeah. So that's what's, that, that's what's happening under the hood. So you, you have the ability, for example, to, to just give the raw model raw inputs, right?
- 18:37
And then if, if you turn off all content filters. And that, that was some of the example that actually Nakumar was showing-
- 18:42
Yeah
- 18:42
... where, yeah. So, so the guardrails are not in the model itself. The model is still the raw model, and the guardrails are actually kind of, uh, outside it.
- 18:51
Does that answer your question?
- 18:52
Thank you.
- 18:52
Okay. All right. Thanks. I think we're out of time.
- 18:55
Yeah.
- 18:55
Um, so, uh, thank you for coming. Uh, if, you know, to, you know, definitely start, get started with the AI red teaming. If you're not doing it today, definitely get started.
- 19:06
Uh, here's a link to the code, um, a-as well as the docs. And yeah, thank you for coming. And if, yeah, so if you have any questions, we will be, uh, you know, at the Microsoft booth, uh, you know, different parts of, you know, today and tomorrow.
- 19:19
So yeah.
- 19:20
Come find us. Thank you.
- 19:22
Thank you. [audience applauds] [outro music]