← All AI Engineer talks

AI Engineer World's Fair 2024

Build, Evaluate and Deploy a RAG-based retail copilot with Azure AI

About this talk

Microsoft advocates David Smith and Cedric Vidal lead a hands-on workshop building, evaluating, and deploying a retail copilot with retrieval-augmented generation, Azure AI Studio, a shared GitHub repository, and GitHub Codespaces. Additional Microsoft technical specialist Miguel Martinez assists participants. The discussion covers vector-backed product retrieval, provisioned Azure resources, application identity with Entra ID, and handling queries for merchandise absent from the retailer's catalog; GPT-4o is mentioned as a model intended for a future iteration of the workshop.

Chapters

  1. 0:00Production RAG workshop and Microsoft presenter introductions
  2. 3:01GitHub Codespaces, workshop repository, and provisioned Azure environment
  3. 59:55Azure AI Studio and application identity questions
  4. 1:43:37GPT-4o plans and retail catalog retrieval edge cases
  5. 1:57:07Microsoft conference resources and workshop wrap-up

Talk transcript

  1. 0:00

    [upbeat music] Uh, this is the workshop on developing a production-level RAG workflow.

  2. 0:20

    So you're in the right place if you wanna learn how to build the backend for a chat application that works also off of OpenAI, and builds its answers based on information, uh, that we draw from databases and vector databases.

  3. 0:33

    We'll see all about that in this presentation today. My name is David Smith. Uh, I'm a principal AI advocate at Microsoft. I've been with Microsoft for about eight years now, uh, after my startup was acquired, uh, back in the big data space, and I've been at Microsoft ever since.

  4. 0:51

    Um, my background is in data science. I also do a lot of work as, as a statistician, uh, and these days I'm a specialist in AI engineering. And I have with me today, uh, two other members, uh, from Microsoft that are also specialists in AI engineering.

  5. 1:05

    Uh, first of all I'll introduce Cedric Vidal, who's on my team in AI advocacy. If you wouldn't mind, introduce yourself, Cedric.

  6. 1:11

    Hello, everyone. Uh, so, uh, like David said, I'm, uh, Cedric Vidal. I'm a principal AI advocate at Microsoft, um, and, uh, I have a background in, uh, AI, self-driving cars, um, um, software design architectures, um, um, everything in between.

  7. 1:29

    Uh, I've been working in the space for twenty years, um, and, uh, today I'm gonna help, uh, David, uh, with the, the workshop. Um, welcome everyone.

  8. 1:38

    Uh, thank you, Cedric. And we've also got, uh, Miguel Martinez, uh, come all the way from Houston. Uh, he's a technical specialist at Microsoft. Miguel, tell the, the crowd a little bit about yourself.

  9. 1:48

    Absolutely. Hello, everyone. Welcome today. My name is Miguel Martinez. I am a senior technical specialist for data and AI at Microsoft. So a lot of our clients, you know, this can be startups, businesses, they hear about OpenAI and ChatGPT and all of those things, so they think about, "Well, how can I actually use it for my business?

  10. 2:09

    How can I use those tools to drive business value?" And that's where me and my team come in, and we help all of our clients develop some new solutions to drive that value.

  11. 2:20

    All right. Well, Cedric and Vidal will be here, uh, for the next two hours helping you as you go through this workshop, so they'll be wandering around if you guys wanna head out.

  12. 2:28

    And if you need any help during the workshop, raise your hand, and one of the three of us will come up and help you. So thanks, guys. All right.

  13. 2:35

    Let's jump right in. If you would like to get started, um, you can use that URL on your screen right there. Just pop up a browser on your laptop.

  14. 2:43

    Uh, I'll give you more information about what's going on when you get there, but, um, you should be able to follow along and get started if you like. Now,

  15. 2:53

    to participate in this workshop, it is going to be hands-on, so you will need to have your own laptop, and I can see, uh, looks like everybody does have their own laptop.

  16. 3:01

    That's great. This is not something you'll be able to do, uh, on your phone or on a tablet, uh, because there's lots of, uh, lots of work that we'll be going through as we work through a GitHub repository.

  17. 3:10

    And that is the second thing that you'll need to have to participate in this workshop, is a GitHub account. Um, if you don't yet have a GitHub account, uh, please go ahead now to [REDACTED:url] and create yourself a brand-new account.

  18. 3:26

    Uh, we'll be using the GitHub Codespaces feature, uh, to provide our development environments, but the free GitHub Codespaces is more than sufficient for the work that we'll be doing, uh, here today.

  19. 3:38

    We are going to be building an application in the Azure cloud, but you do not need to have an Azure account for this workshop. Uh, we're gonna provide you with a login to an Azure account, and we will have already set up all the resources that you need to work with this application.

  20. 3:55

    And that specifically is things like, uh, Azure AI Search, the vector database. Cosmos DB, uh, the database which we're gonna be using for our customer information. Uh, connections to OpenAI, which we're gonna be using for our LLM, and various other resources and tools that we'll be using in Azure.

  21. 4:14

    You can do all of this on your own, um, if you do have your own Azure account and are willing to spend a few dollars in credits, uh, to run the resources for a few hours, uh, for this resource.

  22. 4:24

    By the end of this workshop, you will ha-have, excuse me, you will have, um, all of the information, code, data, and everything you need to recreate everything that I show you here today.

  23. 4:35

    Uh, one of the first things that we will do, in fact, is to clone, uh, to fork, in fact, a repository into your own GitHub account, and everything you need, uh, will be right there in there.

  24. 4:45

    All right. Um, but if you would like to run through this at home, you will need an Azure account, and you can create one at, uh, the link you see on the screen right there.

  25. 4:54

    So before we jump in and I start giving you some demos about how this works, let me orient you a little bit. Um, that link that I gave you on the last slide will have launched you into a virtual machine.

  26. 5:07

    It is a Windows virtual machine, but we're not gonna be using Windows at all. Uh, in fact, the only thing we're gonna be using this virtual machine for is the instructions for the workshop, which you can see on the right-hand side of the screen there in front of you.

  27. 5:20

    You'll be going through that page after page. Um, this instruction workshop feature has some nice tools that you can use. For example, wherever you see the green text, you can click on the green text, and it will paste that directly into your browser.

  28. 5:36

    Um, that'll save you some time when you're entering in the passwords and URLs and things like that, and so forth.

  29. 5:42

    I will mention, um, that we are in a virtual machine environment. It looks like the Wi-Fi's pretty good here. Um, but if you find that the virtual machine is slow, um, or there's lots of lagginess because of the Wi-Fi, or if you just prefer to use your own browser on your own laptop, you can totally do that.

  30. 5:58

    You can just open up a browser and follow all the same set of instructions. The only difference is you'll have to open up that virtual machine every now and again to look at those instructions, and you'll have to manually cut and paste the green parts, uh, from the guide into your own browser.

  31. 6:12

    So it's your choice about which direction you wanna go there

  32. 6:20

    When we actually come to running shell commands and things like that, uh, the dev environment that we're going to have you use is GitHub Codespaces. You could totally run all this on your local desktop.

  33. 6:33

    Um, but we've gotta make sure you've got VS Code installed, you gotta have all the right Python libraries installed, et cetera, et cetera. So to make things easy here, we're just asking everybody to go straight into GitHub Codespaces where everything is set up, and you can do exactly the same process for yourself at home.

  34. 6:49

    If you're not familiar with GitHub or not familiar with GitHub Codespaces, just pop your hand up when we get started. We'll come and chat to you and talk to you about how that works and what's going on there, on there.

  35. 6:59

    But if you have used GitHub Codespaces before, it should be, uh, pretty straightforward right there.

  36. 7:07

    All right. So this is what we're gonna do today. We are gonna build this app right here, or at least we're gonna build the back end to this app and have you interact with that back end through a little UI that you'll be working through.

  37. 7:22

    If you do wanna build the front end to this as well, we do provide the code for the front end and all the data. Uh, that's also linked from the GitHub repository, so you'll have, have access to that.

  38. 7:31

    But I'm gonna give you a little demo, uh, about this website just to set the scene. So let me go over to my browser here. I think I have the website open.

  39. 7:42

    So the, the, um, the sort of the, the idea here is that we are engineers, uh, working for a retail company. You know, something like REI that sells camping equipment and backpacks and trail hiking shoes, things like that.

  40. 7:58

    They already have a website which you can scroll through, and you can see the products that they have available. It's a pretty limited selection just to keep things simple, and we can click through to see product information for each of the products that this company sells.

  41. 8:12

    So for these, these, uh, Trailblaze hik- hiking pants, we find lots of details about the features. There are some reviews, FAQs. There's a return policy, um, some cautions, technical specifications, user guide, care and maintenance.

  42. 8:25

    Lots and lots of information available to the customer about all the products that are available at this store. In addition, this storefront has a customer login feature. In this particular example, the customer, Sarah Lee, is already logged, uh, into the system.

  43. 8:44

    What we're going to be building here is a chatbot that operates on this website. You'll be able to access that, click on the chatbot button and ask the question, for example, [keyboard clicking] "What can you do?"

  44. 9:00

    And this is gonna connect to the system that we're going to build to get an answer to that question. This is a pretty simple question. It's coming straight from the LLM and its context to tell it that as, as an AI agent, I can provide you with information about our products to help you with your purchases, so

  45. 9:15

    on and so forth. Now, we can also ask questions about specific products. For example, I'm gonna pop up my paste here. And let's try one like... Here we go. [keyboard clicking]

  46. 9:29

    "What is a good tent that goes with, um,

  47. 9:35

    the Trail Walker shoes?" Okay. So in this case, it's actually consulting all the informati- product information that's available in that website to formula that out, formulate that answer with the LLM, and it comes up with, "For your upcoming trip to Andalusia, I recommend pairing your Trail Walker hiking shoes with the TrailMaster X4 tent."

  48. 9:59

    We're also gonna provide the LLM with information about the customer themselves, their name, their status in the loyalty program, where they live, and their order history, so it's able to answer questions like this:

  49. 10:12

    "What have I already purchased from you?" So in this case, the LLM is able to consult the customer's purchase history and give back the answer to say that Sarah Lee, a valued customer, who has purchased the Cozy Night sleeping bag, the Trailblazer hiking tent, and so on and so forth.

  50. 10:29

    So this is the system that you'll be building on the back end so that that chatbot can use an LLM like OpenAI to answer those kinds of questions based on the customer's purchase history and the products available from this particular retail store.

  51. 10:48

    Any questions on that so far? Very straightforward. Okay, great.

  52. 10:55

    All right, we are gonna be building this on Azure and using various Azure resources, which we'll provide to you to do that. Uh, we're gonna be using the Azure AI Studio platform to manage the large language models and the various resources that we're working with.

  53. 11:11

    Uh, we're going to be using Azure AI Services, which are going to provide, uh, various tools to us, but in particular the OpenAI models that the chatbot is going to use to generate its responses.

  54. 11:23

    We'll be building an Azure AI project, which is gonna manage all the flows that we have to, to, uh, have the chatbot gather its information and generate its responses.

  55. 11:35

    The product database is going to be stored in Azure AI Search. And Azure AI Search is a, is a vector database, or at least it has a vector database feature, which we're going to be using to match the customer's questions to the nearest or the most relevant products that are related to those customer questions, and using that

  56. 11:55

    to provide context to the chatbot. And lastly, the customer information is going to be stored in a database, a regular, uh, rows and columns type of database, relational database, where we will extract the customer's information and order history.

  57. 12:13

    Yeah.

  58. 12:15

    Turn it on. [laughs]

  59. 12:18

    Okay, so let's have a look at the architecture we're building for the back end in a little bit more detail.

  60. 12:26

    Customer types a question into the chatbot on the website, and that gets extracted out and sent to an endpoint just as a simple string, along with the customer's chat history as well.

  61. 12:38

    That question gets sent to Azure OpenAI, where it will get embedded into a vector format. If you're not familiar with embedding, it's basically the idea of converting a piece of text, piece of string, into a point in multidimensional space in such a way that the other points in multidimensional space we have already defined by embedding the product

  62. 12:59

    information pages, the ones that are closest to the question are the products that are most relevant to that. And we'll use that to extract out the most relevant products and push them into the context of the large language model.

  63. 13:13

    At the same time, we'll also extract out information about the customer and their purchase history from Cosmos DB.

  64. 13:20

    So we do the embedding, do the search in Azure Search to get the product information. We extract out the customer information with Cosmos DB.

  65. 13:28

    We feed all that information into a prompt. So we create a large prompt with information about the problem to solve. We have the customer information about our products, information about the relevant products from the Azure AI Vector search, information about the product's customer history, and then generating a response based on the, the customer's question.

  66. 13:49

    Any questions so far? All right, we'll get right into it just shortly.

  67. 13:56

    I'll come into all this a little bit later on, but we're gonna be using a tool within Azure AI Search called Prompt flow. I'll come back into the details a little bit later on.

  68. 14:05

    But what we're going to be doing is taking that sta-- that same retrieval-augmented generation process that I just outlined to you. The steps are to take the question, retrieve related data, which in this case is our product data, augment the prompt with that information that's come from the knowledge base, generate a response with the large language model,

  69. 14:27

    and then send the results back to the user. We'll be creating a version of that res-- uh, retrieval-augmented generation flow using Prompt flow, which is a development tool that's in Azure AI Studio to streamline this process of putting all that data together.

  70. 14:42

    We take the inputs, we embed it, we retrieve information from AI Search, we look up customer information with a database, put all that together with the customer prompt to generate the response, and that's what the customer sees in the website.

  71. 14:55

    So that's what you will be doing. So with that, let me go back to the start here. Has anybody not yet been able to log in to that website?

  72. 15:06

    Have you-- Anybody having trouble? If so, um, Cedric and, uh, Miguel will help you guys out. Let-

  73. 15:13

    The password?

  74. 15:14

    The password will come up in just a second. All right. So let me actually show you that.

  75. 15:18

    Yeah, it's gonna be [inaudible].

  76. 15:20

    Okay. So I've gone to that website, and you should have rece-- come to a page just like this, where you can go ahead and launch the virtual machine.

  77. 15:32

    Can I just get a admin password?

  78. 15:35

    Yep. It'll-- I'll show you how to get to that in just a second.

  79. 15:38

    Okay. So on the right-hand side of the virtual machine is the instructions. At the bottom right-hand corner is a Next button. That'll take you to the next page of the instructions.

  80. 15:49

    And that is the page where you'll find the password for the Windows machine, which is, curiously, password.

  81. 15:59

    And from then, you should be able to log into the Windows virtual machine right there.

  82. 16:08

    Okay. The next step after that, and I can actually go to the next step here,

  83. 16:15

    is to open up the browser. I'm doing this directly within the virtual machine.

  84. 16:22

    We're gonna browse to a particular GitHub repository that we provided to you for this workshop. You can just click on the link there and have it typed into the browser in the virtual machine, or you can cut and paste it into your own browser as you prefer.

  85. 16:39

    Okay. Um, you'll need to log in to, um, GitHub at this point. Um, just a, a slight warning is that if you use single sign-on, this happens to us at Microsoft, we use single sign-on to access GitHub.

  86. 16:54

    That does not work through the virtual machine, um, so I'm gonna show you actually doing that directly from the browser

  87. 17:01

    where I've already logged in to GitHub there.

  88. 17:07

    And we're gonna go to Codespaces. Oh, it looks like I've actually already started a codespace here, so let me just go ahead. Um, the instructions will tell you to launch a codespace.

  89. 17:16

    Oh, I forgot one very important step. Before you launch any codespaces, we're gonna switch to a different branch. Um, rather than doing the main branch, which is the one that you'll use if you wanna do this at home, we've got a special branch here for this lab.

  90. 17:30

    It's called ms-build-lab-three-two-two. The only difference is this version of the lab skips all the deployment instructions, 'cause we've done that for you already.

  91. 17:41

    And then from that f- from that branch of the, um, of the repository, we're going to launch a codespace on that specific, um, branch. I'll show you that in just a moment.

  92. 17:54

    But to launch the one that I already have here,

  93. 17:58

    this is what happens once Codespaces set up. Takes a couple of minutes for Codespaces to warm up, warm up. This one should already be warm.

  94. 18:11

    And what we have here, in case you've never seen this before, is basically an ins-instance of Visual Studio Code, uh, running in the GitHub cloud directly in your browser.

  95. 18:23

    So it's the same user interface as Visual Studio Code, but it's just running in your browser. Uh, if you're experienced using Codespaces, you probably know that you can also run this directly from within Visual Studio Code on your desktop and connect to that instance.

  96. 18:35

    Uh, if you've got questions about that, happy to show you.

  97. 18:40

    One of the things that we are going to be doing with Codespaces, and it looks like I've already

  98. 18:47

    closed the terminal, is logging in to your Azure account, um, at the terminal. And the instructions to that, if you've been following along, you might be ahead of me already,

  99. 19:03

    is to log in to, um, the portal, and actually do this directly within the, the local browser in this case.

  100. 19:18

    Okay. This is the username we provided to you for this temporary Azure account. It'll only exist, um, for the duration of this workshop. E- everything will be deleted once we're finished.

  101. 19:36

    One thing you'll probably find is because we're in a brand-new Windows virtual machine, it thinks you're a brand-new user to Windows. It thinks you're a brand-new user to Azure.

  102. 19:44

    It keeps on popping up all these hints. Um, if you're familiar with Azure, you can just delete them. If you'd like to do an introduction to Azure, and you can do it here, or you could do it at home.

  103. 19:54

    And the trick is to get back to the starting place is to go to the Home button there, and then I won't go through these steps, uh, on the screen for you, but you'll be following through the steps on here to have a look at the resource groups we have available.

  104. 20:07

    Also, the resource group we've provided to you in Azure, and you'll be able to inspect all the resources that we've already deployed for you that you'll be working with.

  105. 20:16

    Lastly, I'll just show you this last part.

  106. 20:22

    Is when we, um, log into Visual Studio Code,

  107. 20:27

    let me show an example of copying. Oops. Sorry, it don't normally do it this way.

  108. 20:36

    There we go. Where is my copy button? I'm not using my usual laptop here.

  109. 20:50

    All right, so what I will do, this might be a neat trick, is just paste it directly up here

  110. 20:57

    and copy that back into my other Visual Studio window.

  111. 21:08

    Allow. There we go. All right. It's usually easier than that, just to show you that process of how you're going to connect your Visual St- a Visual Studio Code instance in GitHub Codespaces with your Azure account is with this AZ login.

  112. 21:27

    After a moment, it will give me an eight-digit code, which I'll copy, and I'll go to this device login page to actually log in to Azure.

  113. 21:40

    And now all the commands that I run within the Visual Studio Code command line terminal will run against the Azure account that we've set up for you. So that's kind of the main things to watch out for as you get started.

  114. 21:52

    Just continue to work with the, through the instructions as you go. Cedric, Miguel, and I will be wandering around to help you, and put your hand up if you get stuck.

  115. 22:00

    And in about twenty minutes or so, I'll jump in and give you a bit more context about the prompt flow that we're building. Thanks. [background chattering]

  116. 22:33

    Yeah. [background chattering]

  117. 22:56

    Yeah. [background chattering] Yeah. [background chattering] Yeah. [background chattering] Oh, okay. [laughs] Okay.

  118. 23:11

    You figured it out now? Yeah.

  119. 23:12

    Yeah, I think so.

  120. 23:12

    Okay.

  121. 23:13

    I think, uh... So-

  122. 23:13

    Yeah

  123. 23:13

    ... I was just about to try and launch the Codespaces. I don't know if, uh-

  124. 23:17

    Yep.

  125. 23:17

    And so, um, yeah, this is building.

  126. 23:22

    Yeah.

  127. 23:23

    At the same time.

  128. 23:24

    Yep, yep.

  129. 23:25

    It's like having the Visual Studio-

  130. 23:28

    So if you're already logged into GitHub, it'll just continue to use that login. Yeah.

  131. 23:32

    Which is like it already has everything installed all the time. It's in the container.

  132. 23:38

    Yeah, b- yeah, before you do that, though, um-

  133. 23:40

    Read the instructions.

  134. 23:41

    Yeah, yeah.

  135. 23:42

    Okay, so I'm logged in.

  136. 23:43

    Yep.

  137. 23:46

    Yeah.

  138. 23:48

    It looks like you're logged in, yeah. Yep, yep. As long as you're logged in, you're good.

  139. 23:52

    Okay.

  140. 23:52

    Yeah.

  141. 23:53

    Okay, yeah.

  142. 23:55

    Yeah. Yeah. And when you forked it, did you, um, did you uncheck this copy of the main branch only?

  143. 24:01

    Uh, oh, I haven't forked it yet.

  144. 24:02

    Okay. Oh, good. Yep.

  145. 24:02

    So glad you mentioned that.

  146. 24:04

    Yeah.

  147. 24:05

    Oh, wait. This is the, this is what I got to fork, then? The Contoso tech chat?

  148. 24:08

    Yep, yep. So click the fork button first of all. Yeah. Which is up here, yeah.

  149. 24:12

    And then-

  150. 24:15

    And be sure to uncheck that box.

  151. 24:18

    So you-- Which step are you doing? Okay.

  152. 24:20

    Yeah, so you can-

  153. 24:21

    And then I-

  154. 24:22

    Yep

  155. 24:22

    ... there are two ways to do like this. The easiest is to just hit-

  156. 24:28

    And the nice thing is now whatever GitHub account you're using, you'll have this in that account-

  157. 24:32

    Right

  158. 24:32

    ... to work from when you get home.

  159. 24:33

    Oh, yeah, yeah.

  160. 24:34

    Yeah.

  161. 24:34

    Yeah, I love that.

  162. 24:35

    Yep.

  163. 24:36

    I mean, the setup's awesome.

  164. 24:37

    Mm-hmm.

  165. 24:37

    I love that you have the, you have everything-

  166. 24:39

    Yeah

  167. 24:42

    So, for the repo, we just did that. I can check it.

  168. 24:45

    Yep, yep.

  169. 24:47

    And then you pull up again on this.

  170. 24:50

    Just telling you that you just skipped ahead, yeah.

  171. 24:53

    Oh, I did. Right.

  172. 24:54

    Yeah, it doesn't matter. You can-

  173. 24:55

    Um, should launch the cases.

  174. 24:58

    Yep.

  175. 24:59

    Oh, I gotta go to that branch. Okay.

  176. 25:00

    Yep.

  177. 25:02

    Uh.

  178. 25:04

    It's the MS Lab four through whatever.

  179. 25:07

    This one?

  180. 25:07

    Yep, that one. Go to Code, and then Codespaces.

  181. 25:14

    So, I didn't try Codespaces.

  182. 25:15

    Yeah, but you turned it off, yeah. So click on the Codespaces tab up there, yeah.

  183. 25:19

    You can see that I haven't seen [laughs]

  184. 25:20

    Yeah. And then click Create Codespace.

  185. 25:23

    Okay.

  186. 25:23

    Yeah.

  187. 25:24

    So that's available just to my, like, private account?

  188. 25:26

    Yep. Yep. Yeah, it's, um, it's limited in how much you can use, like it's twenty hours worth of computer stuff. But-

  189. 25:32

    Right

  190. 25:32

    ... you know, those kinds of things. But-

  191. 25:33

    Right

  192. 25:34

    ... it, it's a nice consistent environment. And what we can do on our end, which we do for all of our samples, is we make sure that when you launch Codespaces, you get a command line that has everything installed that you need to go.

  193. 25:43

    So there's no-

  194. 25:44

    Right.

  195. 25:45

    Yeah.

  196. 25:45

    Right. Right.

  197. 25:46

    Yeah.

  198. 25:46

    Uh, and then, uh, the-

  199. 25:48

    Yep

  200. 25:49

    ... the... Oh, wait.

  201. 25:52

    Yes.

  202. 25:52

    Oh, it is-

  203. 25:52

    It's, it's doing it. Yep. Yep. So that's gonna take probably sixty seconds or so to, to, to thing. So what we actually recommend is you go to the next step, and then come back to it later.

  204. 26:00

    Like a new market or something.

  205. 26:01

    Yeah. Yeah. All right. I'll let you go to it. [laughs]

  206. 26:08

    Uh, so I mean, it-

  207. 26:10

    Mm-hmm

  208. 26:10

    ... it looks like you gave me an Azure account.

  209. 26:12

    Yep. Yep, and there's the username-

  210. 26:14

    Yeah

  211. 26:14

    ... and password.

  212. 26:14

    I was, I was able to log in. Oh, wait.

  213. 26:16

    Yeah.

  214. 26:16

    It's right here. Yeah.

  215. 26:17

    Yep.

  216. 26:17

    So I had all the resources.

  217. 26:19

    Exactly, all deployed for you.

  218. 26:20

    Nice that they're all provisioned though.

  219. 26:21

    Yeah, 'cause some of them take a little while to provision. So it's, it's for these workshop environments, it's easy to-- for people to be able to jump straight in.

  220. 26:27

    Um, and, uh, but th- none of this is AI Studio though, right?

  221. 26:32

    Um, one of them will be. You should have-

  222. 26:33

    Oh

  223. 26:33

    ... eleven resources.

  224. 26:34

    Okay.

  225. 26:34

    So you got an Azure AI project.

  226. 26:36

    Mm-hmm.

  227. 26:36

    So that's an AI Studio project. And then-

  228. 26:39

    Oh, okay.

  229. 26:39

    Yep, and the next step, we're gonna actually log into AI Studio.

  230. 26:42

    Okay.

  231. 26:42

    Yep.

  232. 26:42

    All right. [laughs] [background chattering] No, no. That one, I'm, I'm not familiar. I think, you know, you gotta-- you will have to register something.

  233. 27:00

    It should be, it should be able to watch you, once you find out. But the, the good news is, like, all of this, it should be available on, in a GitHub.

  234. 27:06

    So-

  235. 27:07

    Yeah

  236. 27:07

    ... today-

  237. 27:08

    It's, it's good to go too.

  238. 27:09

    Yeah. So today, you know, the VM is closed, but you still have the GitHub. So later on, you can [inaudible] everything closed. Yep.

  239. 27:15

    Perfect. Thanks.

  240. 27:18

    Great.

  241. 27:22

    And regarding the, uh, Codespace, uh, how could I view it?

  242. 27:29

    Um, nope, there you go. Yeah. [background chattering]

  243. 28:05

    Sorry, what? I, I don't understand what I'm supposed to look for. Like, the image or?

  244. 28:11

    Hmm, everything going all right so far?

  245. 28:13

    The-

  246. 28:13

    Yeah.

  247. 28:14

    Documentation.

  248. 28:14

    Cool.

  249. 28:14

    Which part? Um.

  250. 28:16

    Isn't clear?

  251. 28:16

    This, look at type column for which...

  252. 28:21

    Oh, because, uh, you see, it says-

  253. 28:23

    Should be the second tab across, I think

  254. 28:24

    ... resources listed. [clears throat]

  255. 28:26

    Codespaces. Yep.

  256. 28:27

    Contoso chat resources.

  257. 28:28

    Yep, that, that. Yep. That one's Codespaces, the one before.

  258. 28:30

    An area of the type.

  259. 28:32

    What's this?

  260. 28:32

    Oh.

  261. 28:32

    So basically saying-

  262. 28:33

    This tab right here where it says Preview README.

  263. 28:35

    Got it. Okay.

  264. 28:35

    That's Codespaces right now.

  265. 28:38

    Okay.

  266. 28:38

    So it's showing you-

  267. 28:39

    Yes, so-

  268. 28:39

    ... the contents of the repository with a command line connected to that file system. [clears throat]

  269. 28:44

    Oh, I see. Okay. Thank you.

  270. 28:47

    I'm, I'm so used to using-

  271. 28:48

    Oh, you don't-

  272. 28:48

    Sorry. Oh, I don't do that. Okay. I'm sorry. I'm so used to just using VS Code, [laughs] I don't even

  273. 28:54

    get it, this other stuff. Sorry. So-

  274. 28:58

    I'm not used. Um, it doesn't seem like you are inside the-

  275. 29:02

    What?

  276. 29:03

    ... the VM.

  277. 29:04

    I'm not? Uh, I'm logged into the VM.

  278. 29:08

    Yeah, I wanna support because I have the blockchain break.

  279. 29:13

    No, but you should, um... It's weird because you should see icons here. You should be in Windows.

  280. 29:18

    Oh, I had it earlier, actually. It showed up. Um, I had it. Yeah, it showed up earlier. I don't know what happened. It, it used to have all the icons, and it looked-- it had showed I Explore and everything else, right?

  281. 29:30

    I guess it just stopped. But yeah, I did it earlier, so I guess it just wiped itself or something. [laughs] That's really odd 'cause, uh, I did that earlier.

  282. 29:39

    Okay. It's okay.

  283. 29:40

    Okay. No, I had it earlier, so-

  284. 29:41

    It was incorrect.

  285. 29:43

    Yeah, it did that the first time. [laughs] You have to, like, actually manually type it on Apple. But I already logged into this, like, ages ago. This is the first thing I did, so...

  286. 29:55

    I mean, that's the reason why it's not loading, potentially.

  287. 29:59

    Oh, no. Here we go.

  288. 30:00

    Yeah, no.

  289. 30:00

    Now you can open the browser.

  290. 30:02

    Oh, I did that earlier. I don't know why it wiped everything. [laughs]

  291. 30:06

    Yeah, and I went-

  292. 30:07

    Yeah, because you know the tab that you opened outside? I don't know. That's what I think.

  293. 30:10

    Oh, that's my other-

  294. 30:11

    Okay, okay

  295. 30:12

    ... GitHub, so that's why. That's it. Uh, I'm just saying, I did this earlier already.

  296. 30:15

    Okay.

  297. 30:16

    Yeah. So that's why I'm confused because I already have this, which is the, um-

  298. 30:22

    But the thing is, you should do that inside the virtual machine, Windows virtual machine.

  299. 30:27

    Oh, I thought I did.

  300. 30:28

    Because it won't run it separately.

  301. 30:30

    Yeah. Well, I did that already. [laughs] that's what I'm trying to tell you. Like, I did that, like, how many steps ago?

  302. 30:34

    Okay.

  303. 30:35

    Yeah. So I mean, I can just do it this way.

  304. 30:37

    Yeah, I guess you need to do it again. Sorry. [laughs]

  305. 30:39

    It's okay. It's just-

  306. 30:39

    And so you should be able to just click on it, and it should-

  307. 30:41

    Oh

  308. 30:41

    ... tab it for you. You don't have to copy paste yourself.

  309. 30:43

    Oh, okay. Well, I already[laughs] I just did this, like, literally. Is it just because I'm in a different window? Like, I wanna just save this file to my computer and then refresh it.

  310. 30:56

    Like, for whatever reason, it's just not loading. It's just pointless. Like, it's so weird. Like... Well, no, the reason I'm saying that is 'cause I had this load. I had everything else loaded, so now I just have to do all of it again.

  311. 31:08

    The only thing I stopped at was after I loaded this 'cause I just made a new tab.

  312. 31:13

    Oh, I see.

  313. 31:15

    I understand that.

  314. 31:16

    Right.

  315. 31:16

    Yeah. Yeah, when I got to Azure from Visual Studio-

  316. 31:17

    Basically, what you're doing right now, you're-

  317. 31:19

    Mm-hmm

  318. 31:19

    ... inside a VM. You're using the Azure subscription.

  319. 31:22

    Mm-hmm. No, I understand that. I'm just saying your stuff is buggy because it shouldn't, like, not work across windows, 'cause I literally just did this. Like, I shouldn't have to go back five times to, like-

  320. 31:33

    Say that again? Oh, no, no, I'm just catching up to you guys. Yeah.

  321. 31:36

    Uh, on the left, what you're about to do, or-

  322. 31:38

    Yeah, well, on your team, did you ever do the same thing? Did you ever do Zoom and GitHub?

  323. 31:40

    Uh, I'm just logging into GitHub right now.

  324. 31:43

    Yeah, that's where it was earlier.

  325. 31:45

    There we go.

  326. 31:46

    I'm sorry. [laughs] I'm just... I have to, like, redo everything 'cause it just-

  327. 31:50

    Yeah.

  328. 31:50

    Basically-

  329. 31:51

    That's what it is

  330. 31:51

    ... yeah, it didn't refresh for me. It just didn't.

  331. 31:54

    It didn't load. It basically-

  332. 31:57

    Like-

  333. 31:59

    Yes. There we go.

  334. 32:00

    Yeah, I did this, and then did this.

  335. 32:01

    Just to avoid confusion, you might want to close the other tabs. Otherwise, it's confusing.

  336. 32:06

    No, that's not what's confusing it. Like, these are other tabs from my doc. They have nothing to do with that. Like, that's not an issue. The issue I'm having is you're saying this thing should be loading from...

  337. 32:20

    Like, literally, I'm just trying to understand what to do after this, because I already have this loaded. Yeah. The MSBuild, I already have that-

  338. 32:29

    Okay

  339. 32:29

    ... which I have in a different tab. So after that,

  340. 32:34

    what I, what I was stuck was on-

  341. 32:34

    Okay, but what I'm saying is that you open these inside your macOS browser instead of inside the VM.

  342. 32:43

    Oh, okay. I guess.

  343. 32:44

    That's what I was trying to do.

  344. 32:45

    Got it. So copy the main branch only.

  345. 32:49

    I did this already. [laughs]

  346. 32:51

    So that's why I was saying that it would be better-

  347. 32:53

    Oh, because I did-

  348. 32:53

    ... to close those tabs to avoid confusion

  349. 32:54

    Got it, got it, got it, got it, got it. Okay.

  350. 32:55

    'Cause otherwise, that's what's confusing.

  351. 32:58

    Oh, I was just hitting reload in the Codespace, and I did it in a different browser.

  352. 33:03

    Yes.

  353. 33:04

    So do I recreate-

  354. 33:06

    No, you could just open the one you al- had already created.

  355. 33:09

    Oh, where?

  356. 33:10

    Here. Yes.

  357. 33:15

    I think it's over here.

  358. 33:19

    Yes.

  359. 33:19

    Okay, so then... Yeah, that works. That's fine. The issue I had is when I bring out-

  360. 33:27

    The AZ login takes a while? Or-

  361. 33:30

    Takes about 60 seconds sometimes for the, um, for the code to pop up.

  362. 33:33

    I keep wondering if I did it. [laughs]

  363. 33:35

    Yeah, yeah. And you do need to hit Enter in the terminal afterwards as well to confirm.

  364. 33:41

    In the new terminal?

  365. 33:42

    Uh, okay.

  366. 33:43

    Yep. Yep, take a look.

  367. 33:44

    Yes, in this browser.

  368. 33:47

    Okay. And then-

  369. 33:47

    And then open the portal.

  370. 33:48

    Were you able to-

  371. 33:50

    Basically redo what you've done, but do it inside the VM instead of your macOS browser.

  372. 33:55

    Yeah.

  373. 33:55

    Is what I'm saying.

  374. 33:58

    Yeah, I just, uh-

  375. 33:59

    Yeah, I think-

  376. 33:59

    ... took a AI architect role at, uh, eighty four fifty one, which is like a data science arm of the program. So-

  377. 34:05

    Okay

  378. 34:05

    ... yeah, we've had... Well, and Azure is our main, uh, place for doing things. No, we haven't used, uh, Azure AI Studio, uh-

  379. 34:13

    I think that's the reason why

  380. 34:14

    ... 'cause we've been, like, a big Databricks shop with, uh-

  381. 34:16

    Okay.

  382. 34:17

    Can you, uh... Yeah, open this.

  383. 34:19

    Oh, you can close that, by the way.

  384. 34:19

    Yeah.

  385. 34:20

    Thank you.

  386. 34:21

    Yeah, it was funny. [laughs]

  387. 34:24

    So, uh, let's see if it's done the raw-

  388. 34:27

    Yeah.

  389. 34:28

    So here...

  390. 34:28

    So I'll just start over?

  391. 34:30

    Yes. It's, uh-

  392. 34:32

    Oh, I literally just copied and pasted into my-

  393. 34:36

    The last command worked, right? So we're here in provision one. That stage provision.

  394. 34:43

    That's it. You're connected, and you can close that tab now. You won't need that one again. [clears throat]

  395. 34:47

    Might work this time since refreshing the credential waited.

  396. 34:51

    Maybe. [laughs] You know, just luck. Error.

  397. 34:55

    It'll show again. And just press Enter on that to confirm. Yep.

  398. 35:02

    Yes.

  399. 35:04

    So right now it's gonna create the Azure DB table and the data.

  400. 35:15

    Yeah, now you should be good.

  401. 35:24

    Where am I entering in any of this information from the-

  402. 35:31

    So now you should... Oh.

  403. 35:33

    Oh, that one worked. Thank you. Perfect.

  404. 35:35

    Can you-

  405. 35:35

    All right. [laughs]

  406. 35:37

    It said something about-

  407. 35:38

    Yeah, but I-- except you are-

  408. 35:41

    Right

  409. 35:41

    ... in your macOS.

  410. 35:42

    So now you should be able to go to Azure and see the-

  411. 35:44

    No. I mean, yes, but inside the VM browser.

  412. 35:46

    Oh, okay. That's what was... It was trying to do that.

  413. 35:50

    Yes.

  414. 35:51

    So recreating, and that's just for... That is for VS Code. To, so basically you clone VS Code within browser.

  415. 36:05

    It's good.

  416. 36:06

    Um, yeah, but you have already-

  417. 36:08

    All right

  418. 36:09

    ... the Codespace opened here.

  419. 36:10

    Let me show you where you'll get to once you get to about step seven.

  420. 36:14

    Then you have open your Visual Studio.

  421. 36:16

    Okay, perfect.

  422. 36:16

    Perfect. There it is.

  423. 36:16

    Yes.

  424. 36:16

    I've gone through the steps of logging into GitHub, cloning the repository, launching Codespaces, logging into the Azure portal and AI Studio from the browser, and also logging into Azure from the Visual Studio command line.

  425. 36:34

    Good.

  426. 36:35

    So that can run some commands, and I've run this post-provision command here.

  427. 36:39

    No, you might want to-

  428. 36:39

    You're welcome to have a look at what that script does-

  429. 36:41

    ... minimize it

  430. 36:41

    ... uh, but let me actually show you

  431. 36:43

    ... to, because-

  432. 36:45

    What we have done-

  433. 36:46

    So otherwise you don't have much real estate

  434. 36:47

    ... is if you go to the homepage of the Azure portal, you can also do this from the Azure CLI if you prefer. You can have a look at the lists-

  435. 36:56

    Easier

  436. 36:56

    ... of resources we deployed into this single resour- resource group for you called Contoso Chat RG.

  437. 37:05

    And you can see that there's 11 resources that we have launched for you. What that script done was-- did was to populate some data into Cosmos DB and to Azure AI Search.

  438. 37:17

    Let's have a look first at, uh, Cosmos DB, which is the Asma- Azure Cosmos DB account.

  439. 37:25

    And here in the portal, I can go into and have a look at the data explorer.

  440. 37:38

    Okay, there's a video to watch if you're interested. I'm not.

  441. 37:42

    Okay. And what you can see here is now within our Contoso Outdoor Cosmos database, we have a customer's database, and you can actually drill in there if you're familiar with using databases and have a look at the tables.

  442. 37:54

    There's about 12 customers, uh, in the, in the table and all of their, um, purchase history. Likewise, if I go back to the resource group and have a look at the Azure AI Search resource, which is called Search Service, [clears throat]

  443. 38:12

    I can go to the indexes in the sidebar

  444. 38:18

    under Search management. By the way, if you can't see this sidebar, it happens if you're using a very small laptop screen or if you've got a large font. It might be hidden behind this hamburger menu here in the corner.

  445. 38:34

    Go to Indexes. Oh, that's a problem. I've got no indexes found, so I must have missed a step. Hopefully, you've got indexes there, and, uh, if not, we'll come back and check.

  446. 38:47

    I don't think I went through all the steps to actually pre-provision things there, so we'll go back and have a look at that again. Okay. I just want to give you a preview of where you're coming to.

  447. 38:55

    Again, always any questions, just, uh, pop your hands up.

  448. 39:04

    No, that one is not necessary because [inaudible] that one.

  449. 39:10

    Hmm.

  450. 39:15

    We got a bunch of errors.

  451. 39:16

    Did I? Yeah, probably. I skipped ahead a bunch of stuff. Oh, there we go, a bunch of it. Did you get errors too?

  452. 39:22

    Yeah, I got errors as well.

  453. 39:23

    Oh, okay.

  454. 39:23

    Permission errors.

  455. 39:25

    Hmm.

  456. 39:26

    On the [inaudible].

  457. 39:29

    Hmm. Let's, let me see if I can figure what's up, what's going on here.

  458. 39:32

    Yeah, yeah. [crosstalking]

  459. 39:34

    Okay.

  460. 39:35

    Right?

  461. 39:35

    Yes.

  462. 39:37

    Um, [inaudible]

  463. 39:46

    Hmm. I wonder why that happened. Oh, you know what? I know. I skipped a step.

  464. 39:58

    I know which step I skipped.

  465. 39:59

    We have one good question from the audience.

  466. 40:02

    Yes. What's the question?

  467. 40:03

    Um, so I was just looking at the, the prompt flow example on step-- sort of end of step seven-

  468. 40:09

    Yep

  469. 40:09

    ... where we're looking at the, uh, the graph of the preexisting prompt flow-

  470. 40:13

    Yeah

  471. 40:13

    ... that's been created. One of the questions that I had is, can the graph also be, be cyclical?

  472. 40:18

    Can the graph be cyclical? Um, no. It's a directed acyclical graph, and that is-- 'cause I don't think there's any support for any kind of looping like that, so yeah.

  473. 40:32

    Is there a reason why-- Is there a reason why you'd want it to be cyclical?

  474. 40:35

    Uh, like-

  475. 40:36

    Yeah

  476. 40:37

    ... for example, react agents.

  477. 40:38

    Yeah. Um, well, like within a single component of that prompt flow is anything you want. It could be any Python code. So if you need to do that kind of interactivity within one of those nodes, you can do that.

  478. 40:51

    But the flow itself is, um, is, is acyclic. [clears throat]

  479. 40:56

    Thank you very much.

  480. 40:56

    Yep.

  481. 40:59

    Um, it worked on the second time.

  482. 41:02

    Yeah, I realized the step that I skipped because I was hitting ahead was I didn't do the bit where we create the .env file. So let me check. Yeah.

  483. 41:12

    Dude, I'm trying to remember which step that was. Done that one.

  484. 41:23

    Mm-hmm. Ah, this is the one I didn't do. Yeah.

  485. 41:32

    Yeah, I just skipped this step. Oops. Clicked too late. [keyboard clicking]

  486. 41:55

    Mm-hmm.

  487. 41:57

    We had to run the post-process, uh, script twice.

  488. 42:00

    Okay. [inaudible] Yeah, wait, did I log into this one?

  489. 42:24

    Just make sure I did. [inaudible][clears throat]

  490. 43:18

    We have another good question here.

  491. 43:21

    Sure, go ahead.

  492. 43:22

    So ag- again, about the, um, the, the prompt flow DAG. So this time, uh, you know, there's a lot of power in the DAG description that we can-

  493. 43:30

    Uh-huh

  494. 43:30

    ... that we can build. Um, so one thing that we'd like to explore, that we're currently exploring with a couple of our clients, is how to sort of hit these kinds of DAGs from Enterprise strategy PT-

  495. 43:42

    Mm-hmm

  496. 43:42

    ... to use the, uh, the actions framework-

  497. 43:44

    Yep

  498. 43:44

    ... and then go to a middleware layer where a DAG resides.

  499. 43:47

    Okay.

  500. 43:47

    So, um, what I'd like to understand is what are our options to expose this as an endpoint to external sources other than-

  501. 43:55

    Yeah

  502. 43:56

    ... you know, like our web app?

  503. 43:56

    So that's actually gonna be about step 11 or 12. We're actually gonna deploy it as an endpoint and then connect it to the website. [chuckles]

  504. 45:28

    Say that again.

  505. 45:29

    When you enter a subscription, can-

  506. 45:32

    Let me take a look. All right. Um,

  507. 45:41

    if you just press Enter at this point.

  508. 45:43

    Did I do anything right?

  509. 45:44

    Yep. Yeah, that's just to confirm you wanna log into that particular subscription. Yep.

  510. 46:05

    Okay, I figured out what I did wrong. I've-- I completely skipped step six, which is an important step. Um, you can confirm if you've done step six by having a look in your Explorer for Visual Studio Code.

  511. 46:16

    Once you've done that, you should have a file called .env, and that is where we've set up all of the endpoints and keys that you will need to access the resources that we've provided for you.

  512. 46:28

    And you'll also have a config.json file, um, which does similar kinds of things for AI Studio.

  513. 46:37

    And then once that is all set up,

  514. 46:39

    I should be able to go back and run this pre-provision script.

  515. 46:53

    Is

  516. 47:52

    it a authentication error or something else?

  517. 47:55

    Um, I don't know. Let me have a look.

  518. 47:58

    It looks like it's something else.

  519. 48:00

    Mm-hmm.

  520. 48:00

    Where's the configure?

  521. 48:03

    I might have-

  522. 48:04

    Or is it... Yeah, it's correct.

  523. 48:07

    Yeah.

  524. 48:08

    It is. Are you in the right directory?

  525. 48:13

    Hmm, looks all right. I think it's because I ran the script already once. I'll just have to do it again.

  526. 48:29

    You said the service, service is, uh, VectorDB or?

  527. 48:48

    Yep, that's right. [clears throat] So one of the things we did in that, um, pre-provision script was to take each of the Markdown files which are in the repository. Uh, there's one Markdown file per product that the company sells, and then we just script indexing that into the AI search database, which essentially converts that entire Markdown file into one

  528. 49:08

    point in multidimensional space.

  529. 49:12

    Yeah, chunking it.

  530. 49:12

    Uh, it is actually chunking it, yeah. Um-

  531. 49:16

    Actually, no. Actually, in this example, we don't chunk it just for simplicity. Um, if you actually do it through AI Studio on the search on your own data, it does do the chunking there.

  532. 49:24

    Yeah. Mm-hmm.

  533. 49:28

    Last question.

  534. 49:29

    Yes, question.

  535. 49:31

    Just while we're getting going, um, how does Prompt flow, Autogen, and Semantic Kernel all come together? Are they all competing projects at Microsoft, or are they-

  536. 49:44

    Uh, not so much competing. Um-

  537. 49:46

    They work with each other or-

  538. 49:47

    ... and they have slightly different-

  539. 49:48

    Yeah

  540. 49:48

    ... versions. Let's start with Semantic Kernel and Autogen. Um, both of those are orches- orchestrators. Um, they come from different parts of Microsoft. Uh, Semantic Kernel we kind of view as the enterprise product.

  541. 50:00

    Uh, that's the one that is designed for use in enterprise settings, um, has strict versioning, you know, API changes, all those kinds of stuff. Um, but like LangChain and other orchestrators, you can use it to connect different tools together in different environments.

  542. 50:16

    Autogen serves a different... a, a similar kind of purpose, but it comes out of Microsoft Research, so it's a little bit more cutting-edge. Uh, it's a little bit more flexible, uh, based on a slightly different paradigm, but that's not the one that we recommend for enterprise applications today.

  543. 50:31

    Prompt flow is a different, different beast again. [laughs] Prompt flow is directly within, uh, the AI Studio product and is purely for orchestrating within an endpoint that you deploy through the AI Studio product. [clears throat]

  544. 50:45

    So one endpoint?

  545. 50:47

    One endpoint. The whole purpose is to create one endpoint that goes through, in this case, a RAG process. But it's designed to be more flexible than just for RAG.

  546. 50:54

    And I've-

  547. 50:55

    But you could use Semantic Kernel to have multiple prompt flows?

  548. 50:59

    You could use Semantic Kernel to manage multiple prompt flows as endpoints, so yes. Mm-hmm.

  549. 51:04

    And, and also you can, uh, embed Prompt flow as a library.

  550. 51:09

    Deploying it as an endpoint is one possibility.

  551. 51:12

    Mm-hmm.

  552. 51:12

    Um, you can... it can be used, uh, to write, uh, integration test, evaluation test, uh, embed in your application. Like, if you have a, a Python native application and you want to embed a DAG in it, it's also kind of interesting.

  553. 51:30

    Can I ask-

  554. 51:30

    Is there any like Prompt flow that sits on top of Semantic Kernel so I can have multiple endpoints managed in one graph?

  555. 51:39

    Is there anything like Prompt flow that sits on top of Semantic Kernel to manage multiple endpoints?

  556. 51:43

    I'm looking at the graphic side of the-

  557. 51:45

    Oh, just the graphic side of it

  558. 51:46

    ... the graphic side. Yeah.

  559. 51:47

    Um, well, I didn't... I don't think so. And honestly, when you work with Prompt flow for any length of time, you'll be working with the yam- YAML files. It's nice to have that picture.

  560. 52:01

    It's great for these workshops because I can point to things and show you how things connect, and it's great for debugging because you can actually see how the data is flowing through the thing.

  561. 52:08

    But in terms of an editing environment, that's not really what it's for. Yeah.

  562. 52:12

    We, we also have the Python DSL version now.

  563. 52:16

    Tell us about it. Yeah.

  564. 52:17

    Uh, i- in addition to, um, the graphical YAML-based, uh, DAG, now you can also use Prompt flow in a more programmatic code-first, uh, way, where you write, you, you write, uh, the Prompt flow, uh, configuration in Python instead of, uh, graphically as an, as an alternative.

  565. 52:40

    That's in Prompt flow documentation or-

  566. 52:43

    Yeah.

  567. 52:45

    Thank you.

  568. 52:45

    Yeah. And I was gonna-

  569. 52:46

    Uh, to, to be honest, I'm not sure how

  570. 52:51

    generally available it is. Uh, to [laughs] I'm sorry, to be verified. [laughs]

  571. 52:58

    But ultimately, this is the representation of the Prompt flow. It's a YAML file which just defines each node, uh, with a bunch of tags associated with its inputs and outputs and how it connects to the various endpoints and the types of nodes that we provide, [clears throat] excuse me, in AI Studio.

  572. 53:15

    But if you're running that as, like, a Python library, uh, then at that point it's fairly similar to using LangChain in Python?

  573. 53:23

    Exactly, yeah. Yeah. So there's a command line way to... which is what we're doing in this, this thing, is to run that Prompt flow with a given set of inputs to generate the outputs. [clears throat]

  574. 53:33

    Cedric, are you saying that, uh, there's a Python alternative to the YAML that, that's possible or?

  575. 53:40

    Correct. Uh, which provides a more long chain-like experience.

  576. 53:45

    Right. Right. Right. Cool. And Autogen could-

  577. 53:51

    Autogen could, in theory, do similar stuff, but, uh, it's more of a labs thing, not, not ready for a prime time.

  578. 53:57

    Yeah, like David was saying, um, the team at Microsoft, uh, works on both projects are very different, and they have very different goals. Um, um, Prompt flow, uh, and, um...

  579. 54:11

    and I'm blanking. Semantic Kernel are very product-oriented, uh, so they follow very strict, uh, software release life cycle. Whereas, um, the, uh, the other team, the Autogen team is really research.

  580. 54:26

    So they, they try the latest cutting-edge AI, uh, things. Um, they also use that for papers. So if you use, um, if you use Autogen, you're literally on the bleeding edge, and, uh, things might break and things are experimental.

  581. 54:46

    So you might wanna do that. It might work for your application, but you need to know where you're going into. [laughs]

  582. 54:54

    Yeah. Yeah. We, we had used it for a hackathon, and we were able to put together, like, an agentic-type flow really quickly with it, so we liked it. But, uh, it's good to know that we need to be a little careful with it.

  583. 55:06

    Yeah, definitely. [laughs]

  584. 55:08

    Okay.

  585. 55:14

    Yes.

  586. 55:14

    I'll, I'll project and if you do think you can do it. So, uh, if, if I understand correctly then the... If I could play back and, and make sure that we're, we're understanding, um-

  587. 55:24

    Stop. Thank you. So, um, Semantic Kernel is more for the orchestration, and this could be more at like a, a, a closer to the app layer where I'm building a back end where I'd be using Semantic Kernel to orchestrate the LLM and the, the calls, uh, on that side.

  588. 55:42

    Prompt flow is more on, I want to build a DAG, manage its life cycle, potentially evaluation. Hopefully we'll get to that in, in-

  589. 55:50

    Yep

  590. 55:51

    ... the, the workfl- the workshop. Um, and then deploy it as one endpoint.

  591. 55:54

    Exactly.

  592. 55:55

    Is that the right way to look at things?

  593. 55:57

    Sure. Mm-hmm.

  594. 55:57

    And like I was saying, you can also-

  595. 55:59

    Mm-hmm

  596. 56:00

    ... embed a prompt flow as a framework, as a library.

  597. 56:05

    You don't have to, uh, deploy it as an endpoint.

  598. 56:09

    But that's the use case we'll be using in this one-

  599. 56:11

    Yeah

  600. 56:11

    ... is as specifically as an endpoint. Yep.

  601. 56:15

    And, um, to make things even more confusing- [laughs] [laughs]

  602. 56:20

    Uh, s-

  603. 56:22

    You were right, Cedric.

  604. 56:23

    Sorry. Semantic Kernel, you can also build agentic applications with it. It also has, uh, some of AutoGen's use cases. Simpler ones, but you can build, uh, agentic applications with it.

  605. 56:40

    Do you have a question?

  606. 56:42

    No, I just wanted to make sure.

  607. 56:43

    Can you give me an example of why I'd want to use, um, the Python library? Why I wouldn't want to use, um, Prompt flow in Python slides?

  608. 56:53

    Uh, personally, I like to use it that way, um, uh, because you're in development. Uh, it makes things, um, easier to write tests, uh, to write, uh, snippets, uh, of code that you can reuse.

  609. 57:06

    Um, so this is very convenient for that.

  610. 57:09

    Yeah.

  611. 57:09

    So it's more like a development convenience rather than anything else.

  612. 57:14

    Yeah.

  613. 57:18

    But if you wanted to build, I don't know, like a, a rich Python desktop app where you happen to be wanting to orchestrate LLMs,

  614. 57:28

    that would also be a relevant use case.

  615. 57:31

    Right. Right.

  616. 57:32

    In terms of the different provision, like, that's still...

  617. 57:37

    Sorry.

  618. 57:37

    Are you able to build on top of that? So when would we build for Semantic Kernel versus when would we build for Prompt flow?

  619. 57:44

    And, um-

  620. 57:45

    Are you able to see the data in the database?

  621. 57:47

    I would say-

  622. 57:48

    I can jump in there, Cedric. Um, Prompt flow you'll find useful when you're managing all your resource within Azure AI Studio. I haven't actually got to show you that yet.

  623. 57:56

    We'll do that in a minute. Uh, but within Azure AI Studio, you can launch... Uh, you can create endpoints for OpenAI models, Mistral models, Llama models, anything from the model catalog.

  624. 58:07

    You can create connections into other Azure resources like Cosmos DB and AI Studio, AI Search rather.

  625. 58:14

    Um, [clears throat] you can hook it up to our evaluation framework. We'll be seeing that a little bit later on. All the features in this for managing the entire life cycle of the endpoint itself and all the resources that are required to make that endpoint search.

  626. 58:28

    That's a situation where you'll mainly be using Prompt flow, is in basically managing that connection between all those resources [clears throat]

  627. 58:37

    with the goal of creating an endpoint. You'll be, in our example here, um, calling from a web app, just regular React, you know, um, API endpoint. But likewise, you wanted to call that same endpoint from within Semantic Kernel or any other orchestrator, you could do that same thing there as well. [clears throat]

  628. 58:56

    And if I may add, um, uh, what I like also about Prompt flow is that when you go to AI Studio, because it's very, very well integrated in AI Studio, so you have the visual aspect of it.

  629. 59:06

    But if you go to the playground of AI Studio, you can, uh, configure a RAG application, uh, and also a bit of a code interpreter, uh, application and export it to Prompt flow.

  630. 59:21

    So, uh, instead of having to come up with a whole design of a Prompt flow-based, uh, application, you can just use the playground, configure it as you want, and export it and have a ready-to-use Prompt flow application that you can reuse and embed in your project.

  631. 59:38

    So you also have a workflow that goes from the UI to the code that way,

  632. 59:44

    which you don't with, uh, Semantic Kernel.

  633. 59:47

    There it is.

  634. 59:48

    Okay. That wasn't very clear. No, that's okay. Uh, Google is worse, so. [laughs]

  635. 59:54

    So are you-

  636. 59:55

    If you'd like to play around with that side of it, it's not part of the workshop, but feel free because you have an instance of Azure AI Studio running already in your virtual machine.

  637. 1:00:05

    Um, you will-- You go through the steps of selecting the one project we have here, which is called contoso-chat-asf-ai-proj. One of the things you might want to play around is the playground.

  638. 1:00:16

    Uh, we provided you with GPT-4 and, um, GPT-35 Turbo endpoints, I believe.

  639. 1:00:22

    And this is the place where you can interface through a playground to test out the endpoints you've created. Same place as when you get into the Prompt flow thing a little bit later on.

  640. 1:00:31

    You can then test out the connections between those endpoints and the databases and everything else you've built your RAG application around.

  641. 1:00:38

    Unless it's somewhere in the UI that I'm missing.

  642. 1:00:40

    From there, you have the Prompt flow button where you can export to Prompt flow. That's what I was talking about.

  643. 1:00:51

    Can I keep asking questions?

  644. 1:00:52

    Yeah, that's why I thought nothing-

  645. 1:00:53

    Yeah, absolutely. That's what we're here for. [laughs]

  646. 1:00:56

    So, uh, Copilot Studio seems to be-

  647. 1:00:59

    Oh, so I anticipated something.

  648. 1:01:01

    Getting upgraded and everything.

  649. 1:01:01

    Yep.

  650. 1:01:01

    Three months or so from here.

  651. 1:01:03

    Mm-hmm.

  652. 1:01:04

    Um-

  653. 1:01:05

    It seems everything else-

  654. 1:01:06

    Where do you s- where do you start and stop in Copilot Studio when you start-

  655. 1:01:11

    There. There it is

  656. 1:01:12

    ... either of these?

  657. 1:01:13

    Yeah. So Azure AI Studio is for mainly the purpose of creating an endpoint that works against an LLM and evaluating those endpoints. They're the main two, two use cases of AI Studio. [clears throat]

  658. 1:01:27

    Copilot Studio by hand is about building applications, not endpoints. Um, either building a complete application like a chatbot application or any kind of user interface that has an LLM element, or of integrating those types of applications into applications like Teams.

  659. 1:01:44

    So Copilot Studio is for building entire apps. AI Studio is just for building endpoints.

  660. 1:01:56

    And as you might have guessed, Copilot Studio was built with AI Studio. [laughs]

  661. 1:02:03

    In terms of, in terms of the integration patterns with both-

  662. 1:02:06

    Yeah

  663. 1:02:06

    ... we were getting quite a bit of the same questions and exploring that space.

  664. 1:02:11

    Mm-hmm.

  665. 1:02:11

    Uh, will you see it then that we build an application in Copilot Studio and then call endpoints that we build in, uh, Azure AI Studio?

  666. 1:02:22

    If, if you want that level of customization. You don't need to with Copilot Studio. It's designed in such a way you can build complete apps without having to customize your own endpoints, but if that's a, the position you're in for a particular use case, then yes, absolutely.

  667. 1:02:35

    You can call any endpoint from Copilot Studio, including ones created with AI Studio. [clears throat]

  668. 1:02:40

    As of yet, one of the reasons would be you can't evaluate stuff in Copilot Studio.

  669. 1:02:44

    I don't believe-- I'm not so familiar with that product, but I don't believe they have evaluation features.

  670. 1:02:49

    So if you wanted to do that-

  671. 1:02:51

    Yeah

  672. 1:02:51

    ... you'd have to start in Copilot and then go to endpoints.

  673. 1:02:54

    Yeah, you would evaluate the, the LLM through its endpoint here in AI Studio. Yep. We'll come to that in probably about twenty minutes or so.

  674. 1:03:01

    Yeah.

  675. 1:03:02

    So the documentation.

  676. 1:03:02

    Thank you.

  677. 1:03:03

    Yep, welcome.

  678. 1:03:04

    Very well.

  679. 1:03:07

    Mm.

  680. 1:03:07

    It's all a little confusing when you saw the website.

  681. 1:03:10

    Yeah. But here is the very simplest prompt flow I just generated from the chat playground. All it does is take an input,

  682. 1:03:18

    passes it through to the LLM called chat there, and then generates the output directly back into the playground. So that's a completely unfiltered AI endpoint. Whereas if you use an application like ChatGPT, there's a whole much more going on with your prompt and the context and everything else in a more RAG style than just passing it directly

  683. 1:03:38

    to an endpoint. [clears throat] I'll keep on going to catch up to where you all are, so I can demo the good bits when we get to the end.

  684. 1:03:53

    You're very welcome.

  685. 1:03:57

    That one.

  686. 1:03:58

    Yep.

  687. 1:04:00

    That's where I would just [laughs] put the documentation to make it a little bit clearer. Okay, cool. Um, and then let's see. What are we seeding into?

  688. 1:04:11

    Seeding into, yeah, provisions.

  689. 1:04:14

    Yep, yep.

  690. 1:04:14

    Provision. Yeah, so you're just missing that in the documentation. It makes it clearer, that's all.

  691. 1:04:20

    Okay.

  692. 1:04:20

    Oh, it couldn't do it. It was doing this.

  693. 1:04:23

    Yeah, so you're already there.

  694. 1:04:25

    It says no provider.

  695. 1:04:27

    Oh, 'cause you're already there. Doing provisions.

  696. 1:04:28

    Oh, no, I am not.

  697. 1:04:29

    Yep, yep.

  698. 1:04:29

    Oh, I already seeded it.

  699. 1:04:30

    Yep.

  700. 1:04:30

    So just set up.

  701. 1:04:31

    Yep. And then it's gonna execute to set up the database and the index.

  702. 1:04:47

    Yep. So it's gonna set up the database with all the data-

  703. 1:04:51

    Okay

  704. 1:04:51

    ... and then the index for the AI search.

  705. 1:04:53

    Got it.

  706. 1:04:56

    Yep. Let me know if you run into any other-

  707. 1:05:27

    Sure.

  708. 1:05:31

    Yeah, just one sec.

  709. 1:05:35

    Are we able to use another vector database besides Cosmos DB?

  710. 1:05:36

    Uh, the question was, are we able to use another vector database besides Cosmos DB? And, yeah.

  711. 1:05:40

    Or do you know, just like we mentioned earlier, we could use like Chroma or something like that?

  712. 1:05:44

    Yeah, absolutely we can. Um, I'll-- when I-- when we get the prompt flow up, I'll show you that. But if we go into the prompt flow, um, node, which is, I think is the retrieve product information node.

  713. 1:05:54

    You can see it's set up directly as a, as a connection to, uh, AI search, but you could set it up as a connection to other vector databases as well, or even just an endpoint. [clears throat]

  714. 1:06:03

    I have a very foolish question.

  715. 1:06:04

    Okay.

  716. 1:06:04

    What do you think is the biggest differ-- oh.

  717. 1:06:06

    Hmm.

  718. 1:06:06

    What do you think is the biggest differentiator between AWS and Microsoft with-

  719. 1:06:09

    What do you think is the-- what do I think is the biggest differentiator between AWS and Microsoft?

  720. 1:06:12

    Like, why should I go with using Microsoft for startups versus AWS for startups? If that makes-

  721. 1:06:18

    That's a different question, because I was gonna go, like, the, the main differentiator between, uh, Azure and AWS is the enterprise scale. You know, Azure is a database that's used by big companies like Microsoft and so forth, and is designed with a lot of features in that around authentication, security, scaling, monitoring, that big companies that run production

  722. 1:06:38

    apps really need. Um, that actually makes it a little more difficult for startups, honestly. Um, because the first thing you do when you start working with Azure is start dealing with things like resource groups and, and security.

  723. 1:06:49

    Uh, what's the... You know the more about this than I do. We have to deal with that a lot. Um, but that's why we set up Microsoft for Startups, is to help startups get into the process of understanding the different kind of process of working with resources in Azure, which is quite different from AWS, in the sense

  724. 1:07:05

    of you can just spin up a single VM in, in AWS, and you're done. Microsoft, when you spin up a VM, there's actually six different resources that get created because it's there to support the enterprise use case, as opposed to just, I want a single VM.

  725. 1:07:20

    Yep. And-

  726. 1:07:22

    I will say Microsoft for Startups, they have a pretty cool program, the Microsoft for Startups Hub. They can give you up to $1,000... $100,000 in credit for Azure. It's like different tiers, so you know, level one they give you this many credits, level two this give you this many credits.

  727. 1:07:37

    So you really get a lot of credits for developing your first startup application, and they also have a pretty good ecosystem. They have, like, different mentors that you can get, get pair up, pair up with.

  728. 1:07:47

    They have, like, different sessions that you can attend to learn, like, how do I use this tool? How do I, like, implement this technology?

  729. 1:07:52

    Mm.

  730. 1:07:52

    So they give you the credits and a lot of the guidance to make that happen.

  731. 1:07:56

    Yeah. And the whole Microsoft for Startups team is here at the booth in, in, um, salon nine. So chat to them about the Startups Hub program, and they can get you in at the right level.

  732. 1:08:06

    Yeah.

  733. 1:08:06

    Yeah.

  734. 1:08:06

    It was in AWS until

  735. 1:08:07

    Mm-hmm.

  736. 1:08:09

    Yeah. [laughs]

  737. 1:08:28

    Everyone does. [laughs] Yeah.

  738. 1:08:38

    You have a question? [whispers]

  739. 1:09:16

    You know, the second label that was delivered, I think that there are equations that are being around the data maps and the CNS, uh, organizations are just speaking out the experiences being, uh,

  740. 1:09:34

    automatic generation before

  741. 1:09:49

    Sure.

  742. 1:09:49

    Um, I don't know. So, so there's a different set of data model. Um, this is based with Azure

  743. 1:10:02

    and, and a DAG that it depends whether there are a bunch of dancing and how are you doing it. Um, instead of being able to have some means to do it, this is your...

  744. 1:10:14

    You have to put the encryption down.

  745. 1:10:17

    So that script loads the search service index?

  746. 1:10:21

    That's correct, yeah. Yeah, I just start over because I ran that script before [laughs] I did all the configuration.

  747. 1:10:29

    Um, basically when you say-

  748. 1:10:52

    There we go.

  749. 1:10:55

    Also, I think I saved the-

  750. 1:11:10

    Yeah. So c- no, don't do that, actually. Ca- can you click on text embedding?

  751. 1:11:16

    One search should do it or-

  752. 1:11:17

    Yeah, it is connected. So yeah, I don't... I'm not sure. Let me talk to David and ask him.

  753. 1:11:23

    But-

  754. 1:11:24

    No. From that it's gonna get the embeds, and from those embeds, you know, you're gonna create your snippets for the embedding to get out. So you have the embedding some data stuff for your document or your-

  755. 1:11:34

    Can you click the deployment?

  756. 1:11:35

    Sure.

  757. 1:11:36

    Yeah.

  758. 1:11:37

    And then from those embeds, the embedding is used to find those embeds.

  759. 1:11:41

    But what's, what's RAG relevance is creating that?

  760. 1:11:43

    If you wanted to store chat history, would you store it in this document?

  761. 1:11:49

    During the talk, do you want to add on-

  762. 1:11:49

    Um, you could, absolutely. Yeah. The way Prompt flow is set up is it actually manages the chat history for you as it comes through the flow.

  763. 1:11:56

    Okay.

  764. 1:11:56

    But if you want to store it beyond that interaction, then yeah, you can store it in the database.

  765. 1:12:00

    That's step three was, right?

  766. 1:12:03

    Um, there might be a bug with the documentation-

  767. 1:12:05

    Okay

  768. 1:12:05

    ... that we missed.

  769. 1:12:05

    That's okay.

  770. 1:12:06

    We are also gonna go in as, uh, inputs to the actual prompt going to the LLM.

  771. 1:12:10

    Yeah.

  772. 1:12:10

    Mm-hmm.

  773. 1:12:10

    Some of, some of them are.

  774. 1:12:11

    Yeah. Yeah.

  775. 1:12:12

    Okay. Yep.

  776. 1:12:12

    In this case, you may ask about, like, products or, like, a specific item, so it's gonna search for those.

  777. 1:12:18

    Here we go.

  778. 1:12:18

    Mm-hmm.

  779. 1:12:18

    Then it's gonna put those as a part of the, the workflow-

  780. 1:12:21

    Right

  781. 1:12:21

    ... for the LLM to like use.

  782. 1:12:22

    Uh, connection default for zero point AI.

  783. 1:12:24

    Mm-hmm.

  784. 1:12:24

    Actually, it's AOAI-

  785. 1:12:27

    Oh, okay

  786. 1:12:27

    ... something.

  787. 1:12:28

    Oh, okay.

  788. 1:12:29

    And by default it is correctly set.

  789. 1:12:31

    Okay.

  790. 1:12:32

    I missed it when I went through it.

  791. 1:12:34

    Okay.

  792. 1:12:34

    But, uh-

  793. 1:12:35

    When-

  794. 1:12:35

    ... a few, a few of them, uh, were... And what's confusing is that in addition to the A- AI, Azure OpenAI-

  795. 1:12:44

    Mm-hmm

  796. 1:12:44

    ... AOAI-

  797. 1:12:45

    Yeah

  798. 1:12:45

    ... connection, you do have a default connection.

  799. 1:12:47

    Right. Yeah. There's two sets of resources.

  800. 1:12:49

    And so when you switch to default, actually you break things.

  801. 1:12:52

    Oh, okay.

  802. 1:12:53

    So [laughs]

  803. 1:12:54

    Oh, okay. All right. When, when folks get to that, I'll have you show them.

  804. 1:12:56

    This thing.

  805. 1:12:56

    Yep. Those of you who haven't done step eight yet, the custom connection for Cosmos DB, I'm gonna go through that now.

  806. 1:13:15

    That is a big step. Is that difference that it allows you to... You don't have to publish the-

  807. 1:13:23

    Go to the connected resources and view them all.

  808. 1:13:27

    And then create a new connection So this is one of the things that the Prompt flow does for you, is it has, um, standardized connections to all these tools.

  809. 1:13:40

    But if you wanna connect to any other service, you can use a custom connection, and that in fact is what we're gonna be doing here.

  810. 1:13:47

    It looks like we... Uh, Cedric, looks like we have a question over the back row, too.

  811. 1:13:50

    Oh.

  812. 1:13:50

    Yep. We're gonna add the key value pairs for our connection to Cosmos DB.

  813. 1:14:04

    So, so, um, uh, like-

  814. 1:14:07

    Yeah. Slow

  815. 1:14:09

    ... from, um, ANN, ANN-

  816. 1:14:11

    I'm gonna grab that value from our .env file

  817. 1:14:15

    ... there's indexing strategies like ANN.

  818. 1:14:17

    Yeah.

  819. 1:14:17

    Um, so, so just depending-

  820. 1:14:19

    There is some pieces-

  821. 1:14:21

    Okay. Yeah, yeah.

  822. 1:14:23

    Veterans.

  823. 1:14:23

    Yeah, yeah.

  824. 1:14:23

    Way worse. Is that- We have another good question here.

  825. 1:14:29

    Yeah, what is it?

  826. 1:14:31

    Uh, how does the vector, uh, search perform at scale? So if you have a million vectors and you wanna perform a nearest neighbor search against a query vector-

  827. 1:14:39

    Yep

  828. 1:14:39

    ... uh, what are some of the, you know, the strategy, strategies to prune that? And then how does that impact recall as well?

  829. 1:14:46

    Yeah. Well, first of all, that's the entire reason why vector databases existed, is to exactly do that search quickly. They're indexed in such a way that they can do that nearest neighbor connection at scale and at speed.

  830. 1:14:59

    So you can certainly do that. It's not very difficult to do yourself, to do, do an embedding for a bunch of documents or a bunch of chunks of documents, and then do a nearest neighbor algorithm to find what is closest to the embedding for the customer's question.

  831. 1:15:13

    But the advantage of doing it within a vector database is you can do that quickly and at scale. And what was the second part of your question?

  832. 1:15:18

    Uh, yes.

  833. 1:15:19

    Yeah.

  834. 1:15:19

    There's different nearest neighbor algorithms, right?

  835. 1:15:21

    Yeah.

  836. 1:15:21

    There's, uh, like ANN as well.

  837. 1:15:23

    Yep.

  838. 1:15:23

    And then you have to... I think there's parameters that you have to tune.

  839. 1:15:27

    Mm-hmm.

  840. 1:15:27

    Um, that affects recall as well, depending on how f- how far you want the tree search to occur.

  841. 1:15:32

    Yep.

  842. 1:15:32

    Um, and so what are, what are some strategies to balancing that, uh, and getting the highest recall instead of, you know, just doing-

  843. 1:15:38

    Yeah

  844. 1:15:38

    ... brute force? 'Cause brute force-

  845. 1:15:40

    Mm-hmm

  846. 1:15:40

    ... I'm assuming, you know-

  847. 1:15:41

    Yeah

  848. 1:15:41

    ... takes, takes a while at that scale.

  849. 1:15:43

    So interestingly, and this might not be the answer you expect, is that at least in our experience working with real world applications, um, vector search by itself actually isn't enough, um, despite playing around with the algorithms and choosing the parameters so you can expand out the search and not miss documents.

  850. 1:16:01

    What we've actually found is actually a combination of keyword search and vector search together, um, actually outperforms either. Um, and that's a feature that's built straight into Azure AI Search.

  851. 1:16:10

    It by de- I think by default these days, it, it actually defaults to a, a hybrid search. In this particular example, just to keep things simple, we're just doing a vector search.

  852. 1:16:19

    But for practical applications, we actually recommend a combo of keyword and vector search.

  853. 1:16:26

    And, uh, I'm assuming that metadata that you can search, like for example, keywords can be updated for a product, in this case, at any point in time for any, any vector, right?

  854. 1:16:35

    Yeah, that's right.

  855. 1:16:36

    Um, and then, uh, is there like a... The last question, is there a upper limit on the dimension size of the embeddings? [sighs]

  856. 1:16:44

    I'm sure there is an upper limit somewhere, but we haven't come across it for AI Search at least. Yeah, I don't think there's one sort of in there by design. [laughs]

  857. 1:17:08

    Because this is supposed to be, um, your creative component.

  858. 1:17:14

    So, so the copy and paste it.

  859. 1:17:16

    Yeah, the copy paste. Okay. I have this in my-

  860. 1:17:19

    One more question.

  861. 1:17:20

    Yeah.

  862. 1:17:21

    Okay. Um, so if you, if you verify-

  863. 1:17:25

    What about using Entra ID with this, um, certainly external Entra ID, if you're trying to develop this for an externally facing application-

  864. 1:17:34

    Yeah

  865. 1:17:34

    ... multi-tenanted-

  866. 1:17:35

    Mm-hmm

  867. 1:17:35

    ... and everybody having their own little space, shall we say?

  868. 1:17:38

    Yeah.

  869. 1:17:38

    Where every, um...

  870. 1:17:39

    As you have it, I think-

  871. 1:17:40

    Is there any sort of frameworks or, uh, s- things like this that are set up so to show how to sort of set that up across all levels?

  872. 1:17:51

    Mm-hmm.

  873. 1:17:51

    'Cause you're sort of across... Like if you are, if you're going from what I'm talking about, you're talking about-

  874. 1:17:55

    Yeah

  875. 1:17:55

    ... Copilot Studio into this-

  876. 1:17:57

    Mm-hmm

  877. 1:17:57

    ... into the AI Search-

  878. 1:17:58

    Yeah

  879. 1:17:58

    ... into Cosmos DB. Is there any sort of... Is there anything written down anywhere? I haven't been able to find it, basically.

  880. 1:18:03

    Okay. That's an area I'm not an expert in, but Cedric or Miguel might be. Either of you two? [laughs]

  881. 1:18:07

    So-

  882. 1:18:07

    Entra ID

  883. 1:18:08

    So, but in ext- using external Entra IDs, I know it's only come out a couple of months ago.

  884. 1:18:12

    Yeah.

  885. 1:18:12

    But being able to create a scope across all of those systems-

  886. 1:18:19

    Yeah

  887. 1:18:20

    ... so that the, the scope that you can see is for that Entra, external Entra ID.

  888. 1:18:24

    Yeah.

  889. 1:18:24

    Is there anything written down to show how to do that across each of the individual systems in the right way, or is there anybody working on that, or is...

  890. 1:18:33

    Do you know?

  891. 1:18:35

    You're... When you, you said Entra ID?

  892. 1:18:37

    Entra.

  893. 1:18:37

    So from an authentication standpoint?

  894. 1:18:40

    So, so say again.

  895. 1:18:41

    Yeah, you're, you're talking about the Entra authentication mechanism?

  896. 1:18:44

    Yeah, Entra and, yeah.

  897. 1:18:45

    Okay.

  898. 1:18:45

    Specifically the external ID, so using external IDs, so social logins and things like that. So if you're facing this to the outside world, enabling external users, external IDs to actually get scope across the whole thing.

  899. 1:18:59

    Yeah, no. It's, um, to be honest, it's not a domain where I've, uh, spent much time.

  900. 1:19:03

    Got it.

  901. 1:19:04

    Yeah, so I'm not gonna be able to talk much about it.

  902. 1:19:06

    I can, I can add a little bit of color.

  903. 1:19:07

    Oh.

  904. 1:19:07

    You, you wouldn't have external users log into AI Studio. Now, that's, that's for the developers-

  905. 1:19:12

    Oh

  906. 1:19:12

    ... and for the IT managers. Um, but what you are exposing to the app, which the end users are then accessing, is those endpoints. And you can manage those endpoints either by tokens or by managed identity.

  907. 1:19:25

    So whatever way that you want the app to talk to the endpoint based on that identity controls obviously what the endpoint is then able to do. Yep.

  908. 1:19:34

    And then all the way through.

  909. 1:19:35

    Exactly, yep.

  910. 1:19:40

    And then this, at the moment, on the... If I'm reading this correctly, the retrieval document, retrieve- retrieved documentation from AI search is retrieving the vector, and then it's handing it off to the, the, the actual document itself.

  911. 1:19:56

    So it's only the vector that's getting passed, not the document itself?

  912. 1:19:58

    No. Well, the vector is, is, um, is then used to retrieve the document, in this case a markdown file, and then the markdown file is actually gets inserted into the prompt so that the LLM can see it, so that OpenAI can see it, and form its answers based on that information.

  913. 1:20:13

    Okay. And so that chunk is being sent?

  914. 1:20:15

    Exactly, yeah, and I'll show you how that gets put together in a sec.

  915. 1:20:17

    Okay. Cool.

  916. 1:20:19

    There is one file there that has the prompt, and then they have the variable for the document. So that variable simply gets replaced by the text that it grabs from the search.

  917. 1:20:34

    So you click, uh, you open the project. Yes. Then you expand the sidebar, and then you're gonna see Prompt flow at the bottom.

  918. 1:20:46

    So just on, uh, to, to continue building on the Entra ID piece, because, uh, one thing that we're looking into is, for example, if we want to do granular access control and make sure that we don't pass to our prompt flow, you know, uh, the ability for it to search any database and retrieve any data from a

  919. 1:21:04

    user-

  920. 1:21:05

    Mm-hmm

  921. 1:21:05

    ... just to make sure that a nefarious actor might not be able to get data from other users by something, through something like prompt injection.

  922. 1:21:13

    Yeah.

  923. 1:21:13

    I noticed that in the authentication type for Prompt flow, we're able to use token-based, so Azure AD tokens. Let's say that we do have a user log into our app, pass through that token to call the Prompt flow endpoint.

  924. 1:21:30

    Can we then use that token authentication of the user to call subsequent, um, components in, of our Azure stack?

  925. 1:21:40

    So to pass through that user's delegated token and just make sure that we're retrieving data for that user.

  926. 1:21:47

    I don't think so. Go ahead, Cedric.

  927. 1:21:49

    I, I was gonna s- I was gonna say yes.

  928. 1:21:53

    Okay.

  929. 1:21:53

    But, um-

  930. 1:21:54

    I'm not a security expert. These guys know a lot more than, than I do. [laughs]

  931. 1:21:57

    So the, the only thing is that, um, I don't know much about Entra, so any Entra-specific things, I don't know much. But, um, 'cause I joined Microsoft not so long ago.

  932. 1:22:09

    But when it comes to, like, normal, uh, security, you could use, uh, any kind of tokens. Let's say Jo- JWT, JWT token, and then, uh, like you would, you know, encode a JWT token for any kind of REST API, uh, you can pass, um, the JWT token to the Prompt flow, um, endpoint and use it inside the

  933. 1:22:34

    Prompt flow definition to pass on to whatever internal service you w- you would like to.

  934. 1:22:39

    So I, I think a, a very concrete example of that would be, let's say we have the Cosmos DB endpoint, uh, and I, I want to ensure that I can only access the specific user data in that Cosmos DB.

  935. 1:22:53

    So I want to actually use like row-based access control, where that user's only allowed to see certain rows in that.

  936. 1:23:00

    Oh, yeah.

  937. 1:23:00

    Yeah.

  938. 1:23:01

    I can, I can actually show you something to do with that right here now that I understand. So the, what we have here is the simple prompt flow for this particular application, where I do the input.

  939. 1:23:12

    There's the embedding to retrieve the product documentation. There's nothing secret in that. What's, what's relevant here to your question is... Let me make this a bit bigger so we can see.

  940. 1:23:22

    There we go. And the customer lookup, which is using

  941. 1:23:29

    information that is provided by the customer through them having logged into the website. So for example, when I go to this website...

  942. 1:23:41

    Do I have it open still? Did I close it?

  943. 1:23:47

    Uh, academics/ai-tour/contoso-web. This link is given in one of the last steps of your...

  944. 1:24:01

    ai-tour-t. That's not right. Contoso-chat.

  945. 1:24:04

    So in, in this case, when-

  946. 1:24:06

    Yeah

  947. 1:24:06

    ... when the LLM gets the prompt-

  948. 1:24:09

    Mm-hmm

  949. 1:24:09

    ... it already knows who you are. It already knows you are-

  950. 1:24:12

    Exactly

  951. 1:24:12

    ... David, and it gives to the LLM only David's information. It will be different if the LLM ask, you know, "What's your name?" And you're like, "David. Never mind, I'm Miguel."

  952. 1:24:22

    You know, because then it will be, it could be jailbreak. But in this case, all of that is set up so that when the LLM gets there, it already has your name and it's authenticated, and it has only your information and only your information.

  953. 1:24:35

    And I think that's just about what you were gonna show.

  954. 1:24:37

    Yeah. This is what I was trying to get to. All right. So when we're at this chat application, we can see Sarah Lee is logged in already.

  955. 1:24:43

    If you go to... Show me the-

  956. 1:24:46

    So as we go through the prompt flow-

  957. 1:24:47

    Show me the last-

  958. 1:24:48

    ... one of the inputs there is the customer ID, and that's come from this app through the token that's being provided to the end... Well, actually, no, it's a parameter to the endpoint in this particular case.

  959. 1:24:59

    It's not set up exactly that way. But when we ask the question, "What are-- What did I order last time?"

  960. 1:25:09

    What's important to understand there is that there's nothing in this app that is searching any database. All it's doing is passing the user ID and that question, "What did I order last time?"

  961. 1:25:23

    into that whole prompt flow. And then as part of that prompt flow process, which is a privileged account, it is then query the database with that user ID to get back her list of product purchases.

  962. 1:25:38

    Then the LLM is operating on that information with that question to generate that bit of text that you see, and that bit of text is the only thing that actually goes back to the app.

  963. 1:25:47

    So the app doesn't have any direct access to databases at all. Is that what you're getting at there?

  964. 1:25:52

    I think, I think that is what we're getting at, and I think you hit the nail on the head when you're saying that Prompt flow is, has privileged access to the database.

  965. 1:25:58

    Yeah, yeah.

  966. 1:25:59

    What I'm trying to avoid is for Prompt flow to have privileged access. What I wanted to do-

  967. 1:26:03

    Oh, I see

  968. 1:26:03

    ... is to inherit the access of the calling user through, for example, his AD token.

  969. 1:26:08

    Right. So I'm speculating here, this is not my domain, but I think the way that would work is through the features of the database, where you give that authentication information with your query that prevents you otherwise, then would that prevent what you're trying to do?

  970. 1:26:22

    Yeah.

  971. 1:26:23

    Yeah.

  972. 1:26:23

    Absolutely.

  973. 1:26:23

    Yeah.

  974. 1:26:23

    I'm just wondering whether or not that's already something that you're looking into with, for example, the Prompt flow connection to Cosmos DB.

  975. 1:26:30

    Um, that I'm not sure of. Um, but I'm sure they are [laughs] 'cause that's the whole purpose for this, this thing existing in the first place. This is what, um, Copilot Chat was built on, for example, and sort of that's all based on enterprise logins and s- and, and things like that.

  976. 1:26:50

    There's questions there, yeah.

  977. 1:26:51

    Is there any support for like dynamic control points in the Prompt flow?

  978. 1:26:56

    Like, uh, like, uh, if you want to like have variable loops in the control DAG, uh, is, is that possible?

  979. 1:27:02

    Yeah, the question on that, do you variable control points in the loop? The answer is absolutely. If you run this within Visual Studio Code, you can set break points in the Python code that runs within each of the nodes.

  980. 1:27:12

    Is that the question you're asking?

  981. 1:27:13

    Uh, I guess, I guess like, uh, depending on the results like o- of, of one node, for example-

  982. 1:27:18

    Oh, conditional results

  983. 1:27:19

    ... you might, you might want to route to like different nodes.

  984. 1:27:22

    Yeah, absolutely. 'Cause within... What's actually being run within each of those nodes, let's actually even take a look at that. Um, if I go over to the Prompt flow itself,

  985. 1:27:32

    let's have a look, for example, at the

  986. 1:27:37

    LLM response node, I think. Actually, that's not... That one's not very exciting. Let's have a look at the customer prompt node.

  987. 1:27:45

    So what's absolutely happening at this point is it's got, um...

  988. 1:27:50

    This is actually just, uh, the prompt that actually gets built together. Um, but you can see like it's got this like meta programming language, you know, for item and documentation and so forth.

  989. 1:28:00

    What that's doing through at that point is looping through all the matched products that are related to the user query, extracting them out from Azure AI Search as vectors, then extracting out the markdown files that relate to those vector indices, putting that directly into this prompt.

  990. 1:28:19

    So when I ask the question of the app-

  991. 1:28:24

    Yep. There it is

  992. 1:28:25

    ... you know, "What's a good pair of shoes?"

  993. 1:28:35

    That's not the only bit of text that is going to OpenAI at that point. What is in fact going to OpenAI is a whole bunch of text defined by this customer prompt here, and telling, including telling OpenAI, "You're an AI agent for the Condors, Condoso Outdoor product retailer.

  994. 1:28:52

    You should always reference factual statements. The following documentation should be used in the response," and this is where the individual relevant products are inserted into the prompt. And to our question earlier on for this particular customer, this is where their previous orders are inserted into the prompt.

  995. 1:29:10

    And then finally, the question, you know, "What's a good pair of shoes?" is sent to OpenAI. So it has all that context from that RAG process to formulate a meaningful response based on that particular customer, their purchase history, the question, and the products that are related to that question.

  996. 1:29:29

    But I guess-

  997. 1:29:30

    Yeah

  998. 1:29:30

    ... I guess in this case, like, I guess, uh, the flow is like kind of static.

  999. 1:29:35

    Yeah. I wanted to give you another example. Here's a better example of that kind of thing. In this case, this particular node is just running Python code. So you could put conditionals into that Python code, for example, based on the inputs to do different kinds of things.

  1000. 1:29:47

    Anything you like, in fact.

  1001. 1:29:48

    May, may I add something?

  1002. 1:29:49

    Yeah.

  1003. 1:29:49

    Because I think I know, uh, w- w- the specific that you're, uh, you had in mind. Um, w- what David showed, he showed two things. He showed in the templates, in the template nodes, he showed, uh, looping logic and conditional, but it's looping and conditional string rendering, um, constructs.

  1004. 1:30:15

    But... And in the Python code, you can have any Python. Like you can have conditions, loops, whatever. But to your point, all the nodes in the DAG are going to be executed.

  1005. 1:30:28

    You cannot have conditional node execution. But what you can have is inside a node in the Python code, you can conditionally execute something, but all the nodes are systematically going to be executed.

  1006. 1:30:47

    It is not a business process orchestration system. It is really tailored towards building LLM applications. So it's a sim- it's simplified. It's not, uh, generic.

  1007. 1:31:03

    Does that, does that make sense?

  1008. 1:31:06

    Yeah. How... It, it looked like in that Python node that you had, uh, that's where it was doing the customer lookup.

  1009. 1:31:14

    Um, so how does that tie together? I mean, I see the line connecting it, and then I see the Jinja template for the prompt. Uh, and the Jinja template was iterating over customers and that, you know, for...

  1010. 1:31:27

    Or sorry, ite- iterating over the orders. Uh, so how does that, how does that tie together? Um, so s- uh, where are you at right now?

  1011. 1:31:38

    This one?

  1012. 1:31:39

    Y- yeah, there was like the customer lookup Python that we were just looking at.

  1013. 1:31:43

    Trying to make it bigger so I can see it. [laughs]

  1014. 1:31:44

    Oh, yeah.

  1015. 1:31:44

    What-

  1016. 1:31:44

    The one on the right. Uh, yeah, the customer lookup.

  1017. 1:31:47

    You have one node which queries the database-

  1018. 1:31:51

    Yeah

  1019. 1:31:51

    ... fetches, uh, all the information from the database, stores it into a variable, into the context.

  1020. 1:31:59

    And then the Jinja template uses the previously set collection of results for the rendering.

  1021. 1:32:08

    Right. That makes sense. So when you say it stores it, uh, into it, is that where the response, the orders, uh, on line thirteen, is that doing it or?

  1022. 1:32:17

    Correct.

  1023. 1:32:17

    Okay. And then if we, if we click on the next one down, the customer prompt, and we go to that loop again, there it is. Well, but, uh, okay, customer.orders.

  1024. 1:32:27

    So that's, that's how it ties then, eh? Or-

  1025. 1:32:29

    Correct. Yeah.

  1026. 1:32:30

    Okay. Thank you.

  1027. 1:32:31

    Using input and output bindings on each node. Yeah. And the arrows that are coming into the top of the representation in this graph, those are the inputs, and the arrow coming out of the back is the outputs, and there can be multiple of those. [clears throat]

  1028. 1:33:04

    So can I actually get, uh... So if I do this [inaudible]

  1029. 1:33:15

    Um, no, the, uh... It's a new, uh, representation.

  1030. 1:33:19

    Okay. So I can't actually edit it?

  1031. 1:33:21

    No. You need to click on a node, and edit the target or the edit page.

  1032. 1:33:28

    Yes.

  1033. 1:33:30

    And then it would look at that [inaudible]

  1034. 1:33:33

    Correct. It's one of those, one way.

  1035. 1:33:35

    Got it. [clears throat]

  1036. 1:33:39

    What?

  1037. 1:33:40

    It's called a visual editor, but it's really more of a-

  1038. 1:33:43

    Yes.

  1039. 1:33:45

    Yes. The comment was, "It's, it's called a visual editor, but it's really more of a visual reader," and that is absolutely true. [clears throat]

  1040. 1:33:50

    I want to highlight a little subtlety too, when you get to step ten, um, when you first run your prompt flow in Visual Studio Code, um, you're gonna be r- clicking on the Run button once you've viewed the prompt flow itself in the Visual Studio Code environment.

  1041. 1:34:05

    You can see the commands it's running. It's just running a little Python command to, to launch in the YAML file. But what I want to emphasize here is that in reality, everything here is running locally.

  1042. 1:34:16

    And in fact, in the usual developer environment, it would be running directly on your laptop or, or a shared machine. In this case, it's running on the GitHub Codespaces environment, which is subbing in for your local environment here.

  1043. 1:34:30

    And the whole idea behind this is you have a very fast, responsive place to try out different prompts, to make sure your connections are working, perhaps testing different types of LLMs, replacing them in the LLM steps, so you can actually figure out, you know, what are the bits of the puzzle that go together to give you a

  1044. 1:34:47

    good experience for the endpoint that you're trying to create just in a local environment. Now, I say local because of course the database is still in the cloud, and the OpenAI endpoint is still in the cloud, but all the orchestration is happening directly on your local machine.

  1045. 1:35:01

    Our next step after this is going to be then publish that prompt flow into Azure inside its own container app, as it turns out, and then that's gonna be a hosted, uh, cloud version of that same prompt flow, which is gonna support the production use, uh, of that endpoint in your application. [clears throat]

  1046. 1:35:35

    Yeah. A side effect of what David just said is that because it's building a Docker container, you can actually customize the environment and add packages. So, um, earlier, we were talking about the differences between Cementin Kernel and Prompt flow.

  1047. 1:35:51

    Uh, one of the nice things with Prompt flow is that, um, it's, it's very interesting for like web developers, uh, because they don't have to care about creating an environment, deploying a Docker environment, scaling it.

  1048. 1:36:05

    Uh, the whole, uh, scaling is done on, uh, automatically by the platform. Uh, so you j- you just need to add packages and you, you, you know, uh, so you can combine an LLM with some packages for some specific processing, um, and the whole deployment is, is done, uh, automatically.

  1049. 1:36:23

    So you can focus on the UI and the user experience.

  1050. 1:36:35

    All right. And then let me get to the evaluation.

  1051. 1:36:47

    I

  1052. 1:36:47

    have

  1053. 1:37:42

    a quick question for you. Uh, when we say that, "What else did I purchase?" Does it also fetch like, um, does it also do the commanding or does it skip that?

  1054. 1:37:55

    Because like you technically don't need that, right? Um, like you just need the purchase history for that question.

  1055. 1:38:03

    Um, so in that specific DAG, we will systematically query both the vector database and the, uh, customer information.

  1056. 1:38:19

    So, um, there are, there are two ways to answer that question. Uh, we will still be embedding the question because we still need to query the vector database. But, um, uh, yes, when it comes to, uh, answering the question of, uh...

  1057. 1:38:36

    Can you repeat the question you-

  1058. 1:38:38

    Yeah. It was, uh, what else did I purchase?

  1059. 1:38:40

    Yeah. What else did I purchase? Then, uh, because we will query, um, the, uh, order history from the relational database, the LLM is gonna pay more attention to that part of the context than, uh, the documen- the product documentation side of things.

  1060. 1:38:59

    But it's really... We are really relying at that point on the feature of the LLM to be able to pay attention to what matters.

  1061. 1:39:12

    One more question. So, so, like, if I asked it, like, what else did I purchase? Let's say, what did I purchase in the last month? So it, it... Can it, like, f- uh, form a relational query to-

  1062. 1:39:23

    Not in that specific prompt flow DAG because that specific prompt flow DAG, uh, just returns the, I believe, the last ten orders from the history, is what we do, I think.

  1063. 1:39:38

    Um, and that's it. That's what we put in the context because that, uh, RAG application, you know, is for, uh, workshops and demos. Um, what you're talking about is to do something else, which is text to SQL,

  1064. 1:39:55

    where you take, you take, uh, a query in natural language, you transform it into a SQL query that you execute against a relational database where you have a filter where date is one month ago or whatever.

  1065. 1:40:11

    That's, uh, a similar use case but a different implementation.

  1066. 1:40:18

    And that's also an area to be wary of too because that's an area where prompt injection could come into the fact. If you're cr- forming a SQL query on the basis of user input, you gotta recognize that there might be malicious input in that process, which might generate SQL.

  1067. 1:40:34

    There's still an intermediate step. It's not directly pasting the string into a SQL query, uh, but there still is an opportunity there for, uh, bad actors to control what happens at that SQL generation step.

  1068. 1:40:45

    And, um, and I believe we have a template for that. Uh, I believe, uh, Pamela, uh, has created a template called RAG on PostgreSQL, which is in our, in the same Azure samples, uh, GitHub account, um, that does exactly what you're saying.

  1069. 1:41:05

    It, it takes, uh, a natural language query, transforms it into a SQL query, and executes, uh, it on a PostgreSQL database. But you could do the same with, um, Cosmos DB.

  1070. 1:41:20

    So that actually leads me into another topic which I wanted to get to before we, uh, run out of time here today, [clears throat] which is about evaluation. So this is an important step anytime that you put any kind of an LM-based application into production where users are gonna be providing input to that.

  1071. 1:41:39

    And in this context of a chatbot, the kind of questions you want to ask are [clears throat]

  1072. 1:41:44

    did my chatbot give a relevant answer to my user's question?

  1073. 1:41:52

    Was the chatbot's answer grounded in the information that is available in my databases that is part of my RAG flow? Was the chatbot's answer coherent? Was it good English?

  1074. 1:42:07

    Was it understandable? Um, and the other metric that is in that list, which I'm trying to remember right now, is, uh, I'll get back to in a minute. But when you get to step number

  1075. 1:42:20

    thirteen, we're gonna take you through a Python notebook which shows you a process for answering these questions, uh, manually, essentially, and then I'm gonna show you how that's built into the Azure AI Studio platform itself.

  1076. 1:42:37

    But we think of debugging in just regular apps tests and tests that we write for application. That's very simple. It's a yes/no. Like, did the, um, did the application return a positive value when it should be a positive value?

  1077. 1:42:53

    Very easy thing to test for in programming style. Much more difficult test to answer the question was, is the answer generated by my chatbot relevant? How would you even program such a thing?

  1078. 1:43:09

    And the answer is, you get an LLM to answer that question. Now, this particular chatbot application we have running on GPT-35 Turbo.

  1079. 1:43:20

    Very cheap, very fast, very reliable LLM. Um, GPT-4o Turbo just came out recently. I haven't played around with it a lot myself, but I imagine that will probably take the place of GPT-35 Turbo in a lot of these applications pretty soon.

  1080. 1:43:37

    Next time we run this workshop, we're gonna switch it over to using GPT-4o.

  1081. 1:43:42

    But you've also heard of GPT-4. They are very large, very powerful, um, LLMs that have reasoning capabilities in some sense. Now, you wouldn't wanna use GPT-4 in a production application like this because every time the user types in a chat, not only are they gonna have to wait quite a long time for a response, but it's gonna

  1082. 1:44:03

    cost you a lot of money on the endpoint. In this RAG architecture, GPT-35 works great, as long as you give it the context it needs to answer that question.

  1083. 1:44:13

    But for this testing paradigm, for asking the question, is the answer

  1084. 1:44:19

    "Trail Master jackets are good" to the question, "What jacket should I buy?" Is that relevant? That is the kind of question a powerful LLM like GPT-4 can answer quite readily.

  1085. 1:44:32

    So think about how you might automate that process. You might use the prompt to GPT-4. Given this question and this context, the stuff that we put into the RAG, and this answer,

  1086. 1:44:47

    ask GPT-4 on a scale of zero to five, how relevant is this answer?

  1087. 1:44:54

    How grounded is this answer in the data that I've also provided here? Is that answer coherent? And these are all things that GPT-4 can do quite readily. And we can use the scores in this case that GPT-4 provides as a ranking of how well GPT-35 is doing in our endpoint for generating its answers based on the RAG

  1088. 1:45:15

    process. And that's exactly what goes on here. In this notebook,

  1089. 1:45:20

    at the top of it, you can put in a question. I just ran it on, "Can you tell me about your jackets?" You can have a look at all the code.

  1090. 1:45:27

    You can even see the prompts that it's using to GPT-4 to answer these questions.

  1091. 1:45:32

    And you can see the actual answers that came back are in the next node up here. Here we go. "Hey, Sarah Lee, let me tell you about our jackets.

  1092. 1:45:40

    We have two awesome things, awesome options that will give you a previous purchase. Summer Breeze jacket," et cetera, et cetera. So that's the answer that the LLM came back.

  1093. 1:45:48

    This is the context that the RAG process was provided by that it used to generate that answer. And then with all the information, we can ask those questions we just asked.

  1094. 1:45:58

    Was that answer about jackets grounded in Contoso's product database? And GPT-4 ranked that a scale of five. Uh, ra-- uh, so a rank of five on a scale of zero to five.

  1095. 1:46:08

    And likewise, we can ask questions about coherence, fluency, groundedness, and relevance, and we get the answers. Uh, this, this particular question is doing really well. You probably also wanna test out your LLM

  1096. 1:46:23

    on some adversarial types of responses. For example, you might ask it the question, you know, uh, "I want a bright toothbrush."

  1097. 1:46:34

    Now remember, this is a hi- this is a camping store. They don't have toothbrushes. Nothing is gonna come up in the database when we do the RAG search. Well, well, actually, something will come back because we always get back some responses that are somewhat close to our query.

  1098. 1:46:48

    But let's see how our LLM actually does here. When I run this notebook, it's gonna run through those scripts. It's gonna pass that question to our RAG flow, generate the response with GPT-35, and then ask GPT-4 to rank it on those four scripts using the prompts that are linked to in this script.

  1099. 1:47:08

    And when we come back to it, we can see the answer it came back with was, "Hi, Sarah. Since you're into outdoor, uh, adventures, I recommend the Fresh Breeze..."

  1100. 1:47:18

    Where's my scroll bar? Here we go. Fresh Breeze travel something, something like... There it is. There's my scroll bar. Fresh Breeze travel toothbrush. All right. This is interesting. Contoso does not sell a Fresh Breeze travel toothbrush.

  1101. 1:47:32

    Uh, GPT-35 just made that up out of whole cloth. But this is what LLMs do, and we have to test to see whether or not they're doing these kinds of things for the type of use cases that we anticipate.

  1102. 1:47:43

    And we can detect that particular test is not going well. Groundedness, score of one out of five. It really wasn't grounded in our data because there was nothing about toothbrushes in our context data that we provided through RAG.

  1103. 1:47:57

    And similarly, coherence, well, it was in nice English, so it got a f- score of four for coherence, a score of five for fluency, but one for groundedness and one for relevance.

  1104. 1:48:08

    And so now you can think about automating this process. You can think about what are the types of questions that we want our application to do well at. What are the types of questions that we might want to say not give any responses to at all and score accordingly.

  1105. 1:48:23

    And I won't go through all the details of this, but when we get into AI Studio,

  1106. 1:48:28

    there is a whole section here on evaluation. Now let me just discard that.

  1107. 1:48:36

    And this is the process where you can actually load into it a bunch of tests, which in this case are not Python code or C# code or whatever. It's responses.

  1108. 1:48:46

    So questions, responses, and context. And then automate the process of evaluating how your endpoint, how your RAG bo-- RAG process does on all those questions. So the next time when you add new products, or next time when you decide to upgrade from GPT-35 to GPT-4o, you've got a series of tests ready to go to evaluate how your

  1109. 1:49:08

    ou-- well your application does in face of those changes. All right. Questions, yeah.

  1110. 1:49:13

    Quick question on that.

  1111. 1:49:14

    Yeah.

  1112. 1:49:14

    Um, so for, um, after you've evaluated the model and you sort of understand the performance of it-

  1113. 1:49:21

    Mm-hmm.

  1114. 1:49:21

    -what typically are your next steps, and what actions do you take to drive, uh, improvement on the measures that you see there?

  1115. 1:49:27

    That is an excellent question. I have a slide just for that. I honest-- This was not, uh, planned. [laughs]

  1116. 1:49:33

    But it's, this is essentially the LLM Ops, LLM Ops process, which is essentially the same as DevOps, but with a fancier name that gets you lots of funding. Um,

  1117. 1:49:47

    here, here's the life cycle. So exactly the same idea as when we build applications. We go through the ideation process, we figure out our use case, uh, we do some exploration, testing it to our data, we build our basic prompt flow in the LLM case, and we develop our first version of that flow.

  1118. 1:50:05

    And then we actually run it against sample data. This is the evaluation step. We're still very early on the process here. If the evaluations don't give the scores that we're looking for before we put it to production, the next step then is to modify our prompts.

  1119. 1:50:19

    You saw that ninja template with a bunch of prompts around do this, don't do that. You would modify those until you get the behavior that you're looking for. Maybe you change the process in RAG.

  1120. 1:50:28

    Maybe you chunk- ... the data differently, maybe present it differently to the RAG process. Then once you get satisfied in that process, you would keep on testing that against perhaps, you know, a live user cohort or bring in some testers, bring in a red team to try and break it, and again, go through that same evaluation process

  1121. 1:50:46

    until you're satisfied. And then finally, you'll be ready to actually deploy that to your production.

  1122. 1:50:51

    You'd put monitoring in. You'd actually monitor live, um, probably a sample of actual user questions, responses, and have, have, you know, real-time charts. Not real-time charts, actually. Probably daily charts of how your model is doing in scores, like groundedness for the types of questions that you can ask.

  1123. 1:51:09

    And that, that might be detecting, you know, maybe things are drifting because your product set has changed and there are trigger words in your products that are making the, the GPT model do strange thing.

  1124. 1:51:19

    Maybe you've got some adversaries that are coming in to try and hack into your system. That might come up in some of your monitoring scores. And then you go back through that reiteration process to go back and build and augment the model for its next deployment.

  1125. 1:51:31

    Is that the kind of question you're looking for?

  1126. 1:51:33

    Yes, that would be great.

  1127. 1:51:34

    Great. Other questions?

  1128. 1:51:37

    I have one question.

  1129. 1:51:38

    Yeah.

  1130. 1:51:39

    You know, right now when we look at the input, you know, from the user, you, you put text. You know, you put what you have purchased or something like that.

  1131. 1:51:47

    Mm-hmm.

  1132. 1:51:48

    Can this be improved to take like, you know, graph or PDF file?

  1133. 1:51:56

    I'm not sure my question is clear.

  1134. 1:51:57

    Yeah. I, I don't see how graph fits in there, yeah.

  1135. 1:52:00

    Let's say, you know-

  1136. 1:52:00

    Oh, a PDF file. Yeah

  1137. 1:52:01

    ... PDF file. Let's say-

  1138. 1:52:02

    Yeah

  1139. 1:52:02

    ... I want to input my PO-

  1140. 1:52:04

    Yeah

  1141. 1:52:04

    ... rather than I type something.

  1142. 1:52:05

    Yeah.

  1143. 1:52:06

    Can this accommodate that kind of re-

  1144. 1:52:08

    Absolutely. A- Azure Search, for example, can index PDF files, and then you can do a search to find the PDF file that is most relevant to that user's question.

  1145. 1:52:15

    Mm.

  1146. 1:52:15

    You can then extract out from that PDF file context-

  1147. 1:52:18

    Mm

  1148. 1:52:19

    ... which is put into the prompt, which is then used to generate the response. And then you can put references back to those source files if it's a trusted user kind of a situation-

  1149. 1:52:27

    Uh-huh

  1150. 1:52:27

    ... so they come back to see them. That's how Copilot works, for example.

  1151. 1:52:30

    Okay.

  1152. 1:52:30

    Yeah.

  1153. 1:52:31

    But still, you know, from the, you know, prompt-

  1154. 1:52:33

    Mm

  1155. 1:52:34

    ... you can only input text, right?

  1156. 1:52:36

    For the prompts? Well, yes and no. [laughs] [laughs] Um, this particular example, everything is converted into text. But today we have what we call multimodal models.

  1157. 1:52:45

    Yeah.

  1158. 1:52:46

    Uh, GPT-4o, for example, as the prompt you can input not just text, but also images, even audio. Um, not video yet. Um, but-

  1159. 1:52:55

    Not video yet

  1160. 1:52:55

    ... but you could set up that RAG application to insert into the prompt the images or the audio or whatever it is you want the LLM to be able to reference.

  1161. 1:53:03

    Great. Great.

  1162. 1:53:04

    Um, that's still relatively new, 'cause 4o doesn't have all of its multimodal capabilities out yet, but the, the principle exists.

  1163. 1:53:12

    Okay.

  1164. 1:53:12

    Yeah.

  1165. 1:53:12

    Good.

  1166. 1:53:13

    And we have, uh, Florence that, uh, we can use.

  1167. 1:53:17

    Yeah. Mm-hmm.

  1168. 1:53:18

    Um, and, uh, Florence version two was actually released earlier this week, which is a model which allows you to do image to text. So you can, uh, analyze an image, uh, generate text out of it, and then take the text and give it to GPT, uh, 3.5 or something else.

  1169. 1:53:34

    Aren't you giving a talk about that tomorrow? [laughs]

  1170. 1:53:37

    Yes. Well, I mean, briefly. It's one of the things I'm talking about, yes.

  1171. 1:53:41

    Okay. And what time is your talk tomorrow?

  1172. 1:53:43

    Uh, that's a very good question. I forget. [laughs]

  1173. 1:53:46

    It's in my calendar. I'll take a look. [laughs]

  1174. 1:53:49

    Uh...

  1175. 1:53:49

    Yeah, I'm going to, you know, join us then because that's one, uh, uh, kind of primary request from our team. And right now I have built, you know, this, uh, you know, customized, uh, prompt window, but it can only take text, and now they want to say, "Okay, I want to use PDF file or even JPEG file."

  1176. 1:54:12

    So, well, like we said, for J- for JPEG file, you, uh... when GPT-4o, uh, when we make available on Azure, um, the multimodality, uh, capabilities, then you would be able to use it directly.

  1177. 1:54:28

    For now, uh, you can use a- another image to text model such as Florence 2 or something else. Uh, for PDFs-

  1178. 1:54:36

    Yeah

  1179. 1:54:36

    ... so it really depends exactly what your use case is. Like, is it a transient use case? Are you storing the PDF, uh, long term? Because if it's the latter-

  1180. 1:54:48

    Transient

  1181. 1:54:49

    ... okay. So that's... Okay. So we... If it's transient, then, uh, an alternative approach is instead of taking the PDF and storing it into the vector database and indexing it long term, what you can use, we can...

  1182. 1:55:02

    Ah. One, two, three. What you can use is do that, um, uh, you, you upload the PDF, you chunk it, and you can actually... The, the algorithms, uh, for embedding the chunks, um, you can actually run them in Python in memory.

  1183. 1:55:18

    You don't, you don't have to do it, like, you know, in a long-term vector database. So you can do the chunking and the embedding in memory and actually the vector similarity search function that to, to find what's relevant.

  1184. 1:55:32

    You can, uh, execute those functions in memory to provide your users, um, with a transient, uh, FMR experience, where they upload the PDF and you query the P- the, the PDF just for the sake of the current conversation.

  1185. 1:55:49

    Okay. I think I need to take this offline with you-

  1186. 1:55:52

    Yeah

  1187. 1:55:52

    ... because my s- uh, use case is a little bit different. I'll explain a little bit, you know, in detail and then, uh, you can-

  1188. 1:55:58

    Sure. Uh-

  1189. 1:55:59

    Oh, yeah

  1190. 1:55:59

    ... and you can do it in memory or in a PostgreSQL database or a Cosmos DB. That works, too. And delete the data after, once you're over with it.

  1191. 1:56:09

    And the last thing I wanted to say regarding, 'cause that's a good question, is... The last thing I wanted to say is right now, today, you can go to Azure OpenAI, you can deploy the GPT-4 model.

  1192. 1:56:20

    And in the chat, they have like a chat section where you can chat with your model.

  1193. 1:56:25

    Yes.

  1194. 1:56:25

    There, you can enter images. There, today, right now, you can go there. There you can enter in a picture, and it can do things with it. Like, "Hey, here is a picture of a website.

  1195. 1:56:34

    Can you write the code for the website?" And it will write code according to the picture.

  1196. 1:56:39

    I know this one. I have done that already. And this is more related to RAG, the prompt window.

  1197. 1:56:46

    Mm-hmm.

  1198. 1:56:46

    And the one I, one I... So far, you know, what I have w- you know, developed can only take text. It cannot take picture-

  1199. 1:56:54

    And-

  1200. 1:56:54

    ... or any PDF file.

  1201. 1:56:56

    And not to cut you off, 'cause we love these detailed questions, but I've been told that I'm gonna get cut off up here in just a minute. [laughs] Um, and before I do that, I just want to let you all know that, um, we'll be here for a few minutes for in-person questions around, but also come to the

  1202. 1:57:07

    Microsoft booth in Salon Nine. Uh, lots of people there to have, uh, you can ask exactly these kinds of questions of. So please go ahead and do that. Uh, Cedric is also giving a talk tomorrow about multimodal models at, uh, 12:30 PM in Salon 10, so please come on and check that.

  1203. 1:57:24

    But thank you, everybody, to coming today. Um, if you'd like to do this at home, the repository is already in your GitHub accounts, and if you happen to miss that step, there's a QR code where you can get to it there as well.

  1204. 1:57:34

    But thank you, everybody, for coming today, and, uh, enjoy the rest of the conference. [outro music]