← All AI Engineer talks

AI Engineer World's Fair 2024

Accelerate your AI journey with Azure AI model catalog

About this talk

Sharmila Chokalingam and a second presenter identifying herself as Shubhi demonstrate how Azure AI Model Catalog supports foundation-model selection, standardized inference, retrieval-augmented generation, deployment, and enterprise safeguards. They survey GPT-4o, Phi-3, Mistral, Llama, Cohere, and Jais; demonstrate Cohere Command R, LangChain/LiteLLM integrations, Azure Marketplace billing, benchmarks, playgrounds, and prompt flow; and discuss privacy, security, and [REDACTED:url]ai.

Chapters

  1. 0:00Introductions and Azure AI model-selection strategy
  2. 3:14Foundation-model catalog: GPT-4o, Phi-3, Cohere, Mistral, and Jais
  3. 4:37Cohere Command R deployment, integrations, billing, and benchmarks
  4. 7:19Chat playground and end-to-end prompt-flow demonstration
  5. 16:25Enterprise privacy, customer adoption, and closing

Talk transcript

  1. 0:00

    [upbeat music] Hi everyone. Thank you for joining us for the session on how you can accelerate your AI journey with Azure AI Model Catalog.

  2. 0:20

    This is Shubhi, I'm one of the product managers on Azure AI.

  3. 0:23

    And I'm Sharmi. I'm one of the s- product marketing managers on Azure AI.

  4. 0:29

    Great, so let's get started. So today, we'll be talking about how we offer the best collection of foundation models on Azure, why the model choice matters so much when you have such a huge collection of models, and how does Azure AI model inference API make it standard and easy for you to switch between multiple models, swap one

  5. 0:47

    out for another without disturbing your rest of the code base, and then how you can build generative AI apps on top of it, and how the platform makes sure that all your enterprise readiness needs, like data privacy, security, and content filtering are met.

  6. 1:01

    And then let-- we'll talk a bit about some of the customer success stories.

  7. 1:05

    So Shubhi, as you mentioned, uh, there are a lot of large language models that are being, uh, released pretty much every day, every week. And, uh, I know in the Azure AI Studio and Model Catalog, we are bringing a lot of models, like large language models, foundational models.

  8. 1:22

    I want to understand why are we bringing all, all these models, and why does it matter to a customer?

  9. 1:29

    Yep, absolutely. That's a very valid question. So let's talk about why model choice matters in the real world, right? As Sharmila asked. So the three questions that we all try to answer when we try to build generative AI apps is, can AI solve my use case first, and then what is the best model for my use case,

  10. 1:47

    and then how do I go about scaling this for the production workloads? So the first step is called prototyping, where you try, try out multiple models that are available to you.

  11. 1:56

    You try to establish feasibility, build that prototype, do compare model benchmarks, and find the right one for your use case, or maybe at least shortlist some of them. And that's when you move to the next stage where you try to optimize it for your use case, where you try to optimize for cost, latency, regional nuances, et cetera.

  12. 2:13

    And that's where the Azure AI platform comes into picture, where you can use the techniques like prompt engineering, RAG, and fine-tuning to make sure that you've optimized it. Once you're done through this loop of prototyping with the model catalog and optimizing with the platform, you can go ahead and operationalize it.

  13. 2:27

    So you don't have to worry much about the capacity and cost trade-offs. You have to-- You have got monitoring, you've got scalability, you've got data privacy, content filtering, and all your enterprise needs are met.

  14. 2:39

    So once you're through this, you have your generative AI application production.

  15. 2:44

    So another question I had here is, uh, have you seen ca- situations where a customer is using more than one model for a specific use case and even for multiple use cases?

  16. 2:54

    Yeah, absolutely. Like, that's the whole point of the model catalog, that you have this wide range of models. You have your use case, and you're able to plug in any model in from that model catalog.

  17. 3:06

    So you can swap out one for another without disturbing anything with quick prototyping and quick comparison.

  18. 3:14

    So as we mentioned, let's talk a bit about the model selections that we offer on our platform. So we offer a wide range of flagship LLMs and SLMs. Recently, we launched the Azure OpenAI models.

  19. 3:24

    The GPT-4o is already on the platform. We launched Mistral models, Llama models, Cohere models, as well as small language models like Phi-3 and the Mistral OSS models. Along with that, we make sure that your multimodal requirements, image generation requirements, and specific needs such as embedding model requirements are also fulfilled on this platform.

  20. 3:44

    So GPT-4o, for example, has function calling and JSON support, so you can make sure you use that for your agent-centric workflows. Along with that, we also make sure that we cater to your region or language-specific needs.

  21. 3:55

    So for example, the Mistral language is really good with the European languages or, uh, the Cohere embedding multilingual ma- model is very good for your multilingual requirements. And we also recently launched Jais on our platform, which is an Arabic LLM.

  22. 4:09

    Along with all these flagship and premium models, we also have hundreds of open models from Hugging Face, and we've been actively partnering with Meta, Databricks, Snowflake, and NVIDIA to make sure that we get their models on our platform as soon as possible.

  23. 4:25

    Okay, so now that we have, uh, seen what all we can do with the catalog, let's try to see a live example of how to actually go to the catalog and deploy your models.

  24. 4:37

    So once you type into your URL ai.azure.com, it's as easy. You land on this AI Studio page where you can go to the model catalog on this left nav bar, and you land on this page that has the list of all the models that we offer.

  25. 4:50

    It, it has sixteen hundred plus models right here. And to make it easy for you to filter it for your use case, you can filter by the dif-different deployment types, different inference tasks that you want to do, or even if just by the model collection families.

  26. 5:03

    So let's start out by filtering for the Cohere models. Like click on Cohere right here. Let's try to see Cohere Command R, for example. Once you click on the model, you land on this model card page where you have all the information about the model, how, uh, you can customize it for your own use case, tool use

  27. 5:20

    capabilities. This one specific talks about RAG capabilities because the Cohere Command R model. It-- all model catalog pages also have these inference samples that you can use as starter codes to get started with, for example, a LangChain SDK or a LiteLLM SDK.

  28. 5:37

    As we mentioned, like, we try to standardize the APIs across all use cases, so you can just plug in your R APIs into any third-party application like the LiteLLM and have it working in no time.

  29. 5:49

    So once you've gone over this, check the pricing, we go ahead and click on Deploy. And this is where we make sure that we are connecting to the Azure Marketplace.

  30. 5:59

    So we use Azure Marketplace just for the billing side of things to make sure that you are billed correctly based on your token usage. And this is the step where you actually subscribe to that offer.

  31. 6:09

    Here, I've already subscribed to that Microsoft subscription, so it's giving me the option to continue to deploy. It's as simple as choosing a deployment name, checking if you want to enable content filter or not, and clicking on Deploy.

  32. 6:22

    So under a minute, you'll have your URL and key ready to get started. So while this is happening, let's look, uh, at other capabilities that we have, like model benchmarks.

  33. 6:33

    So when you're in the Azure AI Studio, you also have the ability to check model benchmarks, which is right here on the left under Model catalog. Once you go in here, here I'm showing all the models that we have.

  34. 6:44

    You can see that we try to benchmark on certain common characteristics like the model accuracy, model coherence, groundedness, et cetera. And this is the perfect place for you to filter out which model you want to choose based on the extreme selections of models that we offer.

  35. 7:00

    So let's go back to check... Yeah, and we see... Go back to the deployment that we created, and it got created within a few seconds. We have our target URL right here and the key ready to use in any code base that you already have.

  36. 7:12

    So now you may be thinking that before I move on to using my IDE, I want to try it out a bit, right? Is there something like a playground?

  37. 7:19

    And that's why we also have this playground capability. So once... It's also in the Azure AI Studio on the left if you see we're right under the Chat Playground.

  38. 7:27

    So let's see a live example right here. You can choose the model that you want to use in the playground. So in this deployment section, I've chosen a Mistral Large deployment that I already have in this project.

  39. 7:38

    Let's try to, uh, chat with this model. Right here, it's not customized on any, any data. I'm just directly asking the model. So yeah, I'm trying to ask, how does Microsoft promote the culture of giving?

  40. 7:48

    So this will, in general, give me a generic response about how it has a culture of giving through various initiatives. It has employee match programs and some generic information that's available online.

  41. 7:59

    But what if you want to s-specialize it for your-- for our own data? So here we can go ahead and use the add your data functionality, where you can choose an a-available index.

  42. 8:09

    So let's choose an available index that's called Microsoft Give. That tells it in specific that what, uh, what are the specific things that are very particularly known internally or may not be available generic circumstances.

  43. 8:24

    If we send out the same question right here, we should get a more targeted response based on the documents in that index.

  44. 8:33

    So just wait for a few seconds.

  45. 8:37

    So Shubhi, while this is happening, I had a quick question. It's great that we are doing all this. I'm just curious because you mentioned data and data source and everything.

  46. 8:46

    Are we using any of the data from our customers to train the models, or is Microsoft using it? Are our model providers using any of the data that a customer brings in?

  47. 8:57

    So that's a great question, Sharmila, because that's a very common question we get from our customers. And no, we have very strict data privacy and policy rules in place.

  48. 9:06

    Your prompts and your completions are not shared with the model provider, nor your data is used for training any of the models.

  49. 9:15

    So yeah, looking back at the results, we see that it gave us a very specific response that says, "You get fifty USDs to start off with the new hire credit for the giving program."

  50. 9:24

    And that wasn't in the response earlier. So with just the click of a button, we were able to link it to an index and get that response. So that's how the playground works.

  51. 9:36

    So talking about the different ways of deployment, the one that we just saw was a serverless API option. So in the model catalog, there are two ways you can deploy a model.

  52. 9:45

    One's called the managed compute, and one's called the serverless API option. With managed compute, the user is responsible for getting their own GPU. So you basically pay for the VMs per hour, and you're responsible for the quota management, capacity management, and you can use hundreds of open source models with this.

  53. 10:02

    The second, uh, way you can deploy models is by getting a serverless API, and this is available with both Azure OpenAI service and Models as a Service. And this is what has about thirty-plus flagship models, premium models that you pay for based on your usage.

  54. 10:17

    So you get ready-to-use APIs, and you only pay for the input tokens or the output tokens that you use.

  55. 10:24

    We've also put in a lot of effort to make sure that we standardize the schema and the APIs of these models for you. So we've worked with the model providers to make sure that we build an SDK on top of a very standardized REST API system.

  56. 10:37

    And such an, such an SDK works with common, uh, open sourced applications, things like LangChain, as well as the model provider specific SDKs. So all you have to have is a different endpoint, and every endpoint has the same API structure and the same SDK structure.

  57. 10:55

    So you can just swap in one for another, evaluate, create multiple evaluations, compare the results, and choose the one that's perfect for your use case.

  58. 11:07

    Yeah.

  59. 11:08

    Awesome. So, um, we've been seeing everything about the model catalog and models. Can you show us an actual use case example?

  60. 11:17

    Yeah. So, uh, let's briefly talk about how you can actually use these APIs in your IDE. Let's talk about the function calling example, and let's take the Mistral Large model for, for that use case.

  61. 11:28

    So here I have a simple function calling example where I'm trying to que- use this model as a chatbot, for example, of an shop. The shop has-- sells certain stationery items, it has certain specific pricing, and may have certain ongoing discounts.

  62. 11:41

    So if you just use a model as a black box, you will not get a specific pricing for the model, uh, for the shop or any of the ongoing discounts.

  63. 11:49

    But what I'm doing here is using the function calling capability of this Mistral Large model to define a function called Get Bill Amount that can take in the specific information that we fed to it, recognize that it needs to call this external function based on the prompt, and smartly make that call, query that result, and give you

  64. 12:07

    the exact information. So right here I've defined that function, I've defined the tool for that model for Mistral Large. We send in a prompt that says, "You're a helpful assistant that helps users find how much they have to pay."

  65. 12:19

    And we also make sure that we tell it that you also care about the environment, and you also have to help users understand possible things they should be careful of when using these items.

  66. 12:28

    So this is just to, uh, add more context to the response and see how the model can adjust based on the requirement. We go ahead, we send this response in.

  67. 12:37

    We can see that the model has intelligently identified that it is calling the function get bill amount with the right arguments to identify that we queried for a stapler, and we'd asked-- tried to ask what is the price for the 10 staplers.

  68. 12:49

    And if we see the chat response, it says, "The cost of 10 staplers, including any ongoing discounts, is $45," which is very specific, and it also makes sure that it res-re-reminds the users to be mindful of the environment and try to use staples when possible.

  69. 13:03

    And this is end result of the system message that we sent to it when we asked it to be environmentally friendly and give users the right context. So similarly, if you see right here, you can swap any code base with, uh, any endpoint that you have, and without putting mu-much time into it or much effort into it,

  70. 13:22

    you have a running API app right here. We also make sure that we use, uh, uh, model provider, um, fields like the safe prompt setting to true. So on top of...

  71. 13:34

    We always build on top of what the model provider capabilities are already existing.

  72. 13:39

    Awesome. Thank you.

  73. 13:42

    So now that we've talked about how we can set it up in the IDE, let's talk a bit about how you can set it up in the UI and how we can create a generative AI app using prompt flow.

  74. 13:51

    So when you try to... Here I'm trying to create a shopping assistant chatbot using prompt flow, where it's a simple RAG application where we take in the user prompt, we try to get retrieval, uh, we retrieve context-specific information from our index, and then, uh, send it to the LLM to generate an output.

  75. 14:07

    Here we've created the lookup step for it, which is basically doing the RAG part of it. The generate part is going to generate the output from the LLM. But we've added this extra step of rephrasing, where we're using the query transformation technique, where we take in the user con, uh, prompt, which is generally very succinct, but we

  76. 14:25

    try to make it more verbose by rephrasing it, because we've seen better results of RAG with that.

  77. 14:31

    So let's look at a live example of this right here. I have this prompt flow running right here, my compute session's running. We see that the first question that we ask is, "Do you have any new hiking shoes?"

  78. 14:43

    But the rephrase step rephrases into a longer verbose output that says, "I'm looking for hiking avail-- hiking shoes available, and if so, what materials and features?" So it basically elongated that question.

  79. 14:56

    We check the output of the lookup step, and we see that the prompt was able to get the context-specific information from the index that we provided to it, so it identified certain amount of information that we can now send to our LLM in the next step, and the LLM generates a response.

  80. 15:11

    So based on that specific information, we were able to get this output that recommended the Fleece Fit, uh, Flex Jacket for the women. So we can see that we are able to generate a prompt flow end to end.

  81. 15:24

    But you may be wondering that how do I make sure that I'm able to plug in different models into this flow? And this is where you can try to create variants.

  82. 15:31

    So here you can see that in the generate step, I'm using a connection from the Cohere Command R model, but you can go ahead and choose any other connection to any other model and use evaluation to try to compare the results from the sa-- for the same flow for different models.

  83. 15:46

    So example, for the, um, first step when we're trying to create the embeddings, here I've used the Ada model, but you can go ahead and try to see, okay, how does the Command R model work with the Command Embed model?

  84. 15:57

    So you can create these variants and try to see the evaluation results. In the interest of time, I already ran some evaluations, as you also saw in the previous, uh, demo, and here we're trying to compare the Cohere Command R versus Llama Three versus the Mistral Large.

  85. 16:11

    And the evaluation capability helps you to compare the same model for the same flow on different parameters, and you can see how one fared against another.

  86. 16:25

    So this is all great, and I think you touched upon data privacy a little bit. So can you go a little bit more into the details of what else do we have in the AI Studio or, um, Model Catalog to ensure customers' privacy and data security?

  87. 16:40

    Yeah, absolutely. That's the key. So let's talk a bit about how we ensure that the data privacy and security compliance needs are met. So as I mentioned, there are three pillars to this.

  88. 16:49

    So one's the data privacy part, second is the security and compliance, and the third is the responsible AI and the content safety. Talking about data privacy, for both managed compute and serverless APIs, your prompts and completions are not shared with the model provider.

  89. 17:04

    Your prompts and completions are not used for training the models. No data is shared for training or with the model providers, so you can be assured that the AI platform makes sure that your enterprise needs for data privacy are met.

  90. 17:18

    We also have this additional feature of adding hidden layer to our model scanning. So we make sure that we are finding the embedded malware and backdoors. We-- It scans for common vulnerabilities and exposures and detects tampering and corruption across model layers.

  91. 17:32

    So for any model that's labeled curated by Azure AI, you can be assured that it's passing through the required checks.

  92. 17:40

    Talking about security and compliance, in addition to the data privacy norms that we mentioned, we also offer the capabilities of adding private networking so that your data is not exposed to the internet.

  93. 17:50

    So you have the control over routing your ingress and egress traffic through the VNets, and you also have the ability to set up FQDN rules so you can approve outbound access to non-Azure resources.

  94. 18:03

    Um, in addition to this, you can also regulate access to models with Azure Policy integration. So you can have allow list or deny list patterns, and you can split out which model collections you want access to or not.

  95. 18:15

    And you can also use these different policies for separating out the dev, test, and production environments.

  96. 18:25

    Awesome. So, um, Shubhi, thank you so much for going through the model catalog and AI Studio, what, what are the features available in there and all, all these great demos.

  97. 18:34

    So now I just wanna go into a few customer success stories and, uh, the... One of the main, uh, kind of underlying theme for all the customer success stories that I'm gonna show is that these customers are not just using one model in-- for their use case, they're using multiple models from the model catalog, and it could

  98. 18:52

    be, uh, in one use case or across multiple use cases, similar to what Shubhi has shown in the demo. [lips smack]

  99. 18:58

    And the first customer we're gonna talk about is [REDACTED:url] And, uh, they've been using, uh, our large language models. They've started off with the OpenAI models that was available in the Azure OpenAI service.

  100. 19:09

    Um, they are doing that today, and they're also looking into Llama models, uh, where they are looking into Llama models for, like, really, uh, uh, task-specific use cases, like for documentation and for summarization and all that.

  101. 19:23

    And, um, they're using our model catalog. They built [REDACTED:url]ai, which is a generative AI platform for EY professionals, which addresses the need for enhanced pro-- uh, productivity and accuracy in professional tasks.

  102. 19:36

    And one of the key things that we wanna show is what's the result of what they've been doing. The EYQ chat that they built has been adopted by two hundred and seventy-five thousand employees internally and allowed the employees to perform a wide range of tasks efficiently and with great accuracy.

  103. 19:53

    And what... Some of the lessons that they have learnt is it's not AI... Using AI in their use cases is not, like, a one-time thing. They need to do continuous evaluation of AI performance, and they wanna stick to all the responsible AI practices.

  104. 20:07

    And that's one main reason why they've been using Azure AI, uh, uh, Model Catalog and Azure AI Studio. It's because they feel that they can easily do this evaluation and, um, make sure whatever they're doing, putting in production is going to be, uh, really an, uh, adhering to safe and responsible AI. [lips smack]

  105. 20:27

    And then the next customer I want to talk about is CMA CGM. Again, they are a big, uh, global player in sea, land, air, and logistics solutions. And they are also building a kind of like a robotic process.

  106. 20:40

    Uh, they've been doing traditional robotic process automation, and they've decided to use Mistral model from a model catalog, and they have built, again, a similar chatbot-like scenario, uh, for their customer care agents.

  107. 20:52

    And, um, one thing that they have seen is they have seen a reduced response latency and increased customer satisfaction in their chatbot, uh, use case. And they plan to, again, extend the application of LLMs to encompass specific products for core business activities like invoicing, customer document analysis, uh, interpreting free client text, and, uh, writing emails and all

  108. 21:16

    that. And finally, the last customer story I want to share is Bridgestone. Again, Bridgestone is a very popular name. Um, their, their use case is, uh, they have been using the Nixtla model from a model catalog.

  109. 21:29

    We launched Nixtla, uh, time series model at Build last month, and it's a time series forecasting model, and they're using it specifically to predict monthly demand for a vast portfolio of products.

  110. 21:41

    And, um, one of the things that they want to do is streamline their forecasting pipelines, enhance accuracy, and reduce operational complexity. And again, what they have seen is that they have seen that using the forecasting model like TimeGen from Nixtla has helped them reduce, uh, errors by nearly thirty percent on average in forecasting errors, uh, which is

  111. 22:02

    huge. And, uh, again, in all these customer use cases, the time it took for them to start using LLMs in their applications to seeing the results and impact has been reduced significantly because they were being able to use a model catalog and Azure AI Studio where we provide, as Shubhi showed in the very first slide, we do--

  112. 22:23

    we have tools for prototyping, um, optimizing, and operationalizing. So whether you're just starting off with, "Let me try this for a prototype project," to realizing, "Okay, I need to put it into production," the time it takes from going from that to the last step has been reduced significantly, mostly because we have streamlined all the different, um, foundational

  113. 22:43

    models. We have provided the right tools for all our customers to kind of go through that whole, uh, LLM life cycle. I think that's pretty much it. Uh, thank you.

  114. 22:54

    Okay. Thank you. [upbeat music]