← All AI Engineer talks

AI Engineer World's Fair 2025

Critical AI Inference Your CIO Can Trust

About this talk

Sahil Yadav and Hariharan Ganesan examine why enterprise AI inference must be trustworthy in telecom, industrial IoT, and other safety-critical environments. They outline three operational foundations—plain-language explainability, adaptive guardrails with human escalation, and digitally signed traceability—then describe XTOps trust dashboards, the MTTRE incident metric, and a Guardhat industrial-safety example involving GPS drift. The discussion closes by acknowledging the difficulty of quantifying the financial value of AI trust.

Chapters

  1. 0:00Introductions and the enterprise AI trust problem
  2. 1:12Governance gaps and safety-critical failure scenarios
  3. 2:57Three pillars: explainability, adaptive control, and traceability
  4. 8:30XTOps guardrails, trust dashboards, and MTTRE
  5. 12:18Guardhat case study and detecting GPS drift
  6. 18:39Closing discussion: quantifying the value of AI trust

Talk transcript

  1. 0:00

    [on-hold music] Hey, guys.

  2. 0:15

    Thanks for being here. I'm Sahil. Uh, I am here with Hari. We're, we're presenting on AI. Um, obviously we're talking about trust, but let me give you a little background about us.

  3. 0:28

    Um, over the past ten years, we have deployed AI in various industries from, from health, um, monitoring to industrial IoT to, uh, network automation in telecom networks. And, uh, there's been one question that has been asked all along all this time:

  4. 0:46

    Can we trust AI? Some of these systems are used in mission-critical applications, but the question really is: Can we trust the inferences of the, these AI systems because they're impacting businesses, they're impacting business decisions and the bottom line at the end of the day.

  5. 1:05

    So we're gonna explore that topic today. With that, let me get us started.

  6. 1:12

    Uh, all right. So, you know, just like any presentation, we'll start with some stats. Um, so McKenzie is saying seventy-eight percent of the companies are adopting AI. There's another research from EY says ninety-five percent, uh, investing in AI.

  7. 1:29

    But here's the problem, only eleven percent of the companies are focused on AI's governance, ensuring safe practices with AI. So that sixty-seven percent gap is gonna be a huge problem because it's, it's not just about, uh, implementing the AI in the right way.

  8. 1:47

    It's also about understanding the impact of that AI in the long run. And let me quantify that with some examples. So you see a couple of examples here. Telecom disruption.

  9. 1:59

    Um, what really happened here was AI made a-- ma-made some decision. Based on that, there was a network disruption. Now, if you look at AT&Ts and Verizons of the world, they're spending millions of dollars each minute the network is, um, not working.

  10. 2:15

    Another example here is, uh, a gas sensor misinterpreted data. That put human lives at risk. Third company lost millions of dollars because a supply chain, uh, company, uh, AI screwed up the SKUs and, you know, ended up in losses.

  11. 2:34

    So what, what we're, what we're trying to say here is that these are silent failures. You cannot quantify the impact of these failures ahead of time. You cannot see them coming ahead of time.

  12. 2:44

    But they are, they, they are worth millions and billions of dollars over time. So this is extremely, extremely impactful as, as AI is getting adopted.

  13. 2:57

    So taking a view of what trustworthy AI looks like, there are three main pillars. You talk about explainability.

  14. 3:06

    When you're talking about explainability, I think the most important thing is, um, having a view of what's really under the hood. Um, otherwise you're just flying blind. You have to understand why those inferences are being made, on, on what basis.

  15. 3:20

    The other thing is traceability. Um, it's like a flight recorder. It's capturing all the audit trails. It's ensuring that you can retrace the steps and based on that, you can understand, uh, the particular situation, recreate the situation, and being able to solve it again.

  16. 3:38

    Guardrails, extremely important. They're to ensure that you don't end up in millions of dollars of losses. Uh, there, there, there is some threshold where, where it stops. The AI's gotta stop.

  17. 3:49

    So together, all of these build trust in a real system. More importantly, when you ta-talk about real-world scenarios where you're implementing this, you're talking about scalability.

  18. 4:01

    These are the pillars to think about when you're scaling them in the real world.

  19. 4:07

    I'll have Hari talk about the pillars of trust.

  20. 4:09

    Hey, thanks, Sahil. So every mission-critical system that we rely on today, be it aircrafts, be it energy grids, or be it even the simple banking financial systems, are built on principles of safety and understanding, right?

  21. 4:23

    Our AI systems should be no different. So looking at the first pillar, right?

  22. 4:29

    First, AI has to show its work. Every important decision shouldn't be a mystery. It should come with a simple English explanation so that a end user, a decision maker, somebody who is auditing the system, is able to act on the information and not look for a data scientist to explain or translate what the system actually means.

  23. 4:51

    That's the first pillar. The second pillar, adaptive control. What do we mean by that? It's about building smart guardrails. If the AI system starts to veer off, makes a wrong decision, the system should be able to slow down, change its course, or at least call a human for help.

  24. 5:07

    Think of it as a lane assist for your AI. The third pillar is always have human in the loop. What do I mean by that? This is basically setting up the rules and the playbooks so that the right experts get pinged in the right time with the right information without causing an overhead for both the system as

  25. 5:27

    well as the person, right? But all of these things are built on the bedrock foundation of traceability.

  26. 5:34

    Every data, every change is digitally signed and is trackable. Think of the concept of it like software bill of materials or even simple. Think of it like a FedEx package.

  27. 5:45

    From the time it leaves the warehouse till it reaches your doorstep, you can track every single step of it. So this was our-- these is the three pillars of a trustworthy AI.

  28. 5:55

    But with these pillars in place, the larger question is: How do we actually weave them into the AI systems we are building and running today, right? Let's look into the journey.

  29. 6:10

    So like I said, how do we make these pillars reality in day-to-day AI operations? [clears throat]

  30. 6:18

    This is where, um, we move beyond the standard MLOps and what we call it as XTOps. Think of it as an MLOps, but with built-in conscience and a direct line of human oversight.

  31. 6:32

    This diagram isn't just a flowchart. It's the blueprint for the entire life cycle of AI. Let's begin with the verifiable traceability. Right from the data stage, know where your data comes from, understand what are all the changes and how it is changing.

  32. 6:49

    No more guess works. When we train the models, we just don't train them for accuracy, right? We are embedding actionable intelligibility. What does it means is the model also learn to explain itself so that we can spot when its reason starts to drift.

  33. 7:06

    When we deploy, right, we put those adaptive cruise controls that we talk about. This is where the guardrails kicks in, automatically adjusting to new situation, new data, and pausing to look at things if they drift.

  34. 7:22

    And when we deploy the model, this is where the human AI teaming comes in, right? This is where the actual real-world feedback kicks in so that we could quickly improve the system and humans can step in when needed.

  35. 7:37

    XTOps is not about creating a sy-- is about creating a system where every AI decision has a clear why, a when, and a who, and attached to it. It is about moving from just launching an AI system to launching an AI which we can truly trust.

  36. 7:56

    So let's pause here. Right? Now, you all might be thinking, "Hey, we do MLOps day in and day out. Most of these modules that we spoke about is already there."

  37. 8:06

    So what is unique, right? What is the big difference in doing this? The challenge is adopting an XTOps is like a journey. It's not a flip of a switch.

  38. 8:16

    XTOps is also taking about all those foundational pieces that we have and giving them a serious integrated upgrade, especially for trust. I'm not going to go through all of this, but let me touch upon a couple of things.

  39. 8:30

    Let's think guardrails and policies. We have IAM policies. We have security policies, MLOps providers, everything, right? But XTOps gives you dynamic AI-aware guardrails that you can actually understand the context and block a risky AI decision.

  40. 8:45

    Let's talk about monitoring and metrics. We do have standard MLOps metrics, right? But XTOps gives you dedicated trust-specific dashboards that both your leadership and the boards can understand. Human in the feedback.

  41. 8:59

    We do have human in the loop, but it is mostly ad hoc when it comes to MLOps. XTOps, think of it it is creating a fast lane. You click to fix workflows where human can look at some of these quick changes and go back and fix it, right?

  42. 9:14

    The larger context is XTOps is not reinventing the wheel, right? It's about adding advanced safety and transpar-- transparency features needed for the high-stakes enterprise of AI.

  43. 9:27

    And what is in for us? We spend less time firefighting unpredictable AI behaviors and spend more time actually building more innovative products.

  44. 9:43

    So if you are serious about managing AI trust, we also need to measure what matters, right? So we talk about two metrics here, MTTRE and trust-adjusted risk in dollars.

  45. 9:55

    The MTTRE stands for mean time to resolve explainable errors. Fancy name, but very simple idea. It's basically the time that takes for us to fix something unexpected when it happens to how quickly can we understand the why and respond with a fix.

  46. 10:13

    The faster your MTTRE is, the team is more agile, less defects in the product, and quicker to solve the problems. Second, trust-adjusted risk in dollars.

  47. 10:25

    This idea is basically to put a price tag on what happens when the trust breaks, right? What is actually the business cost? Is it fines? Is it lost customers?

  48. 10:36

    Is it damaged reputation? Right? Or-- And if your AI system keeps failing or remains a black box, this metric makes value of the trust.

  49. 10:48

    Well, uh, I think we... Yeah. So let me again pause here. We spoke about metrics. So why obsess about all these metrics, right? This is-- We have enough of metrics in MLOps.

  50. 11:02

    We have enough of metrics, but why obsess? The challenge is this. Look at the first table, right? On an average, an MTTRE takes several months in some of these cases to even find a resolution.

  51. 11:14

    Right? Now, imagine this. Imagine your AI is making a biased decision for months. The damage escalates quickly, and sometimes it also es-escalates exponentially. Now, look at the second table.

  52. 11:28

    It actually shows the fallout, right? It is not just not one parameter.

  53. 11:33

    It starts with your direct fines. It starts with your engineering effort, regulatory scrutiny, and above all, the loss of trust and brand value of the products that we stand day in and day out for, right?

  54. 11:44

    These aren't small figures. A serious incident like a privacy bug or a bias in a credit card system could quickly escalate up to seven hundred millions of dollars, right?

  55. 11:53

    So this is why these metrics are not just about defense. These are about building resilient, reliable, and ultimately AI-powered products that the end users can trust. All said,

  56. 12:07

    and we are not talking out of thin air. So

  57. 12:11

    Sahil, uh, is gonna present a case study on a real incident and how we went about building this whole framework.

  58. 12:18

    All right, perfect. So that's the right stage for-- Let's bring it all together. Uh, I'm gonna talk about a company called Guardhat. Uh, this is a company that I used to work for.

  59. 12:28

    It's, uh, focused on, uh, worker safety. So more specifically, uh, it has an AI-driven platform that is geared towards, uh, solving worker safety problems in hazardous environments. So what we're doing here is we built, uh, we built IoT devices, wearable devices that would be worn by the workers.

  60. 12:50

    And at some point in time, uh, uh, these devices would get deployed, activated, and they will, they will collect data. They'll collect health data as well as, um, as well as environmental data.

  61. 13:04

    And then that data is sent to the back-end system where the AI analyzes this in real time. And based on that, it is able to identify when an incident-- predict when an incident is about to happen and, and, you know, you can pr-prevent that incident from happening.

  62. 13:19

    So very mission-critical application. Um, it was great because, uh, while we were saving lives in a way, um, there was, there were enormous challenges. One of the inputs to the, uh, to the AI platform was the GPS.

  63. 13:36

    And as a result, uh, seventy percent of the cases were false positives. And, um, it's easier to say this now because it's after the fact, but back then we didn't know that.

  64. 13:49

    And so w- the behavior of the user was that at that point in time, uh, the user stopped, uh, reacting to the alerts. They started ignoring the alerts, and that caused a huge safety risk, not just for the people, of course, their, their, their lines at st-- uh, lives at stake, but even for the company from the

  65. 14:10

    liability point of view that, uh, workers were not, uh, reacting to alerts. So we went back to the drawing board. Um, we started identifying the issues. Um, and, you know, if we were to do this without the XTOPS framework,

  66. 14:27

    uh, we would probably do an MTTR, mean time to resolution. If you look at it, um, you know, seventy percent of the time is spent in identifying the problem.

  67. 14:35

    Another twenty is spent in, um, finding a solution, and then you deploy it. But I think the most critical part is that there is no system to identify the GPS drift.

  68. 14:46

    We wouldn't know about it. And then because it's such a complicated model and code that it's really hard to identify what's causing that problem.

  69. 14:55

    So if you were to apply this model, um, day zero, uh, you get an alert that was ignored, um, during an incident. Day two, there is an attribution telemetry that will flag the anomaly.

  70. 15:10

    And day seven, you have a solution deployed which, uh, which fixes the GPS drift, uh, or at least finds a reroute to the GPS drift. Sorry.

  71. 15:20

    Now, having said that, um, to be real, this problem did not get solved in seven days. It took eight months for us. But it-- this was a model problem that actually helped us to build this framework.

  72. 15:34

    And once we were able to build this framework, these kind of problems can be solved in seven days. We, we tested it across our enterprise and eventually became an enterprise standard.

  73. 15:44

    So all this is great. Um, you have the impact, you have, uh, y- you, you can see the value in this. But here's the big question:

  74. 15:55

    What do... How do you convince the CIOs? Uh, how do they look at all of this and find value in this? What is the language that they talk? The answer is...

  75. 16:07

    What? [laughs]

  76. 16:09

    So you gotta convince the CIOs that this is saving money. And if you were to look at the left side of the slide, uh, you'd see that the risk exposure that we're looking at is approximately two point five million dollars per site per, per year.

  77. 16:23

    Now, some direct impact with this is that we were able to solve the, uh, with this, with this product, we were a-- or this structure, we were able to solve the fines, and we saved five hundred K in fines every year per site.

  78. 16:38

    Beyond this, some of the indirect, uh, benefit was that this system in, if it w-were to be working correctly, was supposed to prevent incidents, all of them, but it was preventing X percentage of incidences because it wasn't working correctly.

  79. 16:53

    But with this structure, it did work correctly, and then after that it wa-- you got the, uh, rel-- um, the, the remaining value as well.

  80. 17:03

    So I just want to wrap it up real quick. Um, I think the outcome, you can see a lot of value there in terms of, you know, false alerts came down, uh, trust score went up.

  81. 17:14

    That means, uh, people started using those alerts. Uh, sh- they were able to see value in those alerts. I think more important things were related to the telemetry itself, um, understanding why a particular inference was made, uh, having the control where you can, if there's a GPS drift, you, you're able to switch.

  82. 17:34

    And the most important thing, human in the loop. So if these things happen, someone is notified. We created a dashboard where someone is notified and someone is able to take action and retrain the, the, the, the model.

  83. 17:48

    So with that, um, I mean, this slide is just, uh, a high-level overview of what we presented to you today. Um, thank you for being here.

  84. 18:00

    Uh, we'll just leave it at that.

  85. 18:05

    Thank you so much for the fantastic talk on the trust gap. Um, I actually had a question for you because we have a minute or so. Um, so the phrasing that you had around like the trust ad-- trust-adjusted risk cost premium, how do you advise people to think about like reputational damage?

  86. 18:21

    Like is this something that you have thought about measuring or investigating at all?

  87. 18:25

    Uh, do you want to take it? Sure. So it's, uh, like I was saying in the beginning, these are silent failures. You cannot... It's really hard to quantify the impact, and reputational damage is, is again, right on top of that list.

  88. 18:39

    So, um, it's really hard to measure that, to be honest. But all you can do in, in this kind of a case is, you know, really you can, you can find, you know, people are creative.

  89. 18:50

    You can find ways to quantify some dollars to it. But, you know, it, it's really hard to predict. Let me-- There's no short answer to it, [laughs] let me just say that. [upbeat music]