The 6 Pillars of an Agentic Harness for Production — Varun Krovvidi, Resolve AI
Read the talk
The 6 Pillars of an Agentic Harness for Production
Varun Krovvidi explains how Resolve AI combines model routing, selective context, causal evidence, permissions, learning and evaluation—and demonstrates an investigation that resists a tempting explanation for a production failure.
From a talk by Varun Krovvidi
At a glance
Ideas worth remembering
Production investigations combine code, infrastructure, knowledge and telemetry across teams; a capable model needs a harness that manages those connections.
Model orchestration must both reassess new models and route individual tasks. Context engineering must supply enough information to explore useful paths, then narrow retrieval through precise tool calls.
A root-cause claim requires a causal chain of evidence. When the chain cannot be established, the system should lower confidence and indicate another investigative direction.
Permissions determine which proposed fixes can become actions; learning and evals determine whether later investigations improve, including their reasoning paths and confidence.
The demo’s central test is whether a diagnosis remains grounded when an engineer suggests coincident outages or deployment errors, while preserving a shared investigation for teammates.
Why writing code and running it demand different systems
Using AI during on-call is a much lower bar than letting it do most of the work. Varun Krovvidi of Resolve AI opens by asking engineers whether they have tried AI for on-call, then whether they can comfortably say it handles 90% of that work. That second question sets the ambition for the talk: agents that run and fix software, with engineers directing them.
Coding offers a comparatively accessible starting point. In Krovvidi’s framing, code describes much of its own behavior, divides into modules that a model can inspect separately, and gives the model a relatively contained domain. Those properties help an assistant find the next relevant piece and help developers define benchmarks for whether it solved a task. Production investigations cross domains: understanding a function is only part of understanding what happened when that function ran.
Krovvidi puts running and fixing software at 70% of engineering time. Treat that as his framing of the opportunity; the talk does not establish how the figure was measured. The work itself spans three distinct needs, illustrated through a healthcare analogy:
- Routine on-call and maintenance: recurring alerts and smaller problems—the Tylenol or Advil of operating software.
- Major incidents: systems shutting down and multiple responsible teams gathering to intervene—the equivalent of surgery.
- Preventive work: infrastructure reviews, cost analysis and platform engineering that keep a system healthy—the daily vitamins.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The last mile includes tokens, tools and other teams
“The era of bottomless AI is over” introduces a practical constraint: every investigation consumes tokens, and its architecture determines how efficiently it uses them. Pulling more data and making more model calls can increase cost without bringing the system closer to the right explanation. A domain-specific harness—the surrounding system that selects models, supplies evidence and controls actions—has to make those choices deliberately.
A generated image or draft can satisfy many acceptable interpretations of a request. A production investigation has a tighter target: identify the steps that produced the failure so an engineer can fix the responsible issue. Krovvidi calls reaching that target “the single longest mile.” Fluency can make an answer seem finished before the investigation has established the cause.
The work also belongs to several people and systems. SREs, platform engineers and service engineers hold different pieces of the explanation, so an agent needs to support a shared investigation. Resolve packages that work into on-call agents for everyday issues, incident agents for complex incidents and incident channels, and ambient background agents for deployment monitoring, analysis and condition-triggered investigations. All three use the same underlying architecture.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Five cracks that appear beyond a convincing demo
Pointing a powerful model at production data can produce a good demonstration for a carefully chosen use case. The harder test comes when the same system serves more teams and more kinds of problems. Resolve’s experience motivates five failure modes, each requiring a response in the surrounding system.
-
Anchoring and the model treadmill: an investigation can settle too early on a plausible explanation. Meanwhile, new models change which capabilities are available. A workflow may contain hundreds of tasks, including open-ended reasoning, deterministic steps, image interpretation, SQL and log analysis. Choosing one model for the entire workflow hides those differences.
-
Too much or too little context: abundant context can encourage exploration of unsupported hypotheses; sparse context can hide relevant paths. Logs and metrics can keep growing, and combining them with code and infrastructure makes indiscriminate collection especially expensive.
-
Coherence without causality: a model can assemble a persuasive answer or follow a user’s preferred theory without establishing the events that caused the incident. That risks producing a patch for the symptom rather than a fix for the cause.
-
Actions without guardrails: a model might select deletion as the simplest fix available to it. The operational question is what privileges the system should have in that situation, rather than whether the proposed action sounds locally reasonable.
-
Investigations without learning: different teams need the same context, and later investigations need useful knowledge from earlier ones. If the hundredth investigation cannot use what the system already learned, the organization pays for repeated discovery.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Pillars 1–2: choose the model and the next piece of context
The first pillar, model orchestration, operates at two timescales. As models evolve, evaluation determines which ones remain suitable. Within an investigation, routing matches a model to the task at hand. Krovvidi offers Gemini for image-specific reasoning, OpenAI models for deterministic steps and Claude for open-ended investigations as examples of that matching problem. These are illustrative assignments, not a permanent ranking: the engine must keep evaluating the best fit at the time.
The second pillar, context engineering, starts with a question about the investigation: what information does the model need to make this next decision? Graph RAG and knowledge graphs are possible implementation choices, but selecting one does not answer that question. A combination of techniques may be useful because the information needed to begin an investigation differs from the information needed to test a specific hypothesis.
Krovvidi sketches a progression from graph-based retrieval that helps start an investigation to precise tool calls that retrieve logs, metrics, dashboards or code as the work narrows. The mechanism controls both exploration and token use: enough initial context exposes plausible paths, then targeted queries supply evidence for the next step. A larger context window alone cannot decide which paths deserve attention.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Pillars 3–4: require a causal chain and govern the response
The third pillar, causal reasoning, sets the standard for a root-cause claim: identify the steps that happened and show how they led to the issue. A coincident event is a candidate explanation, not yet a causal link. Resolve’s stated policy is to base root causes on a chain of evidence; when the data stops short of establishing that chain, the system should return low confidence and point the user toward another investigative direction.
That policy changes what completion means. An investigation can end with a supported cause, but it can also produce a useful statement of where the evidence runs out. The latter gives an engineer a place to continue instead of making a coherent explanation carry more certainty than the available data supports.
The fourth pillar, governed actions, controls what happens after reasoning. Teams define whether the agent can read or write and, for writes, the conditions under which an action is allowed. Those rules depend on organizational access policies, restrictions and industry requirements. Least-privileged access makes the allowed action set part of the design, so a model’s proposed fix does not by itself grant permission to execute it.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Pillars 5–6: learn from interactions and evaluate the path
The fifth pillar, learning systems, treats user interactions as information about the investigation. Positive feedback, negative feedback and an engineer steering the work in a new direction provide different learning signals. The goal extends beyond remembering the final answer: later investigations should benefit from how people corrected, supported or redirected earlier ones. Krovvidi does not specify whether these signals update prompts, stored knowledge or model weights.
The sixth pillar, evals, is both the starting point and the ending point of the architecture. Resolve describes five levels of evaluation:
- Positive reinforcement: capture feedback that supports the agent’s work.
- Negative reinforcement: capture feedback that identifies a problem.
- Investigation-path assessment: trace how the agent reached its solution and score that process as engineering work.
- Expert comparison: compare the investigation with how a strong engineer would approach it.
- Confidence calibration: assess how the agent expresses confidence in its answer.
Evaluating the path matters because an apparently correct answer can still come from weak reasoning. Expert comparison asks whether the investigation took useful steps; calibration asks whether its stated confidence fits what it established. New models, new use cases and architecture changes all trigger another pass through these evaluations. Model selection and harness design therefore remain connected to observed investigative behavior.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The demo: turn an alert into an inspectable investigation
The concrete example begins with Grafana sending an alert into a shared Slack channel. Resolve picks it up automatically and returns a short investigation result: a root cause, evidence, recommendations, contributing factors and rejected theories. Opening the investigation UI moves the engineer from that summary into the work behind it.
The alert itself becomes the prompt that starts agents. Investigator agents first determine what is failing, gather metrics and traces, and use traces and logs to locate the issue. They then connect those observations with change events, deployments, code and infrastructure. The investigation draws across four source categories: code, infrastructure, knowledge bases and observability platforms.
How does information move from a notification to a diagnosis an engineer can question? The diagram shows the alert initiating work, the four source categories supplying evidence, and the result returning both an explanation and alternatives. The important relationship is that the Slack notification starts the investigation; it does not contain everything needed to explain the failure.
The reported diagnosis traces the high failure rate to stale integrations that remain live in the system. A GCP outage happened at the same time, and traffic spikes offered another plausible direction to investigate. Resolve presents these as alternatives it ruled out. The demo illustrates a diagnosis and its interrogation; the spoken explanation does not expose the individual log records or intermediate causal links, and it does not show a remediation being executed.
Resolve automatically picks up the alert and treats it as the investigation prompt.
The alert starts investigator agents. Evidence comes from several systems, and the result retains rejected theories for engineers to question.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Hold the explanation against pressure, then bring in the team
The next observable change is in the engineer’s interaction with the result. Instead of accepting the diagnosis, Krovvidi highlights material and asks whether the issue relates to the GCP outage. He also describes pushing Resolve about deployment errors that occurred at the same time. These challenges revisit tempting correlations after the system has already selected stale integrations as the cause.
Resolve’s intended response is to keep its answer tied to the established causal evidence rather than follow the user’s preferred explanation. This completes the example’s progression: an alert triggers investigation, investigators assemble information across systems, a diagnosis includes rejected alternatives, and an engineer can challenge those alternatives without merely prompting agreement. The value of the interaction is that disagreement becomes a request to examine evidence.
The same investigation can bring a teammate in directly from Slack, creating a shared virtual war room. That returns to the earlier multiplayer requirement: the explanation, supporting context and questions need to travel with the work when another engineer joins. Collaboration belongs inside the investigation rather than requiring everyone to reconstruct it independently.
The closing distinction is architectural. These six pillars address how a general-purpose model becomes useful for a specific operational problem. Putting that system in front of customers also requires a dependable platform to run it. The talk develops the investigative harness; ordinary product and platform engineering remain a separate axis of the work. A useful root-cause investigation is one part of shipping an agent that people can rely on.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
Related talks
- Beating RL With Reflection: GEPA and Optimize Anything
Develops the learning-and-evaluation side of agent harnesses: execution traces guide changes to prompts and programs, while human feedback can shape the evaluator used to improve the agent.
Read the complete timestamped transcript
- 0:12
Uh, can everyone hear me okay?
- 0:16
Awesome. Thank you so much, uh, for attending the session. Um, my name is Varun. I'm part of the team at Resolve AI. We are building AI for prod. Uh, basically, we're building agents to run and fix your software. So in this session, we'll cover what are the six pillars of an agentic harness that you'll need in a system that is there to fix and run your software. Uh, before I get into the session, just a quick show of hands just to understand-- for me to understand who's there in the audience. I'm
- 0:46
assuming most of you are engineers. Uh, how many of you have been on call before? Uh, quick show of hands. Awesome. Have you tried to use AI to help you with your on-call? Show of hands. Awesome. For the people who raised their hands, uh, can you tell me that, uh, ninety percent of your on-call work is being done by AI? Uh, can you comfortably say that? Awesome. Yes, that's what we realized when we started the company as well. This is a really hard
- 1:16
problem to solve. So, and in this session, I wanna share, like, our journey, what is the evolution we went through, what are some of the learnings we had, specifically in building such an agentic harness. Before I get in, um, I wanna set context on, uh, what we are experiencing today. So the first wave of AI really was focused on coding, and, uh, truthfully, I think most of the-- most-- for most of us, the way we code has fundamentally changed. And I also wanted to share
- 1:46
context on why that is the case. Uh, firstly, code in itself is self-documenting. It's easy for AI to actually parse and understand what's there to help you out with the next step. Second, inherently, code is very modular, so it's easy for, uh, for AI to actually break it down as chunks and navigate to the next step and also help you extend and build and bey-- build beyond it. Third, and the most important thing, is code is single domain. Um, you do not need to teach
- 2:16
AI multiple domains to reason across them. It's easily able to navigate to the next step, and that is one of the reasons why we have such defined benchmarks as well. You can objectively assess how good is AI in solving this problem for you. But in reality, for engineering, seventy percent of our time is spent on the other side, not generating new software, but rather in pr- in production systems, basically running and fixing software. Uh, and this ranges through a
- 2:45
broad range of issues as well, right? Like operating and running software. Typically, you can actually put it in three broad buckets. Like, one is, uh, think of the healthcare analogy. That's the easiest way I like to think about it as well. One is the regular on-call and maintenance. Think of this as regular small problems that are coming your way. Maybe you need a Tylenol for, uh, some, uh, for some pain that you have. Maybe you need an Advil. So these are the regular on-call style alerts that you need to
- 3:15
deal with. Second category is the incidents, where you need all hands on deck. It's equivalent to a surgery. Systems are shutting down. You need everyone's attention who's, uh, who's involved or responsible, uh, to come in and fix it. And the third category is more like your daily vitamins. So you need to do a lot of things on a regular basis to keep-- to maintain the health of your system going on. So be it things like taking a look at your infrastructure, taking a look at your costs, or maybe some platform engineering
- 3:45
work that you'll do to scale. So these things are fundamentally hard because they're not exactly like code. They cut across multiple different domains, multiple different tools, and that is the repercussions that we are starting to see when the euphoria of initial-- initial euphoria of AI is starting to settle in. And that's what you're seeing in the news right now. So there's a lot more talk about the number of issues maybe AI-generated code is starting to create in production, and there is talk around what are the right frameworks to deal with these, uh, issues in code.
- 4:15
Second, uh, the era of bottomless AI is over. Like, the unlimited AI is no longer there. You can increasingly see conversations about tokens, token optimization, token efficiency. How do you basically improve those architectures, those rigorous engineering architectures that you need to efficiently and precisely use AI? And third, uh, this is the conversation that is, uh, probably made popular by Satya Nadella, who's the CEO of Microsoft, where he started talking about the architecture you
- 4:45
need on top of your regular models for any domain-specific AI to work. So there is a broad classification, right? Like, whenever you look at a generation tasks, you don't have an output or an outcome in mind. Uh, you don't want-- You're not specifying to AI exactly how an image should look like, how a blog should look like, or how a website should look like. So that's where the euphoria sets in. But when you think of a flip side of a task where you have a single correct answer that you want AI to get to, that is when the disappointment starts to set in,
- 5:15
which, uh, in most cases, uh, most people dub it as a last mile, but that is the single longest mile that you have to run with AI. That is why you, uh, you need these architectures for you to focus on a specific problem. Now, beyond all this, uh, operating production systems is also a multiplayer and a multi-system problem. You're not just doing it on your own. You need to pull in SMEs for, for different teams. Over time, uh, all of us have specialized in different categories of engineering. We are put in different teams. Uh, we need to talk to,
- 5:46
to solve-- to resolve any issues, uh, across AI. So, sorry, across your production systems. All in all, uh, our company, Resolve AI, started with a single thesis. So we genuinely believe that all of this, uh, all of this work that you do on production systems like fixing issues or like, uh, running your software on a day-to-day basis should be done by agents in most part. That's one of the reasons why I was asking, uh, can you confidently tell ninety percent of your work or more is being done by agents? And engineers should just be running those
- 6:16
agents where agents take care of like fixing and running your software. So that's why we've built, uh, Resolve AI to-- as three categories of agents. There are on-call agents to fix your software issues that come in on a daily basis, and there are incident agents that'll help you drive incident channels and get you of-- get, get you from a complex incident to a root cause and fix. And lastly, there are ambient background agents that will help you with your production tasks, be it things like deployment monitoring or any of the analysis that you have to run,
- 6:46
or any of the conditions that you trigger where you want a specific investigation to happen. Underneath all this is the same Resolve AI agent architecture. That is what we'll be talking about today, uh, and, and what it takes to get there to build it. So let's start with a question. So why is this even required, right? Like, why can't we just point the biggest model at production data, models are, of course, getting better, the reasoning capacity is getting better, uh, to fix your software issues? It ge-it genuinely demos well. Whenever you focus on a particular use
- 7:16
case and you build your harness on top of it, it shows up with a very good demo. But slowly, as you start to scale across use cases and teams, that is where, uh, some-- that is where very familiar cracks start to appear in the system. So what are the kind of cracks that we have seen as you start to scale in the product ac-along the way, right? First thing, you'll notice that, of course, models will have anchoring bias. This is a very strong thing that you have to deal in your harnesses itself. They're, they're designed to give you like a coherent answer,
- 7:46
um, and first-- as in, as in when new models start to come out with better reasoning capabilities, you also need to maintain the treadmill. Uh, you also need to continuously change, uh, which-- evaluate-- firstly, evaluate which model is suitable for the task and also update, uh, the right model for your use case. And second, and a little bit of an overlooked item. So we always associate a model to a workflow or an outcome. But every workflow has hundreds of different types of tasks. So
- 8:16
there are reasoning tasks, there are deterministic tasks that you go through. There are things like image gen-- image reasoning or you're reasoning across images, you're reasoning across SQL or things like logs. Now, all the models, frontier models that you see around, each of them is good at a specific parti-- or a specific, uh, kind of task. So the orchestration engine that you need to imagine for, uh, moving across these models for a workflow, first of all, needs to keep pace with how models are evolving. And second,
- 8:46
it needs to be able to marry the best model for the best task. And second failure mode, uh, is, of course, the context windows. I know the context windows are expanding, but context windows, context windows are not just the problem, uh, for, for solving su- for, for solving a complex problem like this. The second part is, if you provide large enough context, models start to overexplore. It'll start to come up with hypothesis that don't even exist or sometimes might not even get you to the right answer. Now, if
- 9:16
you supply very limited context, of course, they underexplore. They won't be able to get to-- uh, they won't actually have visibility into what other pathways exist or probably what other potential hypothesis can exist in solving a problem, especially for production incidents. This becomes a huge issue because telemetry is literally infinite. You can start to create as many log lines as you want, as many metrics as you want. And now when you start to marry that with your code and infrastructure, that is when real st- real cracks in system, uh, starts to appear.
- 9:47
Third, and my, and my most favorite one is, uh, defining causal reasoning. So models are there to please you. I mean, we've all come across these use cases where you push the model hard enough, uh, it'll start to agree with you in every different direction. So models are specifically designed to give you coherent answers, but not causality. But production incidents, the only thing you're looking for is a causal chain of evidence, like a detective. Uh, what are the steps that took place that led to a particular incident so that you can actually
- 10:17
fix, uh, the right issue instead of creating a patch. So especially when you're solving a hard problem like production reasoning, you need causal reasoning built into the mod- built into the system, not just coherent answers. And the next one is, of course, guardrails. I won't harp on this topic too much, but all of us have read, uh, some news or the other where, let's say, AI is going and deleting a particular file system or like a database. Uh, so AI, it's-- uh, also this is a feature, right? Not a bug.
- 10:47
Uh, AI might have determined that the cleanest possible fix is just deletion of that particular code snippet. You can't fault it for it. You just have to build guardrails, uh, around the system so that it works in the right fashion. What is the least privileged access that AI can operate with in any given scenario? And lastly, learning loops. Any system that you build should be scalable across your team or your organization in the best case. Every time you're using AI for production incidents, like I said at the starting, it is a
- 11:17
multiplayer problem. You have multiple different teams, like SREs, platform teams, uh, backend engineers or service engineers. All of these people should be involved. And most importantly, all of these people should have the same context. Now, when you add temporal discussions on top of it, like, uh, your AI system should have the same context of a previous investigation in the hundredth investigation, or else you're starting the wheel from the same time again. So these are the failure modes where we've seen the cracks appear. So that's why we've
- 11:47
actually built an agent architecture that answers these same questions exactly. So there are six major parts, uh, for any AI architecture that has to work in a specific domain. The reason I'm mentioning it as a specific domain is models are there for general purpose reasoning. Now, when you want to harness these models into a specific answer, you need to take care of six different pillars. The first one is model orchestration. Like I mentioned, uh,
- 12:17
model orchestration is two layers. One, how do you keep up with the model treadmill with the latest models? And second, how do you match the best model for the task? Is it Gemini for image, image-specific reasoning, or is it OpenAI for deterministic steps, or is it Claude for open-ended investigations? Uh, you need to constantly keep evaluating what is the best, uh, possible model for the task at that point in time. So this is an engine that we built, uh, upfront very early. Second is context engineering. Now,
- 12:47
context engineering is often minimized to like, uh, the database or an execution solution. Oh, are you using a graph RAG, or are you using a knowledge graph? Are you using X, Y, and Z? Those are all implementation details. But what matters as a goal is what is the precise amount of context AI needs to solve that specific problem. And more often than-- most often than not, it's a combination of techniques. You might have to use a graph RAG kind of a solution for AI to start investigations. And when you actually further
- 13:17
go into the next steps, you need to determi-- you need to define very precise tool calls so that you don't burn through tokens very quickly when you're making multiple tool calls in, in form of log queries, in form of, in form of metrics or dashboards or code queries, et cetera. The third one is causal reasoning. Like I mentioned, so this is one of the first, uh, principles that we've built into our system at Resolve AI, where, uh, the eviden-- the root cause is always based on a causal chain of evidence.
- 13:47
What are the exact or precise steps that happened that led to this particular issue? If you're not able to establish that chain of evidence, we have to give out that information with a low confidence level and point the user in a different direction. This is true for any domain-specific AI, and this is true for Resolve AI as well, where we are pointing you in a different direction if we don't see the-- if we don't see the data beyond a particular point. The next one, like I said, is governed actions. This is, uh, quite straightforward. Uh, it depends upon
- 14:17
every team, every different organization, uh, depending on, uh, depending on the access levels or probably the restrictions that you have or the industry you operate in. What are those specific guardrails that you need to define for AI to operate under? Is it read? Is it write? And under write, what are the specific conditions for write? Uh, et cetera, et cetera. And lastly, a learning system. So e-- think of it like this, every interaction that, uh, you're doing with an AI system is an opportunity for learning. So, uh, AI system has to learn not just
- 14:47
from an investigation it is going through, but from how, from how a user is actually interacting with it. Like, am I giving you positive reinforcement? So is this turning into a positive eval, or am I giving you some negative reinforcement? Am I g-- am I guiding you or steering the investigation in a specific direction? Those are different kinds of evals. So that leads into the last piece, which-- where evals is also a very critical step. So this is the starting point and ending point for any agent architecture. So the way we think about evals is in five
- 15:17
different levels, like I mentioned. What is the positive reinforcement you can give? What is the negative reinforcement you can give? And can we trace your path? How did you come up with the solution? Can we score you like an engineer based on that? Uh, how would the b-- how would a best engineer perform this investigation, and how do you compare against it? And on top of it, how are you calibrating yourself as, as AI? How are you giving out that confidence level? So all these are systematic eval interes-- uh, uh, evals that you build into one platform that you constantly have
- 15:47
to push the architecture against whenever there's a new model, whenever there's a new use case, or whenever there's an architecture change.
- 15:55
Cool. This is enough talking, so let me actually show you this in action on how it works in Resolve. Like I mentioned, our platform, uh, is specifically there. Uh, so there are two categories of agents that we work with. One is an on-call agent, one is an incident agent, and of course, the background agent. So, uh, since, like I've, uh, since most of you are engineers over here, you might be familiar with the system. Of course, uh, most of your telemetry or observability, uh, alerts are firing into a collaboration tool like Slack or Microsoft Teams or whatever you might be using. In this
- 16:25
case, what we are seeing is Grafana fire-firing alerts into our common Slack channel. Um, so Resolve automatically picks up those alerts and kickstarts an investigation. What you see here is, uh, it is throwing an error about a log scale, and it starts to provide, uh, what is the, what are the, uh, what is the root cause that it found for the, uh, alert that came in, and also the other contributing factors or other theories it rejected. It gave a quick root cause, uh, and evidence and recommendations. But let's
- 16:55
actually deep dive. So in most cases, like I said, this should be a multiplayer experience. So let's actually go into the UI and see what's happening here behind the scenes. So, uh, it-- the work starts off from Resolve AI when it treats the alert that's coming in as a prompt in itself. So it takes the prompt, and it spins up a bunch of different agents, actually two different categories of agents, uh, to get the work done. So first category of agents are basically investigators. So they are, uh, trained to actually, uh,
- 17:25
go through the investigation just like how an engineer would. Firstly, find out what is the issue, gather more information from metrics and traces, et cetera, and figure out where the issue is happening from traces, logs, and now start to correlate that with things like change events, uh, or any of the new deployments that would have gone in, and your code and your infrastructure. So primarily, it's doing this translation or reasoning across four types of data sources: your code, your infrastructure, your knowledge bases, and your observability platforms.
- 17:55
Based on all that, it starts to give out this, uh, root cause or, uh, this root cause what we've seen. There are m-- So one of the log scale is actually experiencing high f-high failure rate. It trace the issue all the way down to some stale integrations that are still live in the system. The interesting part is it's also showing out some other ruled out theories where generally we would start an investigation. At the same time, there was a GCP outage. So, uh, you would actually, uh, this, this would be a perfect correlation
- 18:25
as well, right? Going back to AI that is designed to give you coherent answers. So a GCP outage happened at the same time, so of course, this might be the issue, uh, that we are seeing. Or there are other traffic spikes that it is experiencing, uh, without missing integration, so, uh, that would be the other thread that you would start to pull on. So let's say if I'm a new engineer, if I'm starting in, okay, Resolve, uh, gave, uh, this as a root cause. Of course, it's-- We have also designed to go operate with the least trust policy. So let's say if I, I can start
- 18:55
to highlight, and it'll start to-- it'll pull me in so that I can interrogate this further. Uh, in this case, let's say, uh, let's ask Resolve, uh, "Is this related to, uh, GCP outage?" So it's giving out the answer here, but there is other, other precise questions that I ran I want to share with you. So there was other deployment, uh, errors that happened at the same time, so that is what I was pushing Resolve on. Going back to what
- 19:25
I mentioned, if you push, uh, an AI hard enough, it'll start to agree with you. We've designed the system to go against it. It'll only base its answers based on the causal, uh, causal chain of evidence that it created. Since it is able to establish that, it's able to give out a very coherent answer. So this, uh, also lets me actually pull in my teammates so that this truly becomes like a multiplayer experience. Let's say if I give one, uh, I can pull in Ranvir. Can you take a look at this?
- 19:55
This can pull in my teammate directly from Slack so that we can collaborate on this virtual war room experience.
- 20:02
This was a quick demo. Like I said, uh, the-- currently... And what I talked about today is just one, uh, axis of the architecture, right? Like, now there is the other axis of the architecture where you just need to productize an AI. Like, if you want to put this in, uh, in front of, uh, any user or a customer, you need to make sure it runs on a robust platform. Those are the other aspects that you have to think through, but this is something that we are al-always familiar with. So that's been my time and, uh, if you're interested to
- 20:32
learn more about, uh, about Resolve AI and which customers are currently using us and how are they using it, uh, please check us out at the booth L twenty-eight, I believe. Uh, and by the way, before you leave, take out those cool bags, uh, uh, that are right outside. Thank you so much.