AI Engineer World's Fair 2025
2025 is the Year of Evals! Just like 2024, and 2023, and …
Read the talk
Why agents make evaluation a business concern
As AI moves from supplying model outputs to taking actions, evaluation must connect system behavior to business consequences—and its judges need validation too.
From a talk by John Dickerson
Monitoring starts with something you can measure
How can you monitor an AI application if you cannot measure its behavior? For John Dickerson, that is the common foundation of monitoring and evaluation: observability depends on measurements, and evaluation supplies ways to make them.
Dickerson approaches the problem from six years as co-founder and chief scientist of Arthur AI, spanning observability, evaluation and security across traditional machine learning, generative AI and agents. At the time of this talk, he had become CEO of Mozilla.ai, whose mission was to support open-source AI tooling and give the open-source community a voice in AI’s direction. Here, however, his subject is the evaluation market—and his prediction that its long-anticipated growth is finally arriving.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Executive attention, constrained budgets and agents that act
Dickerson identifies three pressures that brought evaluation beyond the technical organization and onto the executive agenda:
- Accessible capabilities. ChatGPT gave CEOs, CFOs and chief information security officers a direct experience of AI. They did not need to understand its internals to see possible applications.
- Constrained budgets. In his account, late-2022 recession fears froze ordinary enterprise spending just as leaders were setting budgets for the following year. Discretionary exceptions remained possible, and generative AI became an executive-sponsored exception.
- Systems that act. AI was beginning to act for people and teams, rather than merely supply inputs to another application. That change made its behavior more consequential—and harder for decision-makers to overlook.
Together, he argues, these pressures explain the growing demand seen by evaluation vendors.
An agent can make decisions and carry out several steps toward an action, either autonomously or with human involvement. Dickerson explicitly retains people in the loop as a good choice for many systems. But deployment across enterprises, smaller businesses and personal projects means evaluation is already an operational concern, not something to postpone until full autonomy arrives.
The contrast with earlier ML applications is organizational as well as technical. A model might produce a number that a larger application consumes. The application’s complexity then obscures the model’s contribution: technical teams know the output matters, but the evaluation problem remains inside the CIO’s organization. An agent acting on someone’s behalf makes that responsibility much more visible.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Why established monitoring struggled for priority
Before ChatGPT’s November 30, 2022 launch, data-science teams already used statistical methods to monitor ML systems. The difficulty was connecting those measurements to downstream business KPIs: dollars saved, dollars earned or another outcome that could justify a purchase. In Dickerson’s sales experience, executive enthusiasm about AI’s return on investment often left the actual buying conversation concentrated in the CIO’s organization.
This was an established market. He traces an early generation through H2O, Algorithmia and Seldon; specialist vendors through companies such as WhyLabs, Arize, Arthur, Galileo and Fiddler; and broader platforms through Snowflake, Databricks, Datadog and the major cloud providers. Later entrants such as Braintrust extended that landscape. The missing ingredient was not awareness among practitioners.
In purchasing discussions, model behavior could still lose to security, latency or other familiar technology concerns. Buyers acknowledged that monitoring mattered without making it the immediate priority.
Dickerson recalls a recurring pitch-deck prediction from the mid-2010s onward: this would be the year a CEO lost their job over an ML failure. He says he knew of no such case. That observation captures the gap between a technically credible risk and one that senior executives personally felt accountable for.
He illustrates the scale question with Jamie Dimon’s letter in JPMorgan Chase’s 2021 annual report, released in April 2022. Dickerson describes its consumer-business example as $100 million invested in AI/ML from 2017 through 2021, which he considers small relative to JPMorgan’s size.
Editor: The cited amount covers consumer fraud-risk systems using AI, ML and other technology initiatives; it is not JPMorgan’s total AI spending. The letter separately describes annual AI investment in the hundreds of millions of dollars. That broader figure materially qualifies the spending comparison. See the 2021 shareholder letter.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A budget freeze concentrated attention
The next part of Dickerson’s account begins with enterprise budget setting in October and November 2022. He connects recession fears to frozen or reduced IT budgets for 2023. Without ChatGPT’s arrival, he thinks those constraints would have left much less room for new IT projects.
The combination was bittersweet: ordinary spending was constrained, but a small discretionary project could attract extraordinary attention. His image for that concentration is the Eye of Sauron turning toward one project. Generative AI became that project.
The timing mattered. ChatGPT arrived just before the holiday breaks, with a single web interface that let nontechnical executives experience the technology themselves. Dickerson recalls Arthur and its Series A investor, Index Ventures, hosting an event at NeurIPS on launch night. Even that room of machine-learning specialists became absorbed in trying it.
Playful writing requests made the capabilities accessible beyond that audience: have Eminem rap like Taylor Swift, or ask for poetry in a rapper’s voice. These were demonstrations people could understand immediately, without a technical presentation.
In Dickerson’s account, that experience unlocked discretionary money for CEO-sponsored generative AI work. The resulting 2023 enterprise experiments grew inside a wider environment of austerity. This is his explanation of the adoption pattern he observed, rather than a claim that every enterprise allocated its budget the same way.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Experiments became production systems
By 2024, Dickerson was seeing generative AI applications enter production, primarily for internal use. Examples included chat applications and hiring tools. Deployment brought a different set of stakeholders into the conversation, asking about return on investment, governance, risk, compliance and brand consequences.
Those questions require more than an impressive demonstration. A quantitative estimate of risk requires evaluation. Once an application is used in real work, its behavior becomes relevant to people outside the machine-learning team and the CIO’s office.
His sequence is experiments in 2023, production in 2024, then shipping and scaling in 2025. Later budgets could earmark substantial spending for AI instead of relying on exceptions to a freeze. He reads rising usage and provider revenues as signs of that transition.
Improving models, open-source participation and investment from venture capital and large technology companies reinforce the momentum in his account. But the third pressure remains essential: systems are moving toward autonomy. The evaluation demand comes not just from more AI usage, but from AI taking on more responsibility.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Actions connect system behavior to business outcomes
Dickerson’s short description of an agent includes perceiving an environment, learning, abstracting and generalizing, then reasoning and acting. He contrasts that reasoning-and-action capability with the earlier picture of a model supplying an output to a larger application. The action might occur in a virtual environment or in a cyber-physical system, where software interacts with the physical world.
That expands both complexity and risk. Evaluation therefore needs a connection to the outcomes a buyer cares about: mitigating risk, increasing revenue or reducing losses. An isolated model score is not the end of that explanation. The measurement must help establish what the system’s behavior means for the business.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Different executives need different evidence
Executive alignment does not mean everyone wants the same measurement. Dickerson maps evaluation to several distinct responsibilities:
| Role | What evaluation needs to support |
|---|---|
| CEO | Understanding capabilities well enough to allocate budget and discuss AI with boards and shareholders. Knowing how attention works inside a model is not a prerequisite for those decisions. |
| CFO | Estimating bottom-line effects and supplying quantitative inputs for allocation and budget planning—the numbers that eventually enter an Excel spreadsheet. |
| CISO | Assessing security risks and testing defenses. Dickerson reports that security buyers can make smaller purchases faster and with less process than CIO organizations. Hallucination detection and prompt-injection defenses had already brought guardrail vendors into these conversations before the agent wave. |
| CIO | Keeping the organization’s systems operating. This buyer was already engaged with ML tooling and continues to carry operational responsibility. |
| CTO | Making engineering decisions from numerical evidence and standards. Dickerson mentions OTel, or OpenTelemetry, as an example in this context. |
The market thesis depends on these budget-owning roles arriving at a shared need to understand AI behavior, even though their reasons and purchasing processes differ.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Monitor the whole system—and test the market prediction
As vendors move into agentic and multi-agent monitoring, Dickerson emphasizes the scope of the measurement: monitor the whole system, not only the model used by one particular agent. The system is what performs the work. Measuring one component does not cover the behavior of the larger process in which it participates.
He also offers a future check on his growth prediction. A mid-April article in The Information had reported leaked revenue figures for evaluation startups, including Weights & Biases, Galileo and Braintrust. Dickerson says those figures lagged by roughly six to eight months. Based on conversations with peers, he expected newer figures to show a stronger business picture.
His proposed test was to look at reporting in early 2026 and see whether evaluation-startup revenue still appeared to lag. That remains a forecast within the talk: executive attention, budgets and agent deployment are his reasons to expect growth, not a substitute for observing it.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A common interface for trying agent frameworks
Before the questions, Dickerson returns briefly to Mozilla.ai. He distinguishes it from the evaluation vendors discussed throughout the talk and introduces any-agent, presented at the time as an open-source, non-monetized project.
His analogy is LiteLLM for multi-agent systems: a unified interface over multiple agent frameworks. The suggested use is straightforward—try different frameworks through that shared interface. The coda concerns making framework experimentation easier, rather than prescribing a deployment architecture.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Who can judge a financial agent’s work?
An audience question makes the evaluation problem concrete: suppose a multi-agent system produces a discounted cash flow spreadsheet, a valuation based on discounting future cash flows to their present value. Who determines whether it did the work correctly?
The question is about domain expertise. Familiar evaluation methods for structured ML outputs do not automatically settle quality judgments when a workflow interprets unstructured information and performs specialist work. Producing something that resembles a human analyst’s output is not sufficient evidence of correctness.
Dickerson first points to the limits of persona-based agents, invoking his Nature Machine Intelligence work on the risks of using language models as substitutes for human participants. Prompting a model to act as a farmer in Ohio in their mid-forties may have value, but it does not perfectly reproduce that person’s perspective. A role description cannot simply stand in for the expertise or experience the evaluation needs.
His practical answer is to put experts alongside the system. Dickerson recalls a leaked Mercor spreadsheet listing expert rates around $50, $100 and $200 per hour, and offers Google, Meta and large banks as examples of customers. The point of the anecdote is the willingness to pay for expensive human validation in lockstep with agent work.
He likens the agent to an intern who may eventually change the expert’s job, then backs away from a simple prediction of replacement. For the immediate evaluation decision, the stakes matter: when an incorrect result could cause substantial financial losses or serious professional consequences, costly expert review can be justified. What happens over the following five years, as those expert judgments become data incorporated into the systems, remains an open question in his answer.
That investment can also create competitive value. Dickerson places particular weight on dataset creation and environment creation: building the task-specific material and setting in which a system can be evaluated. A strong DCF evaluation environment can become an asset that competitors lack. In this view, expert validation is not only an operating expense for checking today’s outputs; building the evaluation capability can be a capital investment.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
LLM judges still need validation
The final question asks when evaluations will be driven primarily by generative AI or language models. Dickerson does not give a date. Instead, he points to a practice already in use: LLM-as-a-judge, where a language model assesses another system’s output. Adoption is happening despite limitations in the judgments.
He refers to his ICLR 2025 work on differences between LLM and human judgments, Style Outweighs Substance: Failure Modes of LLM Judges in Alignment Benchmarking, while discussing qualities such as conciseness and helpfulness.
Editor: In the ICLR 2025 study, the tested judges favored stylistic features, and their preferences did not track the study’s concrete safety, knowledge and instruction-following measures. That finding concerns the evaluated judges and benchmark setting, rather than every judge or task.
Persona prompting offers an imperfect substitute for a human judge and can partly ease the problem of creating evaluation data. Dickerson describes it as a practical crutch: it makes judgments easier to obtain, but does not establish that they are reliable.
That brings the answer back to the domain-expertise question. The judge’s outputs still need validation against human judgment, with attention to whether the evaluation is moving in a biased direction. Automating the assessment does not remove the need to check the assessor.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
From the talk
The primary document behind the banking-investment example; its$100million figure concerns consumer fraud-risk AI/ML and other technology initiatives.
Mozilla.ai's unified interface for trying and evaluating different agent frameworks, which Dickerson recommends at the end.
Dickerson's Nature Machine Intelligence coauthored study matches his Q&A caution about replacing domain or demographic expertise with prompted personas.
Dickerson's ICLR2025 coauthored study matches his closing warning that LLM judges can reward stylistic properties over substantive quality.
Read the complete timestamped transcript
- 0:00
[on-hold electronic music] Thanks everyone for being here.
- 0:16
Um, I'm gonna give this talk mostly from the, uh, point of view of being a co-founder and chief scientist at Arthur AI where, for six years prior to joining Mozilla AI as CEO.
- 0:27
Um, I do wanna say Mozilla AI operates in the open source world, where we're providing open source AI tooling, and we're supporting the open source AI stack. Our end goal is to, uh, enable the open source community to be at the same table as a Sam Altman when talking about AI moving forward.
- 0:41
So if you're interested in that, that's not what this talk is gonna be about, but we can talk about that offline. This talk is gonna be about 2025 finally being the year of the evals.
- 0:50
And as was written, uh, written, as was spoken, uh, by my introduction, I've been in this space for a very long time. Arthur AI, for example, was in, and it still is in observ-observability, evaluation, and security in both traditional ML and AI, and then into the deep learning revolution, and then into the GenAI revolution, and then into
- 1:08
the agentic revolution. And I think we're finally at the point where, uh, all of these companies are going to start seeing hockey stick growth, which is exciting.
- 1:16
So the thesis of this talk, one thing is that, uh, I, I see AI/ML monitoring and evaluation as two sides of the same sword or ruler, right? You can't do monitoring or observability without being able to measure, and measurement is, uh, the core functionality for evaluation.
- 1:32
This was not, uh, top of mind really with the C-suite, uh, until two things happened concurrently. One is AI became a thing that people who aren't a CIO or CTO could understand.
- 1:45
So the CEOs, the CFOs, the CISOs began to understand it basically when ChatGPT came out. And simultaneously, there was a perfectly timed budget freeze across enterprise, at least in the US, that happened due to a fear of an impending recession.
- 2:01
So this is right before ChatGPT launched. This is like October, November, when most enterprises would set a budget for the next year. At that time, there was a freeze except for money that could be opened up for a specific pet project, and that pet project, because CEOs and CFOs then knew about it, uh, was GenAI.
- 2:19
So that happened, and then now we have the final sort of, uh, vertex on this triangle, which is going to force evaluation to be top of mind this year, which is that we have systems that are now acting for, uh, humans, acting for teams, as opposed to just providing inputs into larger systems.
- 2:37
So these three things together are, as you saw with Braintrust, as we're seeing at Arthur, as you're seeing with, like, Arize AI or Galileo, other big players in this space, we're starting to see big takeoff because of this.
- 2:50
Cool. So right now this is what's happening. Year of the agent, we all hear about it. We're hearing about it here at this conference. Agents are starting to make decisions and take actions, complex steps that lead toward an action, either autonomously or semi-autonomously.
- 3:05
Uh, as a question on the last talk, um, um, uh, brought up, you know, bringing humans into the loop is still obviously a very good idea on many systems, but we're getting closer and closer to full automation.
- 3:15
And agentic systems are going into deployment now, okay? And that's in enterprises, that's in SMBs, that's in pet projects, and so on. And what that means is, by that last slide, this is also the year that we need eval.
- 3:29
Flash back, uh, to then, up to like a year ago, where every year we were asking, "Hey, is this the year of ML? Is this the year of eval- of the eval- eval- evaluations?"
- 3:39
And prior to sort of these agentic systems coming out, we would have machine learning models basically spitting out numbers that would then be ingested into a more complex system, and that complexity, uh, would sort of erase, uh, the top-of-mind need to think about what's coming out of the model itself.
- 3:55
Except for for people in this audience, we know that's very important. But when it comes to decision-makers, it would often get wrapped up into this sort of opaque box.
- 4:02
And that meant that the ML part didn't really bubble up beyond, like, the CIO or whatever org the system was going into. Which meant that typically, that year was not the year for eval because the need for evals was not obvious to the entire C-suite.
- 4:17
Cool. So let's take a quick step back in time. Before November 30th of 2022, the ChatGPT launch, uh, ML monitoring was certainly a thing, right? Like, data science teams have long used statistical methods as part of larger systems to understand what's going on, right?
- 4:30
This is, this is core. Um, like I mentioned though, there was a tenuous connection to sort of downstream business KPIs, and at the end of the day, that's what gets your product bought in the enterprise, is being able to make a sell about dollars saved or about dollars earned.
- 4:43
So being able to connect machine learning, the components specifically to a downstream business KPI.
- 4:50
There was a lot of lip service around AI/ML, around the ROI from the C-suite, including the CIO-- CEOs, um, but that was just lip service i-in our experience at least.
- 5:00
Uh, it was still basically selling into the CIO. So basically, it made it hard to sell outside of that. Now, obviously, this is a, a large space. This has been happening since, you know, about 2012, I would say, is when AI/ML monitoring really started up with, like, H2O and Algorithmia and Seldon, sort of the first generation of
- 5:16
these companies coming around. WhyLabs, Aporia, Arize, uh, Arthur, Galileo, Fiddler, Protect AI, and so on and so on and so on. I put the cutoff here at, uh, like mid-2022, sort of like before the GenAI revolution happened.
- 5:31
There have obviously been companies founded after that. You know, we just saw Braintrust talking as well. And then, you know, the big players here as well, right? Snowflake, Databricks, Datadog, SageMaker, Vertex, you know, Microsoft's products, and so on.
- 5:42
So people have been thinking about it, but it was never the thing. Again, rarely top of mind for the CEO, the CFO, and the CISO. It's never the issue.
- 5:52
So when we would talk to people, it's always, "Yes, we understand that we need this, but security's gonna be a bigger issue," or, "Latency's gonna be a bigger issue," or some of these more traditional technology sort of problems are gonna be the issue.
- 6:03
It wasn't the machine learning model itself. So basically, you know, I do a lot of due diligence for venture capitalists as well now as a, a, you know, um, a, a multi-time founder and so on.
- 6:13
And basically every pitch deck in this space from, like, the mid-2010s onward had a slide that said, "This is the year that a CEO is gonna get fired because of an ML-related screw-up."
- 6:23
Uh, and to my knowledge, it just still hasn't happened. There have been some forward-thinking leaders. So I have here a, um, the annual report from Jamie Dimon, uh, head of JPMC.
- 6:33
Uh, this came out in April of 2022, so it covered basically JPMC up through their fiscal year in 2021. Uh, and he's talking about the spend that they have going into AI.
- 6:41
But if you squint and you look at these numbers, they're still, like, comically small. So basically, one of these is in the consumer world, he makes the statement that from 2017 up through the end of 2021, they had put $100 million into AI/ML, right?
- 6:53
That's not a huge amount of money, uh, for JPMC.
- 6:58
So keep sitting back in time, pre-ChatGPT and so on, but let's now flip to the macroeconomic side of things.
- 7:06
So the economy started getting pretty dicey, um, right up until about ChatGPT, um, uh, um, [lips smack] launched, right? So that was the end of November for ChatGPT. A, a lot of enterprise budgets are set in October, November for the following year, and toward mid to the end of 2022, there were very deep fears about an impending recession that
- 7:27
didn't end up happening. But those fears basically made it so that most enterprises either froze or shrunk their IT budgets for 2023, okay? So what that meant is, were it not for particular tailwinds called ChatGPT, we probably wouldn't have seen a lot of new technology being developed in the IT departments at these large enterprises in 2023,
- 7:49
right? So this is sort of a bittersweet, uh, mix. I guess I just talked about a lot of this. Um, this was sort of a bittersweet mix in the sense that it did set us up to put the Eye of Sauron on a particular, specific, small pet project where, uh, a small amount of budget could be applied.
- 8:08
And that small pet project, uh, came from our friends at OpenAI launching ChatGPT right before the holiday breaks. And I'm convinced just from talking to some C-suite folks across the enterprises here in the US, that basically what happened is now CEOs and CFOs and s- you know, less technical on the computer science sense of the word, were
- 8:26
able to interact swiftly and easily with a single UI on the internet and get wowed by AI, right? So I was also wowed by AI. In fact, we were hosting, um, Arthur was hosting with our, um, uh, Series A leaders, Index Ventures, an event at NeurIPS, the major machine learning conference, the night that ChatGPT came out, and
- 8:44
it basically took over a, a bunch of nerds in a room being like, "Wow, this is very, very impressive." And that happened to everybody else, right? Flip back to November 30th, you know, you can make Eminem rap like Taylor Swift.
- 8:53
Oh, hey, isn't this funny? Oh, hey, uh, my mom did this sort of, like, joke about, you know, I wanna have poetry, but in the, uh, uh, uh, written in the, you know, the words of a rapper or whatever.
- 9:02
This started to happen over and over again, and what that meant is that that discretionary budget, uh, which exists, uh, was unlocked specifically for now the CEO's pet projects, which were called GenAI.
- 9:15
So 2023, uh, we still had austerity forced on us because of those frozen budgets,
- 9:21
um, or even reduced budgets. But the thing that was happening here is that now the only money going around that could be allocated was going to specifically GenAI, and so everybody focused on this.
- 9:31
A, it's a cool technology. Everybody focused on this, and the science projects started to sort of float around within the enterprise.
- 9:39
2024, we started to see GenAI-based applications going into production, right? Chat applications are the obvious one, internal chat applications, internal hiring tools, things like that.
- 9:49
And that's because basically the only budget going into new projects in 2023 was going to GenAI. And now as things go into production, primarily internally in 2024, uh, we have the folks who tend to dress in business suits, uh, asking questions around ROI, governance, risk, compliance, brand optics, and that kind of thing.
- 10:06
So now we're starting to get a little bit closer to people outside of the machine learning, the data science, the computer science world, the CIO's office caring about evaluation, right?
- 10:14
If I need to quantita- quanti- uh, if I need to have a quantitative, uh, estimate of risk, then I need to do evaluation.
- 10:22
2025, we've seen s- you know, scale-ups, right? Look, look at the revenue numbers for any frontier model provider. Look at the revenue numbers for a lot of us in this room.
- 10:30
Just everything is really going up right now, and that's because of usage, which is great.
- 10:34
That also means, uh, that's a function of the C-suite basically becoming comfortable basically talking about and putting large real budget into AI, right? So 2023's IT budgets, 2024's IT budgets, uh, set for the following year, those weren't frozen, right?
- 10:47
Those are earmarked specifically for AI applications and things along those lines. So we had science projects in 2023 go into production in 2024, and they're now shipping and scaling, uh, in 2025.
- 11:00
And also, like, frankly, the mo- the technology's g- just gotten really amazing, right? Like, all of us in this room are, are technical, but even I'm just amazed every time a new model is dropped.
- 11:09
The community has also really gotten behind this. You know, open source has gotten behind this. Venture capital, big tech is writing huge checks into, uh, into frontier model providers, and so on.
- 11:17
So everything is sort of coming together in 2025. And also, remember that third vertex. We have machine learning systems now moving toward autonomy. Okay, so 2025, we all hear it.
- 11:29
It's the year of the agent. I'm, you know, no longer is a question mark needed here. It's clearly the, the year of the agent.
- 11:35
Now, quick 30-second definition of an agent, uh, as defined, you know, in the late '50s onward. Um, agents need to perceive the environment. They need to learn. They need to abstract and generalize.
- 11:46
And unlike traditional machine learning, they're going to reason and act, right? We have reasoning models out there. We have systems that are acting in virtual environments or cyber physical environments.
- 11:56
And what that means is you have a lot of complexity introduced into the system, and you have a lot of risk introduced into the system. And that's great for those of us in, uh, email.
- 12:08
So at the end of the day, the thing that really matters, like I mentioned, is connecting when you're selling any product, not just our products in this room, when you're selling any product into an enterprise or an S-SMB, is being able to attach, uh, your product into some sort of downstream business KPI, risk mitigation, revenue gains, uh,
- 12:25
you know, losing less money, whatever. So now evaluations, you need to be able to do this because you're quantifying things, and they're finally a first-class discussion point, which is fantastic.
- 12:36
So we have the CEO, like I mentioned, November thirtieth, twenty, uh, twenty-two and onward, now at least knows what the tech is. And I'm not saying they know, you know, uh, what attention means, uh, right?
- 12:47
But I am saying that they know some of the capabilities around these generative models. They know some of the capabilities around agentic and multi-agent systems, uh, and they're comfortable, um, you know, talking to experts about it, uh, allocating budget for it, uh, and, uh, talking to their board of directors and shareholders about it as well.
- 13:03
Uh, we have the CFO who, because the CEO cares about this stuff, uh, also obviously needs to care about it, but she's gonna care about the impact to the bottom line, right?
- 13:11
That's what a CFO does. They're doing allocation, they're doing budget planning, uh, and they're going to need to basically write some numbers into an Excel spreadsheet, and those numbers have to come from, in part, quantitative evaluation.
- 13:23
Uh, CISOs, uh, now see this as, you know, a huge security risk and opportunity. And for those who haven't sold into enterprises before, CISOs are typically willing to write checks, uh, smaller checks, especially for startups, uh, more quickly and with less overhead than, like, a CIO would.
- 13:38
CIOs tend to have a bigger org, for one, and also tend to have a lot more process. CISOs tend to be a little bit more scrappy and willing to try out tools, and so on.
- 13:45
And so this actually happened before the agentic, uh, revolution. This actually happened when GenAI started coming up, where CISOs were like, "Hey, hallucination detection, prompt injection," things like that.
- 13:54
That's firmly in the security space, which is why you've seen a lot of guardrail products, including Arthur and including the ones that come from our competitors, going into the CISO's office and, and basically being able to sign a lot of deals.
- 14:06
Uh, the CIO, corporate CIO, has been on board the entire time, uh, and they're just trying to keep the ship sailing, so that's great. They're still on board. Uh, they wanna keep their job.
- 14:14
And the CTOs now, uh, they always want standard, right? They need to make these decisions based on numbers. Those numbers are coming in part from, like, OTEL standards, standards like that.
- 14:25
And that's great, right? So I've listed a lot of the C-suite here. I haven't talked about chief strategy officers or otherwise, but, like, the CEO, the CFO, the CTO, the CIO, the CI- CISO, they control a lot of budget, and now they are all willing to talk about, and they're all aligned about basically the need to understand
- 14:41
evaluation from AI. Great. So, uh, quick remind me, you know, you should hold me truthful here. All the evaluation companies, observability companies, monitoring companies, security companies, whatever you wanna call them, have shifted into agentic and multi-agent systems monitoring, right?
- 14:58
The point around you should monitor the whole system, you shouldn't just monitor the one model that is being used by one particular agent, that's well taken and well understood in industry and in government, and I think that's, that's great to have that at that top-line discussion point.
- 15:11
But, you know, keep me honest here. Uh, there was an article that came out in mid-April, uh, in The Information, uh, showing some leaked revenue numbers for a variety of startups in the evaluation space, Weights & Biases, Galileo, Braintrust, and so on, but they were lagged by about six months or eight months.
- 15:26
Um, and just from talking to friends in the space, uh, those numbers are no longer representative of what folks in this area are making. And so let's see what The Information leaks in early twenty twenty-six about this, and maybe we'll see something like revenue no longer lags at AI evaluation startups because this is the year for AI
- 15:43
evaluation. Great. So I'll leave it with some time for questions. Um, I did mention Mozilla is not firmly in the evaluation space. We do have a very nice open source, not monetized at all, uh, what we're calling a Lite LLM for multi-agent systems.
- 15:56
So if you're playing around with different multi-agent system frameworks, check out Any Agent. We implement a lot of them for you under our unified interface. So for people in this room, that might be a fun project to play around with.
- 16:04
So thank you. I'll, uh, have three, three minutes for questions.
- 16:13
I'll go up to... Sir.
- 16:16
Okay. Thanks. Uh, thanks. Great presentation. I just have a question really about the enterprise value about the-- Most of the evaluations in GenAI require domain expertise. So for example, if you're building a multi-agent system to do financial investment analysis, to do something called a discounted cash flow spreadsheet, is the agent doing it correctly or not?
- 16:38
Uh, I'm just trying to understand is, how is that problem getting solved? Because most of them are, you know, coming from an ML background where it was structured data, but this is a lot of unstructured data, and you have to measure the quality.
- 16:48
Like, is it in acting like a human, right?
- 16:51
Yeah. Yeah. It's, it's great. Actually, I have a paper in Nature Machine Intelligence talking about some of the problems that can come around when you do persona-based agents where I say, "Act like a farmer in Ohio in your mid-forties," and so on.
- 17:00
There's value in that, but you can't do it perfectly. And my gut reaction to this is, um, there was a leaked spreadsheet from Mercuri, which is a company that, uh, can hire in experts, showing, you know, fifty dollars an hour, a hundred dollars an hour, two hundred dollars an hour for experts to be hired by, for example,
- 17:14
Google, or for example, Meta, or for example, large banks to do kind of what you're saying, which is you're gonna have an expert sitting alongside the multi-agent system, uh, basically sitting next to, you know, the intern who's gonna come in and take your job in, like, a year or two, or change your job.
- 17:28
Maybe take is not the right word. But they're basically doing that expensive human validation in lockstep with the multi-agent system, which, you know, if you're gonna be doing discounted cash flow analysis, right, the kind of thing where, A, you can either make or lose a lot of money, and B, lose your job if you get it wrong,
- 17:41
uh, it's worth spending that large amount of money doing the human validation. It's a question for everyone, though, is like, what does that look like in five years once that data is incorporated into the systems themselves?
- 17:50
And that can be a mode as well, right? When you talk to anyone in the eval space, like, it's the dataset creation and the environment creation that matters more than anything, which is the point you're getting at as well.
- 17:58
So if I, if I spend a bunch of money to have a very good competitive, like, uh, DCF, uh, environment or whatever, that can help me versus my competitors.
- 18:07
So there is, like, um, there is CapEx going into that. Uh, yeah.
- 18:13
Thanks for the presentation. Um, do you have a rough, uh, maybe timeline on when you think, um, evals will primarily be driven by, uh, maybe GenAI or even LLMs?
- 18:23
Yeah, the LLM-as-a-judge paradigm is... And I think this was talked about in, in the previous talk as well. Uh, we see it getting used in practice because there are issues with it, right?
- 18:31
We have a paper in ICLR, uh, from last month talking about some of the biases that LLMs as judges have versus humans in things like conciseness or helpfulness and some of those anthropic words.
- 18:40
But the long and short of it is that, like, it, it solves the dataset creation problem in some sense in that, like, you can ask-- You, you give a persona to an LLM, and it, it, it, it is like a, a poor man's version of, like, a human doing the judging.
- 18:52
And so we see a lot of people using that as a crutch right now. But you need to, toward the, toward the last question that was asked, you do need to make sure you're validating this and, and making sure that you're not going off in some weird biased direction.
- 19:05
That's time? Oh. Uh, happy to chat offline. [upbeat music]