AI Engineer World's Fair 2024
E-Values: Evaluating the Values of AI
Read the talk
E-Values: Evaluating AI Against the Goals and Values It Serves
As AI moves from isolated predictions to autonomous workflows, evaluation must connect model behavior to user outcomes, business requirements, and the values those systems express.
From a talk by Sheila Gulati and Nischal Nadhamuni
Before you start: Basic familiarity with machine learning evaluation and large language models is helpful; no knowledge of Klarity is required.
Can we measure performance against our goals?
Can we evaluate an AI system’s performance in relation to the goals we set for it? As more work becomes automated and agents take on more responsibility, that question becomes a prerequisite for understanding what we have built. A system can produce impressive outputs without doing what its users actually need. Evaluation has to connect observed behavior to intended outcomes.
Sheila Gulati and Nischal Nadhamuni approach this problem as an investor and an operator: Gulati founded Tola Capital; Nadhamuni is Klarity’s co-founder and CTO. Gulati opens by noting Klarity’s newly announced $70 million Series B. The company’s ambition is to extend the coordination that software brought to internal business systems into external relationships. Customers and partners still negotiate individual documents, leaving important commitments scattered across one-off agreements. Klarity seeks to make those documents part of an ongoing, automated relationship—a foundation for what Gulati calls exponential organizations.
Gulati’s earlier work at Microsoft included leading database and developer-platform businesses, co-leading enterprise strategy, and advocating for Azure inside a company organized around Windows. That experience shaped Tola’s original thesis: a new generation of applications would be built on the cloud. AI applications now present another platform transition, but with a harder evaluation problem.
Cloud infrastructure favored scale. Data centers required capital that existing search, office-software, or retail businesses could supply. Buyers could compare speeds and feeds, performance, functionality, and price. AI retains many of those constraints—training and inference costs, scarce chips, and specialized engineering talent—but adds a more diverse field of models. Open-source development, academic research, and partnerships with large technology companies create contenders whose usefulness cannot be reduced to infrastructure specifications.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
From a correct prediction to a useful agent
AI reflects the understanding, intentions, and preferences of the people who build and use it. That becomes more consequential when agents represent individuals. Gulati traces a progression from models consuming the good and bad of internet data, through emergent capabilities and curated datasets, toward an anticipated era of self-learning and self-sufficient systems. Her urgency follows from that trajectory: evaluation needs to improve before full automation makes its deficiencies harder to contain.
For narrow AI, the evaluation question resembles a hammer meeting a nail: did it perform the specified task? An image classifier has a relatively bounded target. Broad AI introduces several dimensions that must be considered together.
| Dimension | Evaluation question |
|---|---|
| Capability and intelligence | Can it perform the required reasoning? |
| Domain understanding | Does it understand this field? |
| Values alignment | Whose values does it express? |
| Safety | What does avoiding harm mean here? |
| Context and end-user awareness | Does it understand whom it serves? |
| Modality | Does it work across text, images, video, and their combinations? |
The values question includes individual versus societal preferences and present versus future expectations. Safety also depends on context: people can disagree about what constitutes harm. End-user awareness cannot remain an afterthought when the purpose of the system is to serve that user.
These dimensions enlarge the evaluation problem beyond what simple task scores capture. Gulati frames the next part of the discussion around weaknesses in existing tools; Klarity’s operating experience then shows how those weaknesses appear in a product.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A leaderboard is not the user’s workflow
A fixed benchmark creates an optimization target. If the task is to solve for X, a model developer can concentrate on producing the right answer to X without establishing that the model understands the broader problem. Gulati uses an AP exam as the example: passing may indicate knowledge of the subject, or preparation for the kinds of questions that have appeared on that exam. The leaderboard alone does not distinguish them.
One response is dynamic evaluation data. Gulati describes Microsoft Research work that generates unpublished synthetic images and moves objects between them. Changing the scene changes the correct answer, allowing tests of spatial reasoning, visual prompting, and object recognition without repeatedly exposing the same test images. The description closely matches Microsoft’s A Dynamic Benchmark for Image Understanding project. The mechanism is to regenerate the task instances so memorizing a fixed set of answers becomes less useful.
Even without memorization, benchmark success may fail to transfer. A model can perform well on MMLU and still answer a particular business question incorrectly. Gulati points to FinanceBench, crediting Patronus AI and Stanford, as an example of financial questions exposing weaknesses that general benchmarks miss. The practical response is to put the user’s task at the center of evaluation and redesign UX and feedback systems to capture what that user needs. Evaluating the experience is difficult, but omitting it leaves the product’s purpose unmeasured.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
What a score cannot tell us
Evaluation is a proxy for the task we care about. It seeks evidence about correctness; it does not provide a complete account of what a neural network has learned or how it reaches an answer. Interpretability addresses a different part of that gap by examining internal mechanisms. For Transformer-based LLMs, Gulati describes tracing information through the network to understand how pieces contribute to an answer. She anticipates more of this work alongside reinforcement learning in GPT models.
Technical capability is only part of what a model inherits from its creators. It also reflects values, whether those choices are explicit or not. Personalization makes the issue concrete: a model of Sheila could keep reinforcing Sheila’s beliefs, drawing her further into her existing worldview. The analogy is the news echo chamber, where personalization can intensify polarization. An agent that represents someone must therefore be evaluated for more than how closely it mirrors that person.
Several choices remain open:
- Appease or challenge: Should the system reinforce the user’s view, question it, or let the user choose?
- Present or aspirational values: Should it reflect society as it is or the society people hope to create?
- User or model-maker responsibility: Should individuals select their own values, and which obligations remain with the model’s creators?
Gulati presents these as unresolved design questions. They broaden the definition of success before the discussion turns to a concrete enterprise application.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The document is only the beginning of the task
Klarity automates cognitive, non-repetitive back-office work. These workflows resist a simple fixed sequence of instructions because a person must interpret a document before deciding what to do. Nadhamuni describes roughly eight years of work on this problem, predominantly involving PDFs in enterprise operations.
The tasks include revenue recognition, matching invoices to purchase orders, and processing tax withholdings in different languages. A finance or accounting team must extract meaning from unstructured documents while following a tightly regulated process. Matching requires aligning information across two documents; tax processing may require understanding both unfamiliar layouts and another language.
Document quality adds another layer of difficulty. Rotated pages, poor scans, graphs, images, and tables can all disrupt processing. Simply handing such a document to ChatGPT does not reliably complete the workflow. Nadhamuni reports that Klarity invested tens of millions of dollars in a stack for handling these documents. The accompanying slide shows the settled checklist of document-processing caveats, including searchable text, PDFs, tables, diagrams, and OCR.
Klarity’s customers at the time were primarily B2B SaaS companies, many near the conference venue, with finance and accounting teams as the main users. The company was beginning to expand beyond that base. This customer concentration gave evaluation a specific target: the documents, rules, and outcomes those teams actually handled.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Why familiar ML evaluation stopped being enough
Nadhamuni dates Klarity’s founding to 2016 and describes four pivots between 2016 and 2020, followed by a final move toward finance and accounting that found product-market fit. Rebuilding around generative AI over the preceding couple of years accelerated the business; he reports more than $90 million raised.
The team had spent more than six years working with traditional machine learning. Supervised learning supplied familiar evaluation tools: train/test splits, F1, ROC curves, and shared tasks such as sequence labeling and SQuAD. Initially, it was not obvious why generative AI should require a fundamentally different approach.
Nadhamuni reports that, after replatforming, Klarity had processed more than half a million customer documents, with more than 15 production LLM use cases and more than 10 LLMs under the hood. A single use case could involve several models working together.
The first complication was nondeterministic performance: repeatedly uploading the same cognitively demanding PDF could produce different responses. The next was the expansion of the product itself. Classification, regression, and random forests usually fit established interaction patterns. Generative models made entirely different experiences possible, and those experiences required different definitions of quality.
- Workflow documentation: Klarity Architect takes a recording of someone doing their job and creates a business requirements document. Nadhamuni describes a typical output as a ten-page Word document with a flowchart and images, comparable in form to a consulting deliverable.
- Natural-language analytics: A user can ask how a contract population has evolved over time and request a stacked bar chart. The output combines interpretation of the question, analysis, and presentation.
- Document extraction: A system locates consequential information in tables or legal language buried inside a document. Here, extraction accuracy remains central, but the information can take very different forms.
Each experience creates useful capabilities for customers while making evaluation less reducible to a single predictive score.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
When evaluation takes longer than the feature
Faster implementation changed the economics of evaluation. Klarity’s earlier deep-learning workflow required dedicated annotation teams and GPU training. Nadhamuni contrasts that cycle with generative feature development:
| Development period | Reported feature timing | Evaluation consequence |
|---|---|---|
| Earlier DNN workflow | Five to six months | Two to three weeks for evals was acceptable |
| Generative AI workflow | Days | One to two weeks for evals can become the bottleneck |
Nadhamuni reports that Klarity launched Document Chat within 12 hours of ChatGPT’s launch. Evaluation that once occupied a small fraction of development could now dominate it.
Model selection added uncertainty. Nadhamuni says small MMLU differences could matter for Klarity’s demanding cognitive tasks, but the ranking was not dependable: models advertised as better on MMLU sometimes performed worse on internal benchmarks. General benchmark scores could help identify candidates, yet they could not settle which model worked best for the customer workflow.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Test the experience before building the full evaluation
Klarity’s response starts with permission to be imperfect during feature discovery. Nadhamuni describes the company’s evaluation practice as still developing, rather than a solved framework. High-quality evals are expensive, so requiring a complete suite at the start of every experiment can waste effort on experiences users will reject.
A new generative feature is almost its own product-market-fit problem. Document chat, natural-language analytics, and automatic video watching introduce unfamiliar ways of working. A recommendation engine, by comparison, often starts with greater confidence that the interaction pattern is useful. Klarity therefore front-loads user testing and delays substantial evaluation investment until the experience has shown promise. Features that fail at the UX stage can stop there. That sequencing does not remove the requirement to build evals before entering production and scaling.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Work backward from the outcome
Once a feature is headed for production, the direction of reasoning changes. Start with the experience and work backward toward measurements. Choosing F1 first and hoping it correlates with user value reverses that relationship. An individual eval does not have to represent the whole experience, but the evaluation suite should reflect it in aggregate.
Klarity walks down three levels:
- Business outcome: Identify the end-user result and the business value the feature should create.
- Adoption and utilization: Decide which observable behavior indicates that people use the feature.
- Feature health: Measure lower-level properties such as JSON adherence and variability in generated outputs.
For an extraction feature, a small health check could examine whether repeated responses contain a required string field and whether the extracted values agree:
python
import json
def extraction_health(responses: list[str], field: str) -> dict:
values = []
for response in responses:
try:
record = json.loads(response)
except json.JSONDecodeError:
continue
if isinstance(record, dict) and isinstance(record.get(field), str):
values.append(record[field])
return {
"total_responses": len(responses),
"valid_records": len(values),
"distinct_values": len(set(values)),
}
responses = [
'{"invoice_number": "INV-104"}',
'{"invoice_number": "INV-104"}',
'{"invoice_number": "INV-140"}',
]
print(extraction_health(responses, "invoice_number"))
This illustrative check separates structural validity from output variation. It cannot determine which invoice number is correct: a stable, well-formed answer can still be wrong. That is why health measurements sit below the user outcome in the evaluation stack.
In production, customers supply their own labels and requirements. Klarity annotates customer data during user acceptance testing (UAT) and builds accuracy measures for the particular use case. Matching documents and extracting fields need different definitions of correctness. User feedback then helps close the loop, but silence is not positive feedback. The provider still has to measure rigorously and monitor data drift—for example, whether incoming documents differ substantially from the population already evaluated.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Share evaluation data and constrain the search space
Customer-specific tests are a starting point, but a sufficiently common use case can justify shared evaluation infrastructure. Nadhamuni reports surveying more than six synthetic-data providers and trying several before Klarity built its own generation stack. The investment begins when a use case is large enough and enough customers share it to make customer-agnostic evals worthwhile.
The difficult part is keeping synthetic data representative. Klarity assigns people to monitor whether generated documents resemble customer documents distributionally. Nadhamuni says the company had not made that validation fully automated or quantitative. A second check compares model accuracy: unusually strong or weak performance on synthetic data can indicate that the generated population is not a useful stand-in for real work.
Prompt engineering creates another scaling problem. Klarity has features for free text, tabular extraction, matching, and table composition, each requiring customer-specific prompts. Across a large customer base, Nadhamuni describes the resulting inventory as thousands or tens of thousands of prompts. Automated prompt engineers (APEs) replace manual prompt writing, but introduce their own model-selection problem. Different LLMs can perform differently across APE tasks and customers, multiplying the combinations to evaluate.
Klarity reduces that search space by using one LLM for APE tasks, based on observed commonality in which models perform well. The choice is revisited as new models appear, but it remains fixed while the team iterates elsewhere. Nadhamuni acknowledges that a broader search might recover another percentage point of accuracy; he presents that as a possible gain, not a measured improvement, and judges the grid-search infrastructure too expensive to justify. Fixing one dimension makes the overall project easier to manage.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Evaluate what might become possible next
Production evals measure what the company does now. A separate, lighter set can explore workflows it hopes to support several years ahead. Model capabilities can advance faster than a product roadmap anticipates, so a wish list of future tasks helps a team recognize new opportunities. These exploratory tests do not need MMLU- or BBH-level rigor for workflows no customer has requested yet.
Nadhamuni reports that Klarity formed an initial assessment of GPTV, the vision-capable GPT model discussed in the talk, in about an hour: it seemed good at check marks and perhaps less good at pie charts. The result was a quick organizational mental model of what to investigate next, rather than a production acceptance decision. Lightweight evals serve that discovery purpose alongside the more rigorous tests needed for deployed workflows.
The case study closes with a historical hiring invitation following the Series B: Klarity was recruiting across AI, back-end, front-end, and go-to-market roles.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The responsibility that comes with defining success
Returning to the broader question, Gulati emphasizes the disproportionate influence of people building early in a platform transition. Their choices help establish what the industry will measure and reward. Reinventing misleading benchmarks, adding depth to evaluation, and treating capability, context, and values together are therefore practical responsibilities for the engineering community.
That work benefits from sharing unsuccessful approaches as well as successful ones. Klarity’s account is useful because it exposes decisions, constraints, and unfinished problems. Gulati couples this openness with curiosity, empathy, mutual assistance, and examination of the builders’ own values. Evaluation standards emerge from people making choices together; those people need to inspect their assumptions as well as their models.
Gulati invokes Brian Christian’s The Alignment Problem and paraphrases the possibility of AGI as a mirror that reveals which values are distinctively human. Those values also change over time. The question is therefore not only whether AI reflects us accurately, but what kind of society its builders want to help create. Her closing appeal—to remain curious, empathetic, kind, fair, and intelligent—ties the opportunity to create new intelligence to accountability for what that intelligence serves.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
From the talk
The original financial question-answering study explains its dataset, retrieval setups, human evaluation and configuration-dependent results.
Klarity's June 2024 announcement describes its funding round, document-automation focus and hiring intentions.
Klarity explains how Architect turns recorded workflow walkthroughs into process documentation, diagrams and improvement suggestions.
Brian Christian's book explores machine learning, human values and the challenges of aligning automated systems with human intentions.
Further reading
- A Dynamic Benchmark for Image UnderstandingDocumentation
Microsoft Research's synthetic visual benchmark tests spatial understanding, visual prompting and object recognition with procedurally generated images.
Read the complete timestamped transcript
- 0:00
[upbeat music] Well, thank you all for being here.
- 0:15
It is a small but excited group of people around evals. And, you know, evals may not be the most sexy talk at this conference, but it might be one of the most important.
- 0:27
And Nischal and I will spend time today speaking through why right now we think it's a seminal moment for evals and the import sort of the changes that we're seeing in the market as we move into these agentic frameworks, right?
- 0:41
So as things are more automated, as things happen in a more agentic way, we need to make sure that we understand what's happening with evals. We're all here obviously because we're very long AI, we care a lot about AI, but we have to really understand what's going on with it.
- 0:57
And really, if you think about the crux of it, can we evaluate the performance of this, our systems in relationship to our goals for those systems? And that's what we'll be walking through today.
- 1:07
And so about us, um, my name is Sheila, and I'm joined by Nischal. It's great. We're, we're an investor, investee pair, right? A VC and portfolio company. And I think there should be more talks like this because some of the most fabulous, uh, portfolio companies are not great at giving shout-outs to themselves.
- 1:25
So I get to give the shout-out to Klarity and to Nischal. He's the co-founder and CTO of Klarity. He's sitting up here, but he'll be speaking very soon as well.
- 1:35
Um, and Klarity just announced their massive, uh, seventy million dollar Series B financing on Monday. So we're gonna hear about that journey. It hasn't been a short journey, so I think that can be quite inspirational for a lot of you as you think about what you're building and the size and the scope of what you're building.
- 1:50
And so that company builds what we call exponential organizations. Well, what does that mean, right? Software transformed how we all work, right? This internal systems is always synced, always on nature of software systems for internal work changed the nature of the efficiency and effectiveness of buildings that, of, of c- businesses that we build.
- 2:11
Klarity is doing that for your external world, right? Most of your relationships with your customers, your partners are, are dealt with through documents. Those documents are one-off negotiated. They're one-off pieces of paper, right?
- 2:25
Klarity is automating all of that and then allowing you to build exponential organizations through that type of real-time relationship with those documents. So you'll hear more about that and how Klarity has implemented a number of eval systems later in this presentation.
- 2:42
Um, myself, I founded a venture firm called Tola Capital, uh, well over a decade ago now. And the reason I founded the firm was I was working at Microsoft.
- 2:51
I was running the database and developer platforms businesses at Microsoft. I co-led the company's enterprise strategy and was fighting the good fight against Team Windows to launch Azure. So being Team Cloud at a company that was based off of Windows was a really difficult thing, but it was fabulous as well.
- 3:10
And obviously, Azure's gone on to do pretty okay for itself. So the, the genesis of the firm Tola Capital really was, "Hey, how do we think about this next generation of applications that would be cloud-based?"
- 3:21
And now we're even more excited by the opportunity to bring this next generation of AI-enabled applications to the fore. So I'm gonna walk down memory lane for one quick moment and say, you know, what we saw in the advent of the cloud was very clear.
- 3:36
It would favor scale. The CapEx requirements, the physicality of building out those data centers was so expensive that you had to have a search business or an office business or a retail business to go fund that development, and then we would build on top of that, right?
- 3:54
The evaluation of those platforms was more straightforward. Am I offering you speeds and feeds? Am I offering you performance? And then, of course, you're layering on what functionality at what price I'm offering.
- 4:05
But the evaluation of the physicality of that world was more straightforward. Now, as we enter this AI world, we're saying, "Wow, there's," you know, some of the similar characteristics of large CapEx and large mega cap participation, right?
- 4:20
Where we have obviously these systems running on top and models running on top of clouds and being trained by the clouds and that, that compute cost, that inference, the ever so, uh, difficult to get your hands on chips.
- 4:33
The talent of all of you in the room, right? The AI engineers that bring this to be. But in addition to that, you have the proliferation of open source and open-source models, and just the ability of those models to do an incredible job at delivering incredibly complicated scenarios that are, you know, really, really, really strong contenders.
- 4:55
And so you have a proliferation of players, a proliferation of models. You have a deep academic heritage of those open source, um, and, and AI development, and you have great mega cap partnerships for those models as well.
- 5:08
So what, where are we going, right? We're-- AI is just-- it's, it's so important to think about AI as more than just another tool, right? It is a reflection of us.
- 5:19
It is a reflection of our understanding of the world, our intentions, our preferences, and at the end of the day, our society. And we'll talk a little bit about what that means individually and collectively, especially in this world where agents will represent more of us as individuals.
- 5:36
So let's talk about some of the shifts here. Um, you know, we, we started with AI eating everything on the internet. I like to call this all the garbage and all the gold on the internet was consumed by AI.
- 5:47
Then we moved into, you know, doing that gave us emergent behaviors. That's, of course, a great fancy way of saying we're not exactly sure how it knows what it knows, but we know it knows it, right?
- 5:58
This era then, then we have trained datasets, curated datasets, but we're moving into the era of self-taught, self-learning, and self-sufficient models. And so if you pause and think about that for a second, before we get into a truly self-sufficient era, we really need to fix evals, right?
- 6:15
Because it kinda gets to be too late in that era. And so that transition to full automation is happening faster and more aggressively than any of us thought. So what does this mean, right?
- 6:28
Narrow AI evaluation was, "I'm a hammer, you're a nail. I'm hitting you. Am I doing that right? You're an image. Am I classifying that image?" And I could, I could tell whether I was performing on that task in the right manner pretty s- in a pretty straightforward way.
- 6:45
Then you get into broad AI, right? And this radar chart speaks a little bit to how this evaluation becomes more multifaceted. What's my capability and intelligence? Am I serving a domain?
- 6:57
Do I understand that domain? What are my values, right? What are the... And whose values do I care about? Mine, yours, the user, society's, today, tomorrow? All of these questions come together.
- 7:07
Safety. Okay, do no harm. What does that look like? Pe- that could be different for people. And of course, the context and end-user awareness, which is often not discussed in an eval world, right?
- 7:19
Your end user is who you're delivering and developing these solutions for, but they're often an afterthought in terms of evals and evaluation. And then we have to do that across all of the different modalities of image and text and video, and kind of all of these together is creating a much larger problem around evals.
- 7:39
So today, the tools are still simplistic. We're gonna dive into kind of each of these areas of simplisticness and talk about some of the innovation happening to deliver this forward.
- 7:49
And then what Nischal will do is show us how Klarity has dealt with each of these, um, issues related to evals as well.
- 7:57
So benchmark hacking. I like to call this, you know, what you, what am I solving for? Solve for X is how a lot of these benchmarks and leaderboards work today from an eval perspective.
- 8:08
And it's interesting because scoring high on the benchmarks is pretty easy to do if you know what the, what we're solving for. If you understand X, you can do everything to solve for X, and you can look as intelligent as you want at solving for X.
- 8:21
But the real, the real reason is you may not understand anything about what's happening. You may be just solving for X. And that, that's a pretty scary reality on some of these things.
- 8:30
And so we say, "Well, these ALMs, LLMs passed an AP exam on a particular subject." Meh. Does that mean it was trained well on the questions that have heretofore come on that AP exam, or does that mean it actually understands the subject?
- 8:43
It's a really, really interesting question, but what we're seeing on a lot of the AI leaderboards is the best solvers for X are at the top of those leaderboards, and that's a problem.
- 8:53
So one area where we're seeing research happen, um, this is a Microsoft Research paper around dynamic benchmarks. Rather than saying, "Solve for X, can you identify this image?" They're creating dynamic datasets.
- 9:05
Basically with synthetic data where you can say, "Hey, we're moving objects around." These are not published. They are not, uh, public, right? So you don't know what the answer is coming into it.
- 9:17
They're first... Then, but then I can test you on spatial reasoning, visual prompting, object recognition. The images are changing, so the models have no ability to memorize those benchmarks.
- 9:29
This is an area, dynamic benchmarks in general, where I think we'll see a lot more work.
- 9:34
Benchmarks versus real-world scenarios is, is, is element two on this. You know, it's interesting, there's a lot of model creators that claim that their models perform very well and are very generalizable, and then you actually go and ask them a specific set of questions.
- 9:49
Even if I've done super well on MMLU and these things, and you say, "Okay, I'm gonna go test this logic and test this reasoning," and the answers are s- simply wrong, and they're wrong much and most of the time.
- 10:00
And, and this, this great example from Finance Bench, which was, um, Patronus AI's work and Stanford's work around saying, "Hey, you know, these basic financial questions were not answered when you actually benchmarked it on those real-world scenarios."
- 10:15
And the... It doesn't... The user is not at the center of our evaluation universe, right? We have to put the user at the center. We have to revamp the UX and feedback systems in order to really understand and capture more user needs.
- 10:29
Evaluation of UX is super difficult, right? And, and Nischal will talk a lot more about this in, in the case of Klarity, but that's why we're all actually here, to deliver that end-user value.
- 10:40
And so if we're not doing that evaluation, what are we doing?
- 10:44
Um, black-box models, right? So this is a problem that's, that's sort of hiding in plain sight, obviously. Evaluation is a proxy for the, for a task. Evaluation is seeking truth.
- 10:55
It is not truth. And so how do we really understand what we're doing in a world where we can't ma- mathematically represent what neural networks have learned and how they have learned that?
- 11:08
And so the area of AI interpretability is not new, but it's super important as we think about the rise of these complex AI systems. And so researchers are trying to open up these black-box models, show us the how and the steps in this.
- 11:23
And, and the Transformer-based LLMs, we're really tracing information flowing through the network. And so this is a good example of, you know, sort of asking questions and seeing, "Hey, how am I answering it?
- 11:34
How are these pieces coming together?" We'll see a lot more of this interpretability work in the RL, um, the reinforcement learning work that's happening with the current GPT models.
- 11:43
It's super, super, super important that we get this right now. Now, values, right? When we create AI models, we do instill our own values in them, whether we want to or not, right?
- 11:56
And these evaluations have to understand the technical capabilities, but also the underlying values that we're putting into the models. And, and this is... There's a lot of questions on this, and I have way more questions than answers, as I think we all do.
- 12:10
But it's, you know, are we creating values that benefit humanity? As we go into this agentic world and you have your own model, right, does that just reinforce you?
- 12:22
Is that a good thing, right? So the, the, the model of Sheila is gonna believe more of what Sheila believes and get deeper and deeper into that Sheila-ness. Is that a good thing?
- 12:33
Right? We've seen the echo chamber of ourselves in news. We've seen the polarization that this has caused in our society. We should be asking questions about where we're going to get to as we enter this agentic world.
- 12:47
And this is kind of a, you know, it's a question, right? This isn't, you know, a simple two-by-two to ask the question, do we want to appease or challenge users?
- 12:56
Do we want you to choose whether you are appeased or challenged? Do we wanna ground AI in present or future societal values, aspirational values versus present values? Do we wanna encourage users to select their own values?
- 13:11
Should model makers be responsible for this? Lots of questions, less answers. Now we're gonna turn it into a real-world example, leveraging Klarity to talk about how the company is replicating human cognition.
- 13:25
Nischal? [audience applauding]
- 13:31
Uh, thanks, Sheila. Thank you for having me here. My name's Nischal. I'm co-founder and CTO of Klarity. Uh, as, as Sheila mentioned, what we do is we automate back-office workflows, things that traditionally required large teams of offshore, of offshore humans, and throughout human history have been impossible to automate because they're cognitive and they're non-repetitive.
- 13:50
You can't just write down a simple set of steps and automate these workflows, and that's what we've been working on for the last eight odd years. Um, predominantly these are document-oriented workflows, so it's some kind of PDF that a human being is ha- having to read as part of a company's back office.
- 14:06
Uh, one example of this is revenue recognition, typically part of an accounting team, um, matching invoices to purchase orders, so actually aligning two p- two documents and matching them to each other, um, processing tax withholdings that often come in many different languages.
- 14:20
But you probably get the sense. You get PDFs that are completely unstructured. Somebody has to go through them because it's a tightly regulated process, uh, part of a finance and accounting team, and today this is done completely manually.
- 14:33
Um, there was actually a really great keynote yesterday, uh, by, by a gentleman who pointed out that document processing tasks are generally super tough for LLMs for a variety of reasons.
- 14:42
Uh, if, if the page is like rotated or the scan quality is bad, if you have like graphs or images, if you have tables inside of these, um, it is very, very tough to get this to work, and it's not the kind of thing you can just give to ChatGPT and it happens out of the box.
- 14:57
This is basically exactly what we do, and we spend, you know, a long, long time and tens of millions of dollars building a stack that's able to deal with these kinds of documents.
- 15:06
Um, these are some of our customers, mostly B2B SaaS companies, um, mostly in a five-mile radius around where we are right now. Uh, and we predominantly serve their finance and accounting teams, although we're expanding quite a bit beyond that.
- 15:18
Our journey has been a little bit unorthodox. It's, it's kind of like that, um, you've probably seen the startup curve of like the trough of disillusionment and you get the TechCrunch article in the beginning, something like that.
- 15:27
Um, we, we founded the company in twenty sixteen. Um, we pivoted four times between twenty sixteen and twenty twenty, and when we were kind of at our wits' end about to give up, we did one final pivot, um, and focused on finance and accounting teams and, and found pretty strong product market fit there.
- 15:43
Uh, in the last couple of years we've completely re-platformed around generative AI, and that's really been a shot in the arm for the company, uh, and have gone on to raise, uh, ninety million-plus off of that.
- 15:54
But before generative AI, we were in traditional ML for more than six years. Um, we, we like to say that we started an AI company five years too late.
- 16:02
Um, and so a lot of the concepts around evaluations were, you know, they seemed pretty natural to us, and we didn't really see why generative AI had to be any different.
- 16:10
Uh, in supervised learning, which is a majority of what we did pre-gen AI, you have your train split, your test split, uh, the metrics are pretty well defined, F1, ROC curves.
- 16:20
Um, there are these like shared benchmark tasks like sequence labeling and SQuAD, and these were all very well understood, so it wasn't immediately apparent to us why generative AI has to change any of these things.
- 16:31
Um, and in, in the two years since then, we've, as I mentioned, re-platformed to generative AI, and that's gotten us quite a bit more scale. We've processed more than half a million documents for customers.
- 16:41
Uh, we have more than fifteen unique LLM use cases running in, in production today, uh, and more than ten LLMs under the hood. Oftentimes a single use case has multiple LLMs working together.
- 16:54
But all this kind of begs the question, why does eval for gener- generative AI have to be any different than traditional ML? And this is kind of the journey that we've been on in the last couple of years.
- 17:03
Um, the first, and a, a number of speakers have spoken about this, so I'm not gonna go into too much detail, is non-deterministic performance. Very challenging for us 'cause it's a very cognitively demanding task, so if you upload the same PDF multiple times, very likely you'll get completely different responses.
- 17:17
Um, new user experiences. I, I think fundamentally when you move from like discriminative models, classification, regression, random forest, to generative models, the types of experiences you can provide your users expand quite a bit, and these new experiences are just much harder to evaluate.
- 17:33
So a couple of examples from what we do, um, we have this tool called the Architect where users will record their business workflow, like just them doing their job, upload it to Klarity, and then we'll create a business requirements document out of that.
- 17:46
So it's typically like a ten-page Word document. It has a flowchart, images, very comprehensive, like what a McKinsey would build for you. It's not at all intuitive how to eval something like that.
- 17:56
Another example, part of our product is, um, you can do natural language analytics, so you don't need like a BI specialist. You could just ask it questions like, "Hey, how's my contract population evolved over time?
- 18:05
Give it to me in a stacked bar chart." It'll do that for you. Uh, cool feature, but like how do you eval something like this? And then a more traditional document extraction task.
- 18:14
You are trying to find certain parts of a document that are consequential in some way. They could be tabular, they could be legalese buried inside of a document, and you need to know what accuracy you're doing that with.
- 18:24
So this is why new experiences, while very rewarding to our customers, have been very challenging to us from an eval perspective. Uh, the second is the rate of feature development.
- 18:33
So in our previous deep neural net world, DNN world, a feature took, like, five to six months to build. We literally had teams that would annotate data, dedicated teams to annotate data, GPUs, most of which are gathering dust now to train.
- 18:45
And so end-to-end, it was, like, six months. So if it took, like, two to three weeks to build evals, thoughtful evals, that was completely acceptable. It was a pretty small fraction of the feature development time.
- 18:54
Now what we're seeing is we can get features out of the door in days. Um, within twelve hours of the ChatGPT launch, we launched Document Chat as a feature.
- 19:02
And so in, in that world, it's unacceptable that it takes a week, two weeks to build evals. It now becomes the bottleneck to feature development.
- 19:11
And the third is, uh, benchmarks diverging from performance. Sheila talked about this a little bit. Um, what we've seen at Klarity is even slight differences in MMLU actually make a very big difference to us because we're kind of at the frontiers of human cognition doing something that's very challenging for most human beings.
- 19:27
We've seen many cases where, not gonna name names, a model is supposed to be better in terms of MMLU, and then it's totally not on our internal benchmarks. There's a variety of factors that go into it, but it's made testing very chaotic and challenging for us.
- 19:41
And so the question then is what can be done? And I'm pretty sure most application developers are running into these problems or something similar. Uh, I'm not gonna say that we have the silver bullet, and I would say we are nascent in our eval journey.
- 19:51
But here's a couple of things that have worked for us. So the first is, like, really give yourself the gift of imper-im-imperfection. Don't put the threshold of high-quality evals at the beginning of the feature development life cycle.
- 20:04
Um, good evals are not trivially cheap to build today. They probably are not gonna be for the foreseeable future. And so we really try to front-load user testing. What we found with these generative AI features is we have to think of each feature almost as its own product market fit 'cause we're delivering experiences that people have never
- 20:20
had before, chatting with a document, natural language analytics, uh, watching videos automatically. And so there's a lot of user experience risk actually baked into each of these features compared to traditional machine learning, like a recommendation engine where you have fi-- you have fairly high conviction that the form factor is correct.
- 20:37
So what we try to do from a development perspective is front-load the UX risk, back-load building out evals. Uh, of course, I'm not advocating that you go into production and scale without building evals, but a lot of features die at the UX stage itself.
- 20:50
Let them die before you build out evals.
- 20:53
But once you decide to go into production, kind of our framework for this is you need to think backwards from the user experience, not forwards from what is easy to measure.
- 21:02
So let's not just say we wanna measure F1 and hope that user experience correlates to that. We wanna look at the end-user value, move backwards from there. Not every eval can or should reflect the entirety of the experience, but you want your evals in aggregate to be reflective of the user experience.
- 21:18
Um, a simple exercise that we do for this when we're building features is we kind of walk down the stack of what is the end-user outcome, the business value we're driving?
- 21:24
What is a good indicator of adoption/utilization? At the lowest level, how are we measuring health of this feature? It could be something as simple as, like, JSON adherence, variability in the output, and so on.
- 21:35
So in practice, this is what it looks like. Every customer is basically giving us their own set of labels. You could think of this as, like, bespoke enterprise AI.
- 21:42
And so we'll actually annotate data for that customer as part of our UAT process, build out use case-specific accuracy metrics for them. Are they trying to do matching? Are they trying to do extraction?
- 21:52
Um, and then there's still user feedback as, uh, uh, user feedback as kind of there to close the loop. But i- we think it's very dangerous to assume that the absence of user feedback is positive feedback.
- 22:03
So we try to put the majority of the onus on ourselves to be rigorous about metrics, um, and we have various tools to monitor data drift. Are we getting documents that are very different from the population we've seen so far?
- 22:14
Um, we've invested a lot in our synthetic data generation stack. We actually surveyed, like, six-plus providers in the market, tried a bunch of them out, um, were not too happy, and so ended up building our own synthetic data generation stack.
- 22:26
Everything that you see was synthetically, um, generated. The way that we think about this is once a use case becomes large enough within the company, we have enough customers doing it, we wanna, uh, we wanna invest in customer-agnostic evals.
- 22:38
That's when we go down the, the kind of path of synthetic data. And we have a team where part of their job is just monitoring that the synthetic data is distributionally similar to what we're seeing from customers.
- 22:50
So we haven't yet cracked the problem of doing this in an automated, fully quantitative way. The other kind of sanity check that we have is, of course, looking at accuracy scores and making sure that our models are not excessively or underperforming on synthetic data.
- 23:03
Um, the next, the next little trick that we use is kind of reducing the degrees of freedom. So this is kind of part of our architecture where we have numerous features, each of which require a custom prompt for each customer.
- 23:15
Um, so in this case, these are four features: free text, tabular extraction, matching, table composition. Each of these has a prompt for each customer, so you can imagine over a hundred customers, uh, you could end up with, like, literally thousands or tens of thousands of prompts.
- 23:27
So instead of manual prompt engineering, we have these APE things, automated prompt engineers. Now, the trouble is different LLMs perform
- 23:35
d-- uh, have different levels of performance on different APE tasks and on different customers. So you get, like, this exponential explosion in complexity, and what we did instead is we just said, "Well, we're seeing quite a bit of commonality in which LLMs do well on APE tasks, so let's just use one LLM for APE tasks."
- 23:52
Not permanently, and we'll continuously reevaluate as new LLMs come out, but if we can fix this dimension of freedom, it gets a lot easier to iterate. So could we eke out another percentage point of accuracy if we didn't do this?
- 24:03
Yes. But building that grid search infrastructure is just too expensive, and we don't think it's a good investment of time. So this is kind of another trick that we use to just...
- 24:11
almost like dimensionality reduction at a project, project management level.
- 24:16
And the last thing I'll mention is, like, identify future potential. I think people spend a lot of time, and rightly so, building evals for what their company does today.
- 24:23
But you should almost have a wish list of what are additional use cases that you want to grow into over time, three, four, five years down the line, because we are frequently surprised that technology is evolving faster than what we can see.
- 24:34
But it is too high of a bar to say we wanna have like MMLU or BBH-level metrics for future workflows that nobody has asked us for yet. And so what we do is we have these very scrappy, um, kind of future-facing evals.
- 24:47
Uh, for example, when GPTV came out, in about an hour we were able to say, "All right, this is our mental model of how it's gonna do. It's good at check marks.
- 24:53
It's maybe not so good at pie charts," et cetera, et cetera. And so that, I think, muscle of just building this organizational mental model very quickly in a scrappy way is also something that's been helpful to us.
- 25:05
Um, I wouldn't be a startup founder if I didn't end with a shameless plug. Uh, as Sheila said, we, we just raised a $70 million Series B. We are hiring across the board, AI, back-end, front-end, go-to-market roles.
- 25:15
So if any of this sounds interesting to you, I'd love to chat. Thank you so much, and I'll hand it back to Sheila. [audience applauding]
- 25:26
Thank you, Nischal. Uh, I have to say, working at Klarity is a dream, so, uh, anyone who's interested, please see, see either one of us afterwards. Um, this, so this is my call to action.
- 25:36
I've seen these paradigm shifts happen in the past, and I think one thing that's important is to remember the people that are early to this sort of AI revolution that's happening, these people matter disproportionately, and these people are all of you.
- 25:52
And so understanding and owning your power as you think about what happens with the next generation of evals, of integrating values into things, it's, it's actually, it's more than just sort of words on a slide, right?
- 26:06
There is a real empowerment and opportunity to spend time thinking about this to get this right for the ecosystem and for the industry. So what do we do? We innovate.
- 26:18
We reinvent benchmarks. We don't let the leaderboard benchmarks stick that are, that we know are just kinda BS, right? There's just, there's too much happening today that isn't truly understanding how these systems work.
- 26:31
We bring depth into the evaluations. We bring multifaceted nature of these evaluations together. We do so as a community, right? I think it's incredibly important that we bring curiosity, empathy, help to one another.
- 26:47
Like, you know, I, I love the fact that Klarity wanted to come and talk about their journey, what they did that was right, what they did that was wrong.
- 26:53
We share with one another on that. And I think that we introspect our own value systems, right? I think that today we are looking at, um, such a pace of innovation that we really need to think, what does it mean to drive this?
- 27:08
How are we the trailblazers for this happening? You know, I, I don't know if folks have read this book, The Alignment Problem. Brian Christianson wrote a great book, uh, I think it was last year, and it was one of my favorite reads of the year.
- 27:19
And he, he had this quote around sort of if we do see AGI, it will be an interesting mirror of society. It will tell us which values are uniquely human in nature.
- 27:29
Well, those change over time. How do we want to, how do we want to drive a society that has values that we can be proud of as these trailblazers?
- 27:39
So feel empowered, be curious, stay empathetic, kind, fair, intelligent, right? That's what we're doing. We're printing new intelligence for the world, and this is both your opportunity and your accountability.
- 27:50
And so thank you all for leading the future of AI. [upbeat music] [audience applauding]