AI Engineer World's Fair 2024
E-Values: Evaluating the Values of AI
About this talk
Tola Capital founder Sheila Gulati and Klarity co-founder and CTO Nischal Nadhamuni examine why increasingly agentic AI systems require evaluations grounded in operational objectives, human values, and real-world performance. Gulati contrasts narrow task testing with more adaptive benchmarks, discusses black-box models and Transformer-based LLMs, and situates evaluation within the changing AI investment landscape. Nadhamuni explains Klarity's automation of document-heavy finance and accounting workflows and the need for customer-specific labels in bespoke enterprise AI systems.
Chapters
- 0:00Why agentic AI needs better evaluations
- 1:07Introducing Tola Capital, Klarity, and enterprise AI investment
- 6:28Values, benchmark gaming, and dynamic evaluation
- 10:44Black-box models and Transformer-based LLMs
- 13:31Klarity's document workflows and customer-specific enterprise labels
- 25:26Closing call for responsible AI development
Talk transcript
- 0:00
[upbeat music] Well, thank you all for being here.
- 0:15
It is a small but excited group of people around evals. And, you know, evals may not be the most sexy talk at this conference, but it might be one of the most important.
- 0:27
And Nischal and I will spend time today speaking through why right now we think it's a seminal moment for evals and the import sort of the changes that we're seeing in the market as we move into these agentic frameworks, right?
- 0:41
So as things are more automated, as things happen in a more agentic way, we need to make sure that we understand what's happening with evals. We're all here obviously because we're very long AI, we care a lot about AI, but we have to really understand what's going on with it.
- 0:57
And really, if you think about the crux of it, can we evaluate the performance of this, our systems in relationship to our goals for those systems? And that's what we'll be walking through today.
- 1:07
And so about us, um, my name is Sheila, and I'm joined by Nischal. It's great. We're, we're an investor, investee pair, right? A VC and portfolio company. And I think there should be more talks like this because some of the most fabulous, uh, portfolio companies are not great at giving shout-outs to themselves.
- 1:25
So I get to give the shout-out to Klarity and to Nischal. He's the co-founder and CTO of Klarity. He's sitting up here, but he'll be speaking very soon as well.
- 1:35
Um, and Klarity just announced their massive, uh, seventy million dollar Series B financing on Monday. So we're gonna hear about that journey. It hasn't been a short journey, so I think that can be quite inspirational for a lot of you as you think about what you're building and the size and the scope of what you're building.
- 1:50
And so that company builds what we call exponential organizations. Well, what does that mean, right? Software transformed how we all work, right? This internal systems is always synced, always on nature of software systems for internal work changed the nature of the efficiency and effectiveness of buildings that, of, of c- businesses that we build.
- 2:11
Klarity is doing that for your external world, right? Most of your relationships with your customers, your partners are, are dealt with through documents. Those documents are one-off negotiated. They're one-off pieces of paper, right?
- 2:25
Klarity is automating all of that and then allowing you to build exponential organizations through that type of real-time relationship with those documents. So you'll hear more about that and how Klarity has implemented a number of eval systems later in this presentation.
- 2:42
Um, myself, I founded a venture firm called Tola Capital, uh, well over a decade ago now. And the reason I founded the firm was I was working at Microsoft.
- 2:51
I was running the database and developer platforms businesses at Microsoft. I co-led the company's enterprise strategy and was fighting the good fight against Team Windows to launch Azure. So being Team Cloud at a company that was based off of Windows was a really difficult thing, but it was fabulous as well.
- 3:10
And obviously, Azure's gone on to do pretty okay for itself. So the, the genesis of the firm Tola Capital really was, "Hey, how do we think about this next generation of applications that would be cloud-based?"
- 3:21
And now we're even more excited by the opportunity to bring this next generation of AI-enabled applications to the fore. So I'm gonna walk down memory lane for one quick moment and say, you know, what we saw in the advent of the cloud was very clear.
- 3:36
It would favor scale. The CapEx requirements, the physicality of building out those data centers was so expensive that you had to have a search business or an office business or a retail business to go fund that development, and then we would build on top of that, right?
- 3:54
The evaluation of those platforms was more straightforward. Am I offering you speeds and feeds? Am I offering you performance? And then, of course, you're layering on what functionality at what price I'm offering.
- 4:05
But the evaluation of the physicality of that world was more straightforward. Now, as we enter this AI world, we're saying, "Wow, there's," you know, some of the similar characteristics of large CapEx and large mega cap participation, right?
- 4:20
Where we have obviously these systems running on top and models running on top of clouds and being trained by the clouds and that, that compute cost, that inference, the ever so, uh, difficult to get your hands on chips.
- 4:33
The talent of all of you in the room, right? The AI engineers that bring this to be. But in addition to that, you have the proliferation of open source and open-source models, and just the ability of those models to do an incredible job at delivering incredibly complicated scenarios that are, you know, really, really, really strong contenders.
- 4:55
And so you have a proliferation of players, a proliferation of models. You have a deep academic heritage of those open source, um, and, and AI development, and you have great mega cap partnerships for those models as well.
- 5:08
So what, where are we going, right? We're-- AI is just-- it's, it's so important to think about AI as more than just another tool, right? It is a reflection of us.
- 5:19
It is a reflection of our understanding of the world, our intentions, our preferences, and at the end of the day, our society. And we'll talk a little bit about what that means individually and collectively, especially in this world where agents will represent more of us as individuals.
- 5:36
So let's talk about some of the shifts here. Um, you know, we, we started with AI eating everything on the internet. I like to call this all the garbage and all the gold on the internet was consumed by AI.
- 5:47
Then we moved into, you know, doing that gave us emergent behaviors. That's, of course, a great fancy way of saying we're not exactly sure how it knows what it knows, but we know it knows it, right?
- 5:58
This era then, then we have trained datasets, curated datasets, but we're moving into the era of self-taught, self-learning, and self-sufficient models. And so if you pause and think about that for a second, before we get into a truly self-sufficient era, we really need to fix evals, right?
- 6:15
Because it kinda gets to be too late in that era. And so that transition to full automation is happening faster and more aggressively than any of us thought. So what does this mean, right?
- 6:28
Narrow AI evaluation was, "I'm a hammer, you're a nail. I'm hitting you. Am I doing that right? You're an image. Am I classifying that image?" And I could, I could tell whether I was performing on that task in the right manner pretty s- in a pretty straightforward way.
- 6:45
Then you get into broad AI, right? And this radar chart speaks a little bit to how this evaluation becomes more multifaceted. What's my capability and intelligence? Am I serving a domain?
- 6:57
Do I understand that domain? What are my values, right? What are the... And whose values do I care about? Mine, yours, the user, society's, today, tomorrow? All of these questions come together.
- 7:07
Safety. Okay, do no harm. What does that look like? Pe- that could be different for people. And of course, the context and end-user awareness, which is often not discussed in an eval world, right?
- 7:19
Your end user is who you're delivering and developing these solutions for, but they're often an afterthought in terms of evals and evaluation. And then we have to do that across all of the different modalities of image and text and video, and kind of all of these together is creating a much larger problem around evals.
- 7:39
So today, the tools are still simplistic. We're gonna dive into kind of each of these areas of simplisticness and talk about some of the innovation happening to deliver this forward.
- 7:49
And then what Nischal will do is show us how Klarity has dealt with each of these, um, issues related to evals as well.
- 7:57
So benchmark hacking. I like to call this, you know, what you, what am I solving for? Solve for X is how a lot of these benchmarks and leaderboards work today from an eval perspective.
- 8:08
And it's interesting because scoring high on the benchmarks is pretty easy to do if you know what the, what we're solving for. If you understand X, you can do everything to solve for X, and you can look as intelligent as you want at solving for X.
- 8:21
But the real, the real reason is you may not understand anything about what's happening. You may be just solving for X. And that, that's a pretty scary reality on some of these things.
- 8:30
And so we say, "Well, these ALMs, LLMs passed an AP exam on a particular subject." Meh. Does that mean it was trained well on the questions that have heretofore come on that AP exam, or does that mean it actually understands the subject?
- 8:43
It's a really, really interesting question, but what we're seeing on a lot of the AI leaderboards is the best solvers for X are at the top of those leaderboards, and that's a problem.
- 8:53
So one area where we're seeing research happen, um, this is a Microsoft Research paper around dynamic benchmarks. Rather than saying, "Solve for X, can you identify this image?" They're creating dynamic datasets.
- 9:05
Basically with synthetic data where you can say, "Hey, we're moving objects around." These are not published. They are not, uh, public, right? So you don't know what the answer is coming into it.
- 9:17
They're first... Then, but then I can test you on spatial reasoning, visual prompting, object recognition. The images are changing, so the models have no ability to memorize those benchmarks.
- 9:29
This is an area, dynamic benchmarks in general, where I think we'll see a lot more work.
- 9:34
Benchmarks versus real-world scenarios is, is, is element two on this. You know, it's interesting, there's a lot of model creators that claim that their models perform very well and are very generalizable, and then you actually go and ask them a specific set of questions.
- 9:49
Even if I've done super well on MMLU and these things, and you say, "Okay, I'm gonna go test this logic and test this reasoning," and the answers are s- simply wrong, and they're wrong much and most of the time.
- 10:00
And, and this, this great example from Finance Bench, which was, um, Patronus AI's work and Stanford's work around saying, "Hey, you know, these basic financial questions were not answered when you actually benchmarked it on those real-world scenarios."
- 10:15
And the... It doesn't... The user is not at the center of our evaluation universe, right? We have to put the user at the center. We have to revamp the UX and feedback systems in order to really understand and capture more user needs.
- 10:29
Evaluation of UX is super difficult, right? And, and Nischal will talk a lot more about this in, in the case of Klarity, but that's why we're all actually here, to deliver that end-user value.
- 10:40
And so if we're not doing that evaluation, what are we doing?
- 10:44
Um, black-box models, right? So this is a problem that's, that's sort of hiding in plain sight, obviously. Evaluation is a proxy for the, for a task. Evaluation is seeking truth.
- 10:55
It is not truth. And so how do we really understand what we're doing in a world where we can't ma- mathematically represent what neural networks have learned and how they have learned that?
- 11:08
And so the area of AI interpretability is not new, but it's super important as we think about the rise of these complex AI systems. And so researchers are trying to open up these black-box models, show us the how and the steps in this.
- 11:23
And, and the Transformer-based LLMs, we're really tracing information flowing through the network. And so this is a good example of, you know, sort of asking questions and seeing, "Hey, how am I answering it?
- 11:34
How are these pieces coming together?" We'll see a lot more of this interpretability work in the RL, um, the reinforcement learning work that's happening with the current GPT models.
- 11:43
It's super, super, super important that we get this right now. Now, values, right? When we create AI models, we do instill our own values in them, whether we want to or not, right?
- 11:56
And these evaluations have to understand the technical capabilities, but also the underlying values that we're putting into the models. And, and this is... There's a lot of questions on this, and I have way more questions than answers, as I think we all do.
- 12:10
But it's, you know, are we creating values that benefit humanity? As we go into this agentic world and you have your own model, right, does that just reinforce you?
- 12:22
Is that a good thing, right? So the, the, the model of Sheila is gonna believe more of what Sheila believes and get deeper and deeper into that Sheila-ness. Is that a good thing?
- 12:33
Right? We've seen the echo chamber of ourselves in news. We've seen the polarization that this has caused in our society. We should be asking questions about where we're going to get to as we enter this agentic world.
- 12:47
And this is kind of a, you know, it's a question, right? This isn't, you know, a simple two-by-two to ask the question, do we want to appease or challenge users?
- 12:56
Do we want you to choose whether you are appeased or challenged? Do we wanna ground AI in present or future societal values, aspirational values versus present values? Do we wanna encourage users to select their own values?
- 13:11
Should model makers be responsible for this? Lots of questions, less answers. Now we're gonna turn it into a real-world example, leveraging Klarity to talk about how the company is replicating human cognition.
- 13:25
Nischal? [audience applauding]
- 13:31
Uh, thanks, Sheila. Thank you for having me here. My name's Nischal. I'm co-founder and CTO of Klarity. Uh, as, as Sheila mentioned, what we do is we automate back-office workflows, things that traditionally required large teams of offshore, of offshore humans, and throughout human history have been impossible to automate because they're cognitive and they're non-repetitive.
- 13:50
You can't just write down a simple set of steps and automate these workflows, and that's what we've been working on for the last eight odd years. Um, predominantly these are document-oriented workflows, so it's some kind of PDF that a human being is ha- having to read as part of a company's back office.
- 14:06
Uh, one example of this is revenue recognition, typically part of an accounting team, um, matching invoices to purchase orders, so actually aligning two p- two documents and matching them to each other, um, processing tax withholdings that often come in many different languages.
- 14:20
But you probably get the sense. You get PDFs that are completely unstructured. Somebody has to go through them because it's a tightly regulated process, uh, part of a finance and accounting team, and today this is done completely manually.
- 14:33
Um, there was actually a really great keynote yesterday, uh, by, by a gentleman who pointed out that document processing tasks are generally super tough for LLMs for a variety of reasons.
- 14:42
Uh, if, if the page is like rotated or the scan quality is bad, if you have like graphs or images, if you have tables inside of these, um, it is very, very tough to get this to work, and it's not the kind of thing you can just give to ChatGPT and it happens out of the box.
- 14:57
This is basically exactly what we do, and we spend, you know, a long, long time and tens of millions of dollars building a stack that's able to deal with these kinds of documents.
- 15:06
Um, these are some of our customers, mostly B2B SaaS companies, um, mostly in a five-mile radius around where we are right now. Uh, and we predominantly serve their finance and accounting teams, although we're expanding quite a bit beyond that.
- 15:18
Our journey has been a little bit unorthodox. It's, it's kind of like that, um, you've probably seen the startup curve of like the trough of disillusionment and you get the TechCrunch article in the beginning, something like that.
- 15:27
Um, we, we founded the company in twenty sixteen. Um, we pivoted four times between twenty sixteen and twenty twenty, and when we were kind of at our wits' end about to give up, we did one final pivot, um, and focused on finance and accounting teams and, and found pretty strong product market fit there.
- 15:43
Uh, in the last couple of years we've completely re-platformed around generative AI, and that's really been a shot in the arm for the company, uh, and have gone on to raise, uh, ninety million-plus off of that.
- 15:54
But before generative AI, we were in traditional ML for more than six years. Um, we, we like to say that we started an AI company five years too late.
- 16:02
Um, and so a lot of the concepts around evaluations were, you know, they seemed pretty natural to us, and we didn't really see why generative AI had to be any different.
- 16:10
Uh, in supervised learning, which is a majority of what we did pre-gen AI, you have your train split, your test split, uh, the metrics are pretty well defined, F1, ROC curves.
- 16:20
Um, there are these like shared benchmark tasks like sequence labeling and SQuAD, and these were all very well understood, so it wasn't immediately apparent to us why generative AI has to change any of these things.
- 16:31
Um, and in, in the two years since then, we've, as I mentioned, re-platformed to generative AI, and that's gotten us quite a bit more scale. We've processed more than half a million documents for customers.
- 16:41
Uh, we have more than fifteen unique LLM use cases running in, in production today, uh, and more than ten LLMs under the hood. Oftentimes a single use case has multiple LLMs working together.
- 16:54
But all this kind of begs the question, why does eval for gener- generative AI have to be any different than traditional ML? And this is kind of the journey that we've been on in the last couple of years.
- 17:03
Um, the first, and a, a number of speakers have spoken about this, so I'm not gonna go into too much detail, is non-deterministic performance. Very challenging for us 'cause it's a very cognitively demanding task, so if you upload the same PDF multiple times, very likely you'll get completely different responses.
- 17:17
Um, new user experiences. I, I think fundamentally when you move from like discriminative models, classification, regression, random forest, to generative models, the types of experiences you can provide your users expand quite a bit, and these new experiences are just much harder to evaluate.
- 17:33
So a couple of examples from what we do, um, we have this tool called the Architect where users will record their business workflow, like just them doing their job, upload it to Klarity, and then we'll create a business requirements document out of that.
- 17:46
So it's typically like a ten-page Word document. It has a flowchart, images, very comprehensive, like what a McKinsey would build for you. It's not at all intuitive how to eval something like that.
- 17:56
Another example, part of our product is, um, you can do natural language analytics, so you don't need like a BI specialist. You could just ask it questions like, "Hey, how's my contract population evolved over time?
- 18:05
Give it to me in a stacked bar chart." It'll do that for you. Uh, cool feature, but like how do you eval something like this? And then a more traditional document extraction task.
- 18:14
You are trying to find certain parts of a document that are consequential in some way. They could be tabular, they could be legalese buried inside of a document, and you need to know what accuracy you're doing that with.
- 18:24
So this is why new experiences, while very rewarding to our customers, have been very challenging to us from an eval perspective. Uh, the second is the rate of feature development.
- 18:33
So in our previous deep neural net world, DNN world, a feature took, like, five to six months to build. We literally had teams that would annotate data, dedicated teams to annotate data, GPUs, most of which are gathering dust now to train.
- 18:45
And so end-to-end, it was, like, six months. So if it took, like, two to three weeks to build evals, thoughtful evals, that was completely acceptable. It was a pretty small fraction of the feature development time.
- 18:54
Now what we're seeing is we can get features out of the door in days. Um, within twelve hours of the ChatGPT launch, we launched Document Chat as a feature.
- 19:02
And so in, in that world, it's unacceptable that it takes a week, two weeks to build evals. It now becomes the bottleneck to feature development.
- 19:11
And the third is, uh, benchmarks diverging from performance. Sheila talked about this a little bit. Um, what we've seen at Klarity is even slight differences in MMLU actually make a very big difference to us because we're kind of at the frontiers of human cognition doing something that's very challenging for most human beings.
- 19:27
We've seen many cases where, not gonna name names, a model is supposed to be better in terms of MMLU, and then it's totally not on our internal benchmarks. There's a variety of factors that go into it, but it's made testing very chaotic and challenging for us.
- 19:41
And so the question then is what can be done? And I'm pretty sure most application developers are running into these problems or something similar. Uh, I'm not gonna say that we have the silver bullet, and I would say we are nascent in our eval journey.
- 19:51
But here's a couple of things that have worked for us. So the first is, like, really give yourself the gift of imper-im-imperfection. Don't put the threshold of high-quality evals at the beginning of the feature development life cycle.
- 20:04
Um, good evals are not trivially cheap to build today. They probably are not gonna be for the foreseeable future. And so we really try to front-load user testing. What we found with these generative AI features is we have to think of each feature almost as its own product market fit 'cause we're delivering experiences that people have never
- 20:20
had before, chatting with a document, natural language analytics, uh, watching videos automatically. And so there's a lot of user experience risk actually baked into each of these features compared to traditional machine learning, like a recommendation engine where you have fi-- you have fairly high conviction that the form factor is correct.
- 20:37
So what we try to do from a development perspective is front-load the UX risk, back-load building out evals. Uh, of course, I'm not advocating that you go into production and scale without building evals, but a lot of features die at the UX stage itself.
- 20:50
Let them die before you build out evals.
- 20:53
But once you decide to go into production, kind of our framework for this is you need to think backwards from the user experience, not forwards from what is easy to measure.
- 21:02
So let's not just say we wanna measure F1 and hope that user experience correlates to that. We wanna look at the end-user value, move backwards from there. Not every eval can or should reflect the entirety of the experience, but you want your evals in aggregate to be reflective of the user experience.
- 21:18
Um, a simple exercise that we do for this when we're building features is we kind of walk down the stack of what is the end-user outcome, the business value we're driving?
- 21:24
What is a good indicator of adoption/utilization? At the lowest level, how are we measuring health of this feature? It could be something as simple as, like, JSON adherence, variability in the output, and so on.
- 21:35
So in practice, this is what it looks like. Every customer is basically giving us their own set of labels. You could think of this as, like, bespoke enterprise AI.
- 21:42
And so we'll actually annotate data for that customer as part of our UAT process, build out use case-specific accuracy metrics for them. Are they trying to do matching? Are they trying to do extraction?
- 21:52
Um, and then there's still user feedback as, uh, uh, user feedback as kind of there to close the loop. But i- we think it's very dangerous to assume that the absence of user feedback is positive feedback.
- 22:03
So we try to put the majority of the onus on ourselves to be rigorous about metrics, um, and we have various tools to monitor data drift. Are we getting documents that are very different from the population we've seen so far?
- 22:14
Um, we've invested a lot in our synthetic data generation stack. We actually surveyed, like, six-plus providers in the market, tried a bunch of them out, um, were not too happy, and so ended up building our own synthetic data generation stack.
- 22:26
Everything that you see was synthetically, um, generated. The way that we think about this is once a use case becomes large enough within the company, we have enough customers doing it, we wanna, uh, we wanna invest in customer-agnostic evals.
- 22:38
That's when we go down the, the kind of path of synthetic data. And we have a team where part of their job is just monitoring that the synthetic data is distributionally similar to what we're seeing from customers.
- 22:50
So we haven't yet cracked the problem of doing this in an automated, fully quantitative way. The other kind of sanity check that we have is, of course, looking at accuracy scores and making sure that our models are not excessively or underperforming on synthetic data.
- 23:03
Um, the next, the next little trick that we use is kind of reducing the degrees of freedom. So this is kind of part of our architecture where we have numerous features, each of which require a custom prompt for each customer.
- 23:15
Um, so in this case, these are four features: free text, tabular extraction, matching, table composition. Each of these has a prompt for each customer, so you can imagine over a hundred customers, uh, you could end up with, like, literally thousands or tens of thousands of prompts.
- 23:27
So instead of manual prompt engineering, we have these APE things, automated prompt engineers. Now, the trouble is different LLMs perform
- 23:35
d-- uh, have different levels of performance on different APE tasks and on different customers. So you get, like, this exponential explosion in complexity, and what we did instead is we just said, "Well, we're seeing quite a bit of commonality in which LLMs do well on APE tasks, so let's just use one LLM for APE tasks."
- 23:52
Not permanently, and we'll continuously reevaluate as new LLMs come out, but if we can fix this dimension of freedom, it gets a lot easier to iterate. So could we eke out another percentage point of accuracy if we didn't do this?
- 24:03
Yes. But building that grid search infrastructure is just too expensive, and we don't think it's a good investment of time. So this is kind of another trick that we use to just...
- 24:11
almost like dimensionality reduction at a project, project management level.
- 24:16
And the last thing I'll mention is, like, identify future potential. I think people spend a lot of time, and rightly so, building evals for what their company does today.
- 24:23
But you should almost have a wish list of what are additional use cases that you want to grow into over time, three, four, five years down the line, because we are frequently surprised that technology is evolving faster than what we can see.
- 24:34
But it is too high of a bar to say we wanna have like MMLU or BBH-level metrics for future workflows that nobody has asked us for yet. And so what we do is we have these very scrappy, um, kind of future-facing evals.
- 24:47
Uh, for example, when GPTV came out, in about an hour we were able to say, "All right, this is our mental model of how it's gonna do. It's good at check marks.
- 24:53
It's maybe not so good at pie charts," et cetera, et cetera. And so that, I think, muscle of just building this organizational mental model very quickly in a scrappy way is also something that's been helpful to us.
- 25:05
Um, I wouldn't be a startup founder if I didn't end with a shameless plug. Uh, as Sheila said, we, we just raised a $70 million Series B. We are hiring across the board, AI, back-end, front-end, go-to-market roles.
- 25:15
So if any of this sounds interesting to you, I'd love to chat. Thank you so much, and I'll hand it back to Sheila. [audience applauding]
- 25:26
Thank you, Nischal. Uh, I have to say, working at Klarity is a dream, so, uh, anyone who's interested, please see, see either one of us afterwards. Um, this, so this is my call to action.
- 25:36
I've seen these paradigm shifts happen in the past, and I think one thing that's important is to remember the people that are early to this sort of AI revolution that's happening, these people matter disproportionately, and these people are all of you.
- 25:52
And so understanding and owning your power as you think about what happens with the next generation of evals, of integrating values into things, it's, it's actually, it's more than just sort of words on a slide, right?
- 26:06
There is a real empowerment and opportunity to spend time thinking about this to get this right for the ecosystem and for the industry. So what do we do? We innovate.
- 26:18
We reinvent benchmarks. We don't let the leaderboard benchmarks stick that are, that we know are just kinda BS, right? There's just, there's too much happening today that isn't truly understanding how these systems work.
- 26:31
We bring depth into the evaluations. We bring multifaceted nature of these evaluations together. We do so as a community, right? I think it's incredibly important that we bring curiosity, empathy, help to one another.
- 26:47
Like, you know, I, I love the fact that Klarity wanted to come and talk about their journey, what they did that was right, what they did that was wrong.
- 26:53
We share with one another on that. And I think that we introspect our own value systems, right? I think that today we are looking at, um, such a pace of innovation that we really need to think, what does it mean to drive this?
- 27:08
How are we the trailblazers for this happening? You know, I, I don't know if folks have read this book, The Alignment Problem. Brian Christianson wrote a great book, uh, I think it was last year, and it was one of my favorite reads of the year.
- 27:19
And he, he had this quote around sort of if we do see AGI, it will be an interesting mirror of society. It will tell us which values are uniquely human in nature.
- 27:29
Well, those change over time. How do we want to, how do we want to drive a society that has values that we can be proud of as these trailblazers?
- 27:39
So feel empowered, be curious, stay empathetic, kind, fair, intelligent, right? That's what we're doing. We're printing new intelligence for the world, and this is both your opportunity and your accountability.
- 27:50
And so thank you all for leading the future of AI. [upbeat music] [audience applauding]