AI Engineer Code 2025
Small Bets, Big Impact: Building GenBI at a Fortune 100
About this talk
Asaf Bord explains how Northwestern Mutual is developing GenBI, an AI-powered business-intelligence agent, through incremental projects that earn funding and trust inside a risk-averse enterprise. The approach starts with real enterprise data, existing verified reports, dashboards, metadata, and BI practitioners; rolls out first to BI experts and then business managers; and adds governance, orchestration, and contextual interfaces while managing production complexity and competition from products such as Databricks Genie.
Chapters
- 0:42Introducing GenBI and Northwestern Mutual
- 2:23Enterprise risk, funding, and the path to production
- 6:08Practitioner buy-in and crawl-walk-run rollout
- 8:54Verified reports, dashboards, and BI-agent metadata
- 14:09Investment risk, Databricks Genie, and governed orchestration
- 22:07Productivity, pricing, and closing remarks
Talk transcript
- 0:00
[upbeat music] Doesn't this look like something's gonna drop from the ceiling? [laughing]
- 0:24
Like a ground zero type thing? [sniffs] Be honest, like who has a buzzer that if some- I really suck, they press it and everything falls down through the trapdoor? No?
- 0:34
Yeah. Be careful. Yeah? Okay, who was it? Okay. You tell me if I'm doing okay or if I should take a couple steps back, right?
- 0:42
So hi, everyone, I'm Asaf. Um, and I'm here to talk about GenBI and kinda... First disclaimer, this presentation was not created with GenAI. Um, to be honest, I actually started doing it, uh, with, uh, GPT o3 back in August.
- 0:59
Uh, and then I did kind of a first draft. And then a couple of weeks back, I wanted to come in and refresh it before the conference, and then GPT-5 took over, completely messed up my slides, so I ended up doing it manually, kinda old-fashioned.
- 1:13
So if I'm missing like an em dash somewhere in the middle, let me know after, okay? [laughing]
- 1:19
Uh, so first of all, a bit of housekeeping. What's GenBI? So it's a fusion of GenAI and BI. It's basically an agent that helps people answer business questions with data like a, a business intelligence person would do in real life.
- 1:33
Uh, the reason that we're pursuing GenBI is really because of the data democratization that it can bring, right? So having access to data at your fingertips without having to be reliant on a BI team that helps you find a report, figure out what it means, uh, understand your world before they can even give you any kind of
- 1:51
input. Uh, so that's GenBI. Uh, a bit about Northwestern Mutual. That's where I work. So we're a financial services, life insurance, and wealth management. Been around for a hundred and sixty years.
- 2:04
Uh, some very impressive numbers there. But first of all, I want to say, why is Northwestern Mutual a great place to do GenAI? We got a lot of data.
- 2:13
We got a lot of money. We got a lot of use cases, and we got access to some of the best talent, uh, anyone can dream of. Really, truly humbled by the people that I get to work with.
- 2:23
Um, but on the flip side, why is it hard to do GenAI at Northwestern Mutual? Because it is a very risk-averse company, right? If you think about it, our main motto is generational responsibility.
- 2:37
I call it don't F shit up. Uh, because what we end up selling to people is a decades-long commitment, right? You buy life insurance now, uh, if you stay with us until it comes to term, so to speak, that can be twenty, forty, eighty years down the line, depending on when you buy it and how long you
- 3:00
get to live. And so stability is something that's very important for us because it's important for our clients. So how do we balance stability with innovation? That's what I want to talk about today.
- 3:12
Um, and really the four main challenges that we had when we even came up with the idea, kind of a pie-in-the-sky GenBI concept. Uh, first of all, no one's done it before, right?
- 3:25
Truly, no one's done GenBI in this fashion in the past. Uh, secondly, and this was really a preference for us, we wanted to use actual data that's messy because we knew that those were...
- 3:38
that's where the real challenges are gonna be, right? Understanding actual messy data for a hundred and sixty-year-old company and how can we perform well within that ecosystem. Um, the third was kind of a blind trust bias.
- 3:53
So, um, the bias-- The trust that we had to build was both with the users but also with the leadership of the company, right? How can we bring accurate information, accurate answers to people when, uh, all of these things that we know about and everyone's talked about is, is just out there, right?
- 4:11
No one's blind to the trust barriers. No one's blind to the accuracy barriers. So how do we convince that this is actually something that we can trust in the company?
- 4:21
And lastly, um, but really firstly, when we go to approach this from an enterprise perspective, budget impact, right? How do we convince someone in a leadership, uh, organization where risk-aversement is ingrained in the DNA to even invest in something like this that no one's done before?
- 4:42
We don't really know how we would do it. Uh, we're not even sure how it would look like when it comes to term.
- 4:48
Uh, so I'll start kinda one by one. Uh, and first of all, really talk about why we chose to use actual data, uh, and not synthesized data or cleansed data.
- 4:58
Uh, so really it's about making sure that we understand the actual complexities that we will have to face when we eventually want to go to production, right? We know that, you know, building, uh, POCs and demos is so easy, but the gap from POC to production is so broad, uh, especially in this GenAI space, especially because we
- 5:17
don't know upfront how to design the system, what we would expect it to behave like. So making sure that we operate with real data just gave us that extra confidence that when something works in the lab, it's very likely to also work in reality.
- 5:31
Uh, but also, and, and maybe not, uh, in the least less important, is that we got to work with actual people who work with the data day in and day out.
- 5:41
And that gave us two things, okay? First of all, subject matter expertise, which are super critical for us to be able to validate that the system is actually working, gave us a lot of real live examples of what people are actually asking in a corporate and what people have answered to them.
- 5:56
So basically the evals, right? And all the testing and stuff. Uh, but at the end of the day, it also brought the business to be a part of the research project itself
- 6:08
And they became kind of bought into the idea as part of the process. So we didn't just test something in the lab and then had to convince someone to go ahead and use it.
- 6:18
The end users were part of the research process itself, and so when eventually it matured enough so we can take some of that to production, they were already there, and they actually were pulling that.
- 6:29
They told us, "We wanna take this. How can we wrap it? How can we package it, uh, quickly enough so we can put it into practice?"
- 6:38
Uh, and the next part was really about building trust. Uh, so this is about building trust, first of all, with our management team, right? Now, I don't know about you, but last time that I got a million dollar to do a research project that I wanted in a pie in the sky idea, I, uh, woke up from
- 6:54
the dream, and I realized that this is not how things work in reality. You don't just get a million dollars and go ahead and try something out. Uh, you had to show that you know what you're doing, and part of what we did, it's kinda listed out here, but obviously, you know, we did all the regular stuff,
- 7:11
right? We worked in a sandbox environment. We made sure that we're not using actual client data. We made sure to put in all the security risks aside. But, uh, one of the first approaches that we said we're gonna take is we're not just gonna build a tool that's gonna be, uh, released to everyone, right?
- 7:28
We understood very quickly that, um, how people interact with the tool, their ability to verify that what they're getting is right and also give us feedback, changes dramatically depending on their expertise and understanding of the data.
- 7:44
So we took that crawl, walk, run approach that basically said, "We're first gonna release it to actual BI experts," right? People that would be able to do it on their own and know what good looks like when they get it, and we're just gonna expedite the process for them, kinda like a GitHub Copilot.
- 8:01
The next phase would be to bring it to business managers, and again, people who are closer to the BI team, but when they see a mistake, they can pretty much figure out that what they're seeing is wrong because they're used to seeing that on day-to-day basis.
- 8:16
Um, and they will... might be less sensitive to these types of mistakes and be more inclined to give us that feedback instead of just, you know, dumping it aside and never using it again.
- 8:26
Giving this type of tool to executives in the company, I don't even know when we're gonna get there, right? Like, an executive, they want clear, concise answers that they know they can trust.
- 8:37
We're definitely not there yet. I think that's the vision, uh, at some point in time, but the system is not accurate enough for us to get there. Maybe it never will be.
- 8:46
Um, another way that we... Another lever that we kinda used to build inherent trust in the system is that we said, "Well,
- 8:54
in the get go, we're not gonna even try to build SQLs," right? This is very complex. This is very hard even for a person. So we said, "Step number one, let's just bring information that is already in the ecosystem that's already verified," right?
- 9:11
So we have a lot of s- uh, certified reports and dashboards, um, and actually, in the conversations we had with some of the BI teams that we work with, they told us, "Guys, like, eighty percent of the work that we do is basically sending people to the right report and helping them figure out how to use it."
- 9:27
So the report is already there. Um, and that, again, built some inherent trust into how we architected the system because we said, "We're not gonna make up information. We're just gonna deliver you the same asset that you would have gotten anyway, just in a much faster, much more interactive way."
- 9:44
Uh, and that was the alignment of expectations that we did very upfront with the, uh, users and also with the management team.
- 9:51
Now, [clears throat] the biggest, um, process or kind of the most important approach that we took when, uh, approaching our leadership team and convincing them that we wanna do this was to create a very gradual, incremental process that gave them a lot of visibility and control.
- 10:12
Uh, and it was very important for us to build incremental deliveries throughout that process so that, uh, not only o- they have the, the visibility into what are we funding now, what do we get out of it, they actually had business deliverables they could realize potential from throughout the process and at any point in time, they could
- 10:33
pull the plug, right? And say, "Okay," like, "it's not working well," or, "We got enough out of it," or, you know, "The next phase is so, you know, unknown and long that we don't wanna further invest in it."
- 10:43
And this is how we basically broke it down. So phase one was just pure research, right? We kinda did the shift from natural language to SQL. We figured out how to write responses.
- 10:54
We figured out how to understand questions that's coming in, just kinda setting the stage. Phase two was about really understanding, okay, so what does good metadata and good context look like in the perspective of a BI agent, right?
- 11:08
It looks very different if you're just chatting with something or if you're trying to do a RAG with, you know, unstructured data like documents and, uh, business knowledge and stuff like that.
- 11:18
And this phase on its own already had, uh, impact on the business because when we define what good metadata looks like for NLA and a... an LLM, uh, we could immediately apply that also to just the ecosystem of data users across the enterprise.
- 11:33
Um, and by understanding how to extract LLM from the information, we could also fi-... How to extract metadata, sorry. Here's where the trap door comes into play, right? Um, we could also project that on how or what good metadata looks like for humans interacting with the data.
- 11:52
We have another initiative around semantic layer going on, which tries to model exactly that, and this provided a very valuable input to that initiative as well. But the immediate next step was basically just doing this kinda, uh, multi-context semantic search, right?
- 12:07
People coming in, asking different questions. And having the system figure out what's the right context, what's the right information we need to g- uh, bring them, and this is something that could already be packaged as its own product and delivered, uh, and basically just do kind of a data finder and data owner finder, which is something that
- 12:28
could take anywhere between two to maybe four weeks in an enterprise like Northwestern Mutual, just finding what data exists and who owns it, so I can start talk, uh, the conversation with them.
- 12:39
Um, and the next layer was really about pulling in information and starting to do some light pivoting around the data. Um, each one of these steps, as you can see, also created an input to the ste- to the following step, so that the research itself was kinda self, um, self-propelling, and there were incremental outcomes coming out of
- 13:00
each one of these phases. Uh, the next one is more kinda setting it up for enterprise-level usage, so understanding roles of in, uh, of different users coming in, what they may be asking about, what type of access we wanna give them, et cetera.
- 13:14
And eventually, and this is still some ways to go ahead, uh, building kind of a fully fledged GenBI agent, which doesn't only quote information from existing reports, but, uh, can actually run SQL queries on its own, uh, pull in more data, do more sophisticated joins between different data, so it can answer more complex questions.
- 13:34
So that's the roadmap, right? That's kinda the high level plan. Now why did that work well? Kinda quickly summarizing, we talked about, uh, so we get value, uh, early and we get value often.
- 13:45
Each one of this was a six-week sprint, at the end of which we ha- had a very tangible deliverable coming back to the business that we could decide to productize.
- 13:55
Uh, and at any point in time, we could decide how we wanna move forward. There was transparent progress. There was incremental business value. Uh, each one of these steps allowed us to learn something that helped feed the next step.
- 14:09
And maybe the most important part, and that's the bottom line here, and that's the part that executives really look at, how do we control the risk in continuing to invest in this type of research project?
- 14:21
And this is really about eliminating things like sunk cost bias, right? We already paid, you know, you know, whatever, a million dollar. Let's just get through the project, see what we get at the end.
- 14:31
This eliminates the, uh, uh, fear of, of competitors coming in, and maybe we don't need to continue investing in this, right? So everyone in the industry is researching GenBI, and there are solutions like Databricks Genie that are coming up, and they're getting better and better.
- 14:45
Maybe at some point in time, it's better for us as an organization to actually adopt Databricks Genie. But at that point, again, first, it's much easier for us to pull the plug in the funding, but we already have a good understanding of what good looks like.
- 14:59
We have benchmarks that we used for ourselves when testing our own system that we can test a third-party solution with, and we know what to expect, right? We know what works.
- 15:09
We know what doesn't. We know what a kinda fluffy demo from a vendor would look like, and we know where to drill in to ask the tough questions.
- 15:18
So let's see kinda what it looks like under the hood and how we productize different elements, uh, of this architecture. Uh, and maybe kinda very quickly, why can't we just do it with, uh, ChatGPT?
- 15:29
So, you know, just dumping a schema into ChatGPT doesn't work. Usually, schemas are very messy. It's not, uh, easy to understand the context and the meaning of things. Uh, and eventually, governance is super important, so there was a lot of governance built into the architecture that was very hard to apply on ChatGPT from the outside.
- 15:46
But even solutions like, you know, Databricks Genie, third party, much harder to govern from the outside than from the inside, but still TBD.
- 15:55
Uh, so the stack kinda looks like this. Uh, we have a data and metadata layer that we produced. We have four different agents that are running across the pipeline.
- 16:04
A metadata agent that understands the context, a RAG agent that finds the different reports, an SQL agent that can pull more data if we need that, and then eventually what we call a BI agent that takes all that information and delivers an answer to the question that was asked.
- 16:18
On top of that, we slap governance and trust and orchestration, and eventually some kind of a contextual UI. Um, and this is how the flow goes. So when a business question comes in, we, uh, push it into the orchestrator and basically decides how to facilitate the process.
- 16:37
The first thing that we do is understanding the context, so that's where that metadata agent comes in, works with the catalog, works with all the documentation that we have across the system to understand what we're being asked about and what's the relevant information to share.
- 16:50
Then we go to the RAG agent, which tries to find an existing report, again, out of a list of certified reports that we know are allowed for people to use, and people have spent a lot of time fine-tuning them and making them as accurate as possible.
- 17:05
If we can't find the report or if it's not exactly what we need to, um, to use, that's where we go to the SQL agent that basically tries to create a more, um, exact query or a more elaborate query.
- 17:18
And even if the report that we have is not usable as is, it gives us that initial seed of a query that we can then expand on rather than having to build one from scratch.
- 17:30
So it's kind of like a few shot, uh, example, but in this case, the example that we give is very, very close to the actual result that we're expecting to get.
- 17:41
We then execute it against the database, pull-- and push it into the BI agent, which gen, with, which gen- uh, translate that to a business answer, and not just dumping data back on the user, and this is what goes into the final answer.
- 17:55
Now, there's obviously some kind of a loop that says, "If I'm in the same conversation, I'm probably talking about the same data, so we don't have to talk about this or do this again and again."
- 18:04
Now, each one of these three components, each one of these three agents can be packaged as its own product And delivered to production with a very tangible and actual impact on business metrics.
- 18:20
Okay? And that's the kind of beauty of this, uh, approach, that after we productize each one of these, we could have basically said, "Stop," or, "Let's move forward." Uh, and just some giving bottom line numbers around some of these.
- 18:35
So just the RAG agent that pulls the right report, uh, allowed us to take about 20% of the overall capacity of the BI team that basically said, uh, all we do is just share the right report with the right person.
- 18:52
So we were able to automate around 80% out of those, uh, 20%, and we're talking about a team of 10 people. So roughly two people full-time job, all they do is find the right report and send it to the right person.
- 19:08
Uh, the metadata understandings that we got from learning how to interact with the data through an LLM allowed us to run A/B test in a, in the semantic layer project that we did, and that allowed us to prove back again to the senior leadership in the company that there is value and tangible value, measurable value in enriching
- 19:29
metadata. And we did that basically by running, uh, a, a battery of questions, um, against a database that had good metadata and one that didn't have good metadata, and we show how much better an LLM performs when having the right metadata in place, so basically proving the value of something that can be very fluffy, like, "Hey, let's
- 19:50
bring in more documentation into the code." Uh, right now we're experimenting with, uh, the data pivoting bot. Uh, so once you have a dashboard or a report, be able to change the time horizon, some of the views, some of the segmentations, and the groupings of the data, again, kind of real time without having a person do that
- 20:09
for, uh, a business stakeholder. And some of the next steps is really evaluating the tools that are out there for, uh, GenBI, like Databricks Genie, for example. And we're gonna go into a much more rigorous process of enriching our catalog with metadata and documentation, and that's also gonna come out of a lot of the learnings that we
- 20:27
got from, uh, the research that we've done. So even if we don't end up writing a GenBI agent full-fledged end to end, we already got a lot of value back from this, and this is really what allowed our, uh, senior leadership team to continuously invest in this project quarter over quarter.
- 20:47
One thing that I wanna wrap up with is just a couple of thoughts I had about the future. So, um, I think we talk a lot about how to prepare data.
- 20:56
I think that's gonna be a huge area in the market, and there are gonna be probably a lot of companies and tools that are gonna help us with that.
- 21:03
Uh, building very specific, task-specific models and applications, I think a lot of startups and companies are gonna come up from that area. Uh, Copilots is really ma-making sure that we meet the users where they are, uh, and securing of models, obviously a very big thing.
- 21:20
The last thing that I wanna... the, the one I wanna focus on the most, 'cause that's kind of a recent thought that came to me a couple of weeks ago, how we do pricing of SaaS in the GenAI era.
- 21:31
Uh, this is really about the fact that one individual person today can be 10X more effective, uh, than they used to be in the past. And then do we price, uh, software based on seats, or do we price software based on how much they used it, or do we price software based on the value that they got
- 21:50
out of it? Uh, Salesforce is already experimenting with that, so the, the Data Cloud product at Salesforce is starting to be, uh, usage priced and not seat priced, and I think this is gonna have a big impact on just the, uh, kind of SaaS economics worldwide.
- 22:07
Uh, and it, it, it doesn't even matter if the product itself is GenAI. It's really about what does the person using the product can do, and what can they do in their other time, uh, and whether it still makes sense to price it by how many employees you have or how much work you get done with the
- 22:23
employees that you have. That is me, and thank you very much for listening, and thanks for not opening the door on me. [upbeat music]