AI Engineer World's Fair 2024
AI Platform Engineering
About this talk
Patrick Debois applies lessons from the emergence of DevOps to explain why organizations need dedicated AI platform teams as generative AI moves beyond isolated data-science initiatives. He describes enabling application developers through shared infrastructure, Kubernetes and CloudOps coordination, caching, feedback systems, local experimentation, engineering rigor, risk assessment, and centralized governance with self-service access.
Chapters
- 0:00Why AI needs a platform team: lessons from DevOps
- 2:50AI engineering identity and organizational adoption
- 5:43Shared infrastructure and enabling application developers
- 11:22Caching, feedback, infrastructure, and local experimentation
- 17:33Engineering rigor, testing, and AI risk awareness
- 26:42Audience question and centralized self-service governance
Talk transcript
- 0:00
[upbeat music] Thank you. Especially after lunch, that applause tell, you know, is, tells double, so.
- 0:19
So yeah, you need a platform team. Um, so usually I'd skip over this slide, but this time I'm gonna explain you my background because it's important to understand what the next things is that I'm completely biased in this, what I'm gonna talk about.
- 0:35
So why am I biased? In 2009, I organized a small conference called DevOpsDays. There were 60 people,
- 0:44
and this is how the world kind of got into DevOps as a word. So my position has been privileged to be there from the beginning and seen 15 years of a new kind of phenomenon, discussion, transformation taking root across the whole enterprise.
- 1:02
So that's kinda like why a lot of what I'm talking about is things from the DevOps space, what we learned here, and I'm trying to apply this to whatever is the new thing we're doing.
- 1:13
Um, when you apply your old paradigms to the new world, it could either be correct, and it could be totally wrong. So I'm explaining what I think that I'm seeing.
- 1:26
Um, and one of the particular interest is after seeing DevSecOps and kind of whatever ops in the world, I especially like this part of the GenAI, uh, in particular for that it's not...
- 1:40
Like, for me, it's automation intelligence. So a lot of my work on DevOps was automating pipeline, making them more robust, and kind of making sure we're delivering the value on this.
- 1:51
So that's why kind of I, I get excited about this new world.
- 1:57
So over the years, I have [REDACTED:physical_attribute], right? So I've seen my fair share of new things coming in the industry, and the pattern is usually
- 2:10
you have one team that's the, the pilot team, and something new happens. We're gonna try this out. Then you scale this out into two or three more teams doing this, and then you extrapolate some learning, some patterns, and eventually you wanna scale this out to all the teams.
- 2:28
This is what happened with cloud, the infrastructure. You know, initially it was a small thing, uh, but then eventually this led into an abstraction layer of us kind of making this easier for other teams so they don't have to understand the whole space, but they can move faster within their domain.
- 2:50
Like DevOps, everybody was saying DevOps is a bad name. W- we kind of hate it, but in the end, the industry stuck with us. AI engineer is a bad name.
- 3:03
What does it mean? Nobody knows. But it's also the advantage that you can bring everything to the table and let it grow. So if you make it too defined, then we are not evolving that way.
- 3:16
That... So that, that, that has kind of evolved in... The first there was the agile team. They were, like, really proud as the change agent, happened to ITIL, and we're kind of bringing this in.
- 3:28
I think what I learned over the years of DevOps is the name does not matter that much. In the beginning, it's really important that you have a label that you can search and find all the emerging stories of people doing this thing.
- 3:42
So that's how I look at AI engineer as a term. If you search for this, it's gonna be on your job post. It's gonna be on your website. This is likely the term that kind of fits into f- us finding each other into the new space.
- 3:59
And then with DevOps, we're trying to bridge something in the organization. Two things that had friction, that did not work together. And I've witnessed firsthand in the company after, you know, now two years building GenAI application, the friction was in the beginning.
- 4:17
We all wanna do this GenAI thing. It hit first the data science team, and then they were, like, screaming, "Hang on. We're not used to running things in production."
- 4:27
Right? So all of a sudden you got friction. They know their world. We know their o- the other world, and we started slowly moving some engineers into the data science team.
- 4:38
Eventually, the data science team became smaller of that part because it was less about the data science. We heard it this morning. It was more about integration thing, and the ratio of the data science got smaller compared to the engineering.
- 4:51
And then what we did in the organization is we started scaling this out to the different teams, and eventually moved this into the platform. So that kind of friction of the movement in an organization is something I, uh, you know, that, that I see happen a lot.
- 5:07
The data was left on their own in their own data lake, and then now we're actually binding and bonding back together in that way. So in that way, that's the shift right.
- 5:16
Security was shift left, so we're always shifting around, like whatever the complexities we can do.
- 5:25
Maybe first a question. Who is familiar with team topologies as a concept?
- 5:29
Wow. Okay. Not that many. So first you had DevOps as a concept, bringing two teams back in the same room, having them collaborate. That was one way of collaborating.
- 5:43
But that did not scale across the org. So slowly, uh, they became their own team that abstracted pieces away, the Kubernetes pipeline, all the infrastructure that, uh, developers needed to build, the security tooling.
- 5:59
That kind of moved into a platform team. Team Topologies originally was actually called DevOps Topologies because it was a m- like a movement of the teams, uh, reacting to the you build it, you run it.
- 6:13
So we-- eventually, the teams you-- that are building and running it, we call them the, uh, kind of the feature teams, and they run on top of the platform.
- 6:24
And the platform c- team can work together with those teams either by just giving it an API, right? Run as a service, and you're done. We can do more collaboration to see, hey, you know, what are you building?
- 6:39
Let's build this, make sure that the platform is helping you. So it's a little more like they're our customers, we're building a product, you're building on top of that.
- 6:48
And then we can do facilitation if they're having some issues, and we're helping them. So kind of that interaction team. Uh, there's a whole book on this, Team Topologies, and it helps people organize kind of their infrastructure teams, their platform teams, their SRE teams.
- 7:02
So it's a known pattern how we deal with these kind of abstractions of new te-technology, bringing it into the company.
- 7:10
So what I'm explaining here is this whole GenAI thing, the impact is gonna be like not just the data science team doing this. But what I tried to do in the last two years is bring this to the traditional application developers, so they can scale this out into their own teams.
- 7:30
We heard it this morning. What's the ideal team? It's the application developers, the traditional ones, and some of the data science on top of that. But there's a few hurdles to coming to that.
- 7:41
First piece that I'm talking about is the platform.
- 7:45
So what services go into an AI platform? And, you know, you're here at the fair. There's like, I don't know, how many vendors. There's lots of hardware. So I have no affiliation with any of the things that I'm showing.
- 7:57
I'm just gonna show you some pieces that go into that platform that you're providing.
- 8:03
First piece, access to models, right? It's this first thing they can do. Instead of every team going out to figure out what's the best model, where do you get access, that team kind of figures out what is appropriate for our company.
- 8:16
You pick your poison, what you like, what's your vendor, typically related to your cloud vendor. So there's a relationship with the CloudOps team and kind of bringing this in.
- 8:26
Next piece that you heard, like RAG, let's have a vector database. In the beginning, you had like very specific ones. Now you see they're all getting inside of the traditional vendors of kind of the databases and whatever.
- 8:39
So it's not that special anymore, but you need to understand what it is, what is a vector, what is embeddings, how we're helping this. So that's one of the other pieces you bring when you're having the RAG in.
- 8:51
And having that ve-- uh, vector database is not enough, so you bring in connectors to all your data sources. So that's another piece that that team builds in the infrastructure.
- 9:02
They connect whatever is out there in your company, and they expose that to all the other teams instead of hem-- them having to do this all the time one by one on teams and figuring out how that works.
- 9:14
Some have called this RAGOps, you know, whatever ops, right? There's always something to be, uh, managing and running on that.
- 9:21
And then we had version control for code. Obviously, we wanna have version control for models, so that's another piece they can provide. So instead of them figuring out on their own how we do this, there's a centralized repository.
- 9:33
Other teams can reuse, uh, some of the registry-- uh, some of the models. So it becomes like visible, like a library to all the other pieces as well.
- 9:44
And then especially the larger enterprises, uh, it's almost like they wanna be cloud agnostic. They wanna be model provider agnostic. So I'm not saying, like in storage, we used to have S3 as being the standard protocol.
- 9:58
Now maybe that's the OpenAI protocol that being standardized on, like the one winner kind of takes it all. But it's also kind of building this in the access control, who gets access to what.
- 10:10
So there's like a proxy that these kind of environments are building, uh, so not everybody can just go out and use whatever they want. So there's a little bit of that.
- 10:21
Then when you've been running this, you want to have similar to your observability and tracing, you wanna capture whatever prompts is running in production. Uh, so that observability layer, you wanna enhance your existing observability layer, but it's a little bit different.
- 10:38
Uh, you know, one prompt, it's rarely one prompt. Usually one prompt leads to five, six iterations, questions. So that's another piece of this, and you don't want every team to figure this out on their own.
- 10:49
And then when you run it in production, you want to have monitoring on your data quality. And I know like, you know, LMM-- MLOps had this before, but this is kind of a new thing that the traditional monitoring and metrics providers are not very capable of because like you had a health check for your API calls, now
- 11:08
you have a health check for evals running all the time in production. So you will notice if the model has changed. You will notice if the end user is doing strange things.
- 11:17
You notice-- so kind of that observability kind of is built in.
- 11:22
And then lastly, you know, there's, there's many things like caching services, and you go on, like feedback as a service being also discussed this morning. Uh, not just thumbs up, thumbs down, try again, but also kind of in, uh, inline editing of the solution be-because you get better feedback.
- 11:38
And again, you don't want every team to build this service because it's quite expensive, uh, and you want this to be centrally, uh, managed in a good way.
- 11:47
So these are just a few pieces, uh, to bring it on the board. Um, if you ever seen the Kubernetes ecosystem slide, this one is also expanding rapidly, right?
- 11:58
So there's similarities on whatever hardware we're running. It's just gonna keep growing, uh, whatever solution we have out there. So that gives you a little bit of an idea is it's not just clouds and APIs, but this is a new set of infrastructure that you're running.
- 12:15
So next piece is you have all the infrastructure, and you provide it to the teams. And one of the things we learned is that it's not just enough to say, "Here is a bunch of things you could use."
- 12:25
Like any good company, you guide your customers to use them. That's the enablement to kind of whatever you're providing that they're able to use. This is also the place where you get feedback, what is working and not working from your team.
- 12:38
So what does the enablement look like for that team?
- 12:42
You provide prototyping tools for them to easier do experimentation. And it's been mentioned before, don't forget the product owners. They will learn. They wanna experiment. So it's a simple thing you do to kind of get them excited, having them playing around, find the right use case, uh, of their stuff.
- 13:00
And then you also connect it with the data of your company. So you do that in a secure way so they can experiment whatever data is there.
- 13:09
And frameworks are great, and I'll come back to that later, for learning. How does it work? What can we all do? And I learned a ton through these, uh, things, and there is a flavor for everybody, like large data, kind of like you like more like Microsoft stack.
- 13:24
Definitely framework is, uh, one of the things, uh, to give them and hand them over, uh, to learn from. And it also helps in education, documentation, uh, going from there.
- 13:37
They want a local dev environment. They wanna feel safe to do some coding, experimentation when they travel, kind of like that they don't wanna have the resources. They're very, like, focused on having this local stack of development, uh, when working on this.
- 13:52
There's pro and cons, but, you know, to get them excited, it's one of the things that you can definitely do. And we've seen this morning you can run more and more of these quality models on your laptop, so that kind of helps them in their faster iteration.
- 14:05
So things that I saw that actually go a lot bad is what is the actual use case? Uh, I've seen companies shout like, you know, we need to have the GenAI.
- 14:17
Every part of the product needs to have GenAI. But we're in the phase that we're still figuring out the real use case for a lot of things. And we're...
- 14:26
I, I'm not gonna say we're running on the marketing budget, but sometimes it's a little bit like that, right? And another pitfall is if you're very focused on kind of model training and fine-tuning, I can tell you that companies run by the data science, and it might take a while until we actually get something in production as
- 14:45
well. Um, an over-focus on, uh, kind of,
- 14:51
uh, cost, like run local, not use the perfect model. There's a lot of cost that, you know, goes down the drain that way. So we know that the models are getting cheaper.
- 15:01
So focus on the first thing, get that business case focused, and the rest we'll kind of like, we'll, we'll sort out later. If there's actually ROI, we'll reduce the cost on that.
- 15:09
And then end user feedback, um, the engineers often don't know what to do with the feedback because the product owner is the only one that knows the domain. It was mentioned this morning.
- 15:21
So that's kind of also like, yeah, we get all the feedback, but now what? Like, yeah, we're not used to handling this kind of feedback. We're used to handling error stacks as well.
- 15:30
So those are a few of the observations.
- 15:34
And that brings me to there's a lot of emphasis on developer experience in general and improving the productivity. But what does it take to be building GenAI applications? What kind of pains do you have?
- 15:49
Like, I've yet to see the developer that knows what to pick as a model. They just pick one, and it's r- it's okay, but you gotta start somewhere. Access to the data in the test is hard.
- 16:02
Uh, I hear a lot of stories that people use the framework. They were bitten by a framework. They don't like the framework anymore because the frameworks were moving too fast.
- 16:11
Uh, but... And they start doing things themselves at the lower layer or AVIs. I hope that's not the way we're going. I hope we eventually have an ecosystem of tools we're building on and not doing the DIY.
- 16:23
But it's kind of what I've seen in people in the journey. They get excited with the framework, and then it's like, "Okay, we, we got this. It's just a prompt, right?
- 16:30
Like, it's not that difficult." And in essence, they're right, but think about, like, the tracing, the monitoring, the observability, all the ecosystems you wanna kind of bring together.
- 16:40
And we went through, I think, eight different models over two years for LLMs to get, like, improvement of quality. And me looking as a VP engineering, I was like, "Hang on."
- 16:57
So every time we change the model, we have to rewrite and rework all the application prompts to make sure it works. That's very costly. And if it wasn't costly enough, they were doing this manually. [laughs]
- 17:10
So I had no guarantee that was actually gonna work. So we've talked enough about evals and working with evals, but it's definitely one of the pain points of you need to have testing evals before you do this refactoring.
- 17:24
And this brings me to the testing, which is every time that the engineers go like, "What?" Like, "How do I, like, write test for this?" Like it's, you know...
- 17:33
This morning we're like, "Oh, it's a vibe check. Okay, you know, I, I don't get it." It's like, that's not engineering rigor. So
- 17:41
I'm gonna give you, uh, like the buildup that they're saying. Like, "Well, I can do exact testing. That, that's easy, right? Just does it need to be twenty characters or regex?
- 17:49
I know that stuff. Sentiment check? Okay." Well, you know, all of a sudden you have to run a model, a helper model, uh, to bring that in. We can see a simple thing like similar to the semantic, uh, distance.
- 18:04
Is the question related to answer? That's an easy one. But then you're telling me, how do you check an LLM? Use another LLM?
- 18:13
Right? So if the AI doesn't work, just use more AI. Like, it blows my mind. Like, it's like a vicious circle, and it hasn't been really solved. But hey, it's the best thing we got, and I got it.
- 18:23
Like, and, and eventually there's a human feedback. So it's, it's a compensation of a, a bit of everything there.
- 18:30
Um, so now I'm switching gears a bit. Like-
- 18:35
When you wanna bring this to the engineers, a lot of the engineers are afraid of the AI in the beginning. So how do you get them excited? And the usu- usual trick is we get them something on their coding pilot, because then they see kind of the productivity that this brings.
- 18:52
And so it goes hand in hand. Like, it's a little bit weird that the same team that is building the GenAI models is also being the evangelizing of kinda, like, take away the fear of developers using more AI.
- 19:05
And this one, uh, I know there's a lot of talk about, like, getting them more productive, but what I wanna show here is, yes, the copilot introduced a lot more code.
- 19:18
It was a lot more things to do. But review times went up, and the PRs sent by the actual AI were bigger, right? So we got faster, but then we started slowing down again.
- 19:30
So there's like... I'm not saying it's, like, a hundred percent compensating the other part, but there's definitely... It, it's like there's friction on the line.
- 19:40
And this brings me to a paper that where-- was very instrumental in the beginning of, uh, DevOps. It was called The Ironies of Automation. And now there's a new paper, The Ironies of GenAI Automation.
- 19:52
And so you see, um, the role of the person using any of the AI is changing from the person producing to somebody who's managing and reviewing things, right? So...
- 20:05
And you could say, "Well, you know, that, that's good." Like, you know. But you see you're spending more time in the review and kind of faster of the generation.
- 20:13
But it can only also lead to, "I don't understand anymore what this, this is doing because I don't have the domain model. I don't have the experience anymore." And then I'm just gonna say accept.
- 20:26
There's a funny story about GitHub Copilot being used, uh, and having a higher acceptance rate for suggestions in the weekend because [laughs] the developers are like, "Whatever." [laughs]
- 20:38
So... But we, we've seen this before, like, the automation, that acceptance, so we kind of have to go a-against this. And you would gradually, when you're not doing the producing job anymore, you'll lead, uh, like, lose situational awareness of when things, uh, go, uh, failing.
- 20:56
And this brings me to this slide. So
- 21:00
DevOps automating things. Go away. I'll re- I'll replace you with a shell script. That was kind of, you know, the left part, twenty percent automate. The other percent that grew into the industry was preparing for failure.
- 21:15
First we had CI/CD. We created more tests, evals. Then we had monitoring because we needed to know what was going on in production. And then all of a sudden it's like, "Yeah, but if it goes down, like, can we prevent the failure?
- 21:27
We're gonna assume failure." [laughs] That was the big model shift. So we started design for failure, and then we needed to predict. But we didn't know what was going on.
- 21:37
But we need situational awareness. That was the whole observability. And then in the end we brought in chaos engineering to kind of inject failures so we can keep training when failures happened, right?
- 21:48
So you see kinda this paradigm of, yeah, we won a lot, but kind of that spun up a whole new thing of doing. I'm not saying everybody's doing this.
- 21:58
You can gladly skip all the tests if that's your risk appetite, but that's c- kind of the shift that, uh, it was going to. And it's also reducing cognitive load in when failure happening, uh, and kind of work from there.
- 22:12
So that's, you know, a little piece on enablement. It's not just, you know, kind of making them GenAI. It's also putting them at ease. And then there's the governance piece.
- 22:21
Uh, make a personal awareness program, right? Don't just copy things in. Uh, this is one that records your screen and looks everything on your screen, right? That, that's... It's getting scary in that way of leaking things.
- 22:38
Uh, but it also promises productivity, so it's like a-- it's a tension that you have to overcome in your company. Uh, you wanna have them opt out on training.
- 22:47
It's-- We've talked about this this morning. It's not always that clear where you opt out, when you opt out, uh, how to work from them. So you have to make, make them aware whenever they're putting a new service in there.
- 22:58
You wanna have them check the license, but that's overcome by just restricting the number of models you put in. Uh, it's not just all because it has the word open that it's open, right?
- 23:07
So we all learned that. Um, look at what the origin was for the model. Learn from that as well. So those are things that typically goes into a governance, uh, uh, workflow, um, of approval.
- 23:22
And then there's the [REDACTED:origin]... I come from Belgium, so risk levels assessment. What kind of workflow are you doing? Is it allowed? Should we do this? Kind of make this awareness a little bit of the law.
- 23:31
I'm not making them specialist lawyers, but at least, you know, they kind of have to know the basics of that as well.
- 23:38
And then there's prompt injection, but we learned it's not a solved problem. But you gotta have something, right, at least in place, uh, for the failures. Um, and then guardrails, uh, but I don't know if you ever done, like, intrusion detection alerting or web application firewalls.
- 23:54
You put all the rules and after a while you say, like, the logs are so big, you're like, "Whatever." So it's, it's not always a problem, but we, we kind of tell ourselves that there's guardrails in place, uh, uh, in there.
- 24:06
So let's hope we are, like, on the path of improving those kind of tools eventually, like, for preventing failure and optimizing for failure. And then we wanna, like, focus on the PII that goes over the wire, the metrics, the alerting, because we're now sending a lot more text over the wire, k- uh, kind of uncontrolled things, so
- 24:26
it needs a lot more scrutiny on there. So simple thing for me. These are the steps you bring in. You bring your infrastructure on one team. Uh, you kind of do the enablement on top of that.
- 24:38
And they also set the, the rules for your company to use this.
- 24:42
And I briefly talked about the platform team, which provides the internal focus and the internal enablement to the company. But, um, there's another model called the unfixed model, and it's a little bit like alternative to the team topologies.
- 25:00
And they also talk about the experience crew. And what the experience crew is, they make AI look consistent in the product. So they're a team that goes on to the feature teams.
- 25:12
It's like, "So what are you doing? How does it look like in the product? What's the UX?" And kind of they do this. They have a whole lot of number of crews, but it...
- 25:21
This is also using companies and they call it, you know, the experience, the AI experience team that works kind of hand-in-hand. And so that leads me to CloudOps, SecOps, developer experience, put the data platform in there, put the AI platform infrastructure in there, and on top of that, do the AI, uh, kind of experiences in there.
- 25:39
Um, I would not recommend to do this for a 10-person company. Uh, don't do premature optimization. But if you're scaling out to 10 and more teams, that's probably a pattern that is known.
- 25:52
And the nice thing is you have cross-collaboration. Uh, the SecOps are really good at access control and governance, so they help each other. The CloudOps know what to spin up and know the cloud vendors.
- 26:04
And so there's a lot of kind of working together. And if you bring that into one area, you have a better chance of pulling this in, uh, and being like a collaboration across.
- 26:15
So that was me. And, you know, you can scan, scan the link and connect and, um, yeah, I don't know. What do you think? Wanna have a platform team for this, or are some people doing this already?
- 26:27
Uh, let me know. So thank you. [audience applauding] Questions? Yeah, we have... Literally, we do have a time for a question or two if anyone does want to ask anything. Uh, oh, fantastic.
- 26:42
I have a question on the slide.
- 26:44
On the slide.
- 26:44
Will they be available?
- 26:45
Oh. Sorry, you... Oh, you, you wanna get a copy of the slides?
- 26:49
Yes.
- 26:49
Okay. He wants a copy of the slides. Yeah. Yeah. Sure. Okay, cool.
- 26:55
Oh, sorry.
- 26:56
I have a question.
- 26:56
Yeah, go, go for it.
- 26:57
Okay. So you, you had mentioned guardrails as a service, and this is one of the interesting paradigms that we're looking into as well and trying to, uh, sort of figure out how to best serve guardrails as a service.
- 27:08
Did you see or is, is there a certain leaning towards you put guardrails in front of all models that you serve, and then, like in your model card, you say, "This is the model.
- 27:19
These are all the guardrails in front"? Or do you expect product teams to consume the guardrails but leave it in their remit to go and consume them properly?
- 27:31
Yeah. So, so what we've seen is that the central governance teams will put the generic rules in place. And then depending on the use case and whatever they're putting, the, each of the teams put their own rules on top of that.
- 27:43
So there's a kind of self-servicing, but also like a centralized component of rules. Not every team has to duplicate, uh, to go from there. Cool. Thank you very much.
- 27:55
All right. Now shall we get set up? [upbeat music]