AI Engineer World's Fair 2025
The 2025 AI Engineering Report — Barr Yaron, Amplify Partners
Read the talk
The 2025 AI Engineering Report: Broad Adoption, Uneven Production
A survey of 500 respondents shows how AI engineering spans job titles, use cases, and customization methods—and where evaluation and production practices still lag.
From a talk by Barr Yaron
Who counts as an AI engineer?
Does AI engineering describe a job title or the work people actually do? The early findings of the 2025 State of AI Engineering Survey give that question a concrete setting. Barr Yaron, an investment partner at Amplify backing technical founders, introduces the results with a promise of less biography and more charts. The survey has 500 respondents, including conference attendees and livestream viewers. Software engineers and AI engineers make up the largest group.
Then comes a quick test in the room. Yaron asks who actually holds the title AI engineer; few hands rise. She asks people with other titles to raise their hands, then keep them up if they do the same work as AI engineers. The response makes the distinction visible: the work extends beyond the title. Yaron expects the label to gain ground as this broad technical community grows.
Her Google Trends observation places the emergence of the term around the launch of ChatGPT: AI engineering barely registered before late 2022, then attracted sustained interest. A new discipline does not necessarily mean inexperienced practitioners. Among surveyed software engineers with at least ten years of software experience, nearly half have worked with AI for three years or less, and one in ten started in the past year. Even veteran developers are learning a new practice.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
LLMs spread across multiple applications
More than half of respondents use LLMs for both internal and external applications. For customer-facing products, OpenAI supplies three of the five most-used models and five of the top ten. These rankings capture a historical snapshot preceding Claude 4 and GPT-4.1.
Code generation and code intelligence lead the applications, alongside writing assistance and content generation. But the more consequential pattern is breadth: among respondents using LLMs, 94% use them for at least two use cases and 82% for at least three. Adoption here means using models across several kinds of work, often both inside the company and in products customers use.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
RAG and fine-tuning both have a place
Beyond few-shot learning, retrieval-augmented generation, or RAG, is the most popular system-customization approach: 70% of respondents report using it. Fine-tuning is also more common than Yaron expected, with researchers and research engineers doing it most often.
An open-ended question asks fine-tuners which techniques they use. Forty percent of those responses mention LoRA or QLoRA, which Yaron reads as a strong preference for parameter-efficient methods. The responses also name Direct Preference Optimization, or DPO, and reinforcement fine-tuning. Supervised fine-tuning remains the most popular core training approach, and respondents describe hybrid approaches as well. The resulting picture is a mix of customization methods rather than a single standard recipe.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Prompts change faster than models
Just as one model integration settles, another release arrives with better benchmarks and a breaking change. Respondents report frequent changes to both models and prompts:
| What changes | At least monthly | Faster cadence |
|---|---|---|
| Models | More than 50% | 17% weekly |
| Prompts | 70% | One in ten daily |
Prompt revision is especially frequent; Yaron jokes that some engineers have not stopped typing since GPT-4 arrived.
A post from Simon Willison can be enough to make yesterday’s trusted prompt feel inadequate. Yet 31% of respondents report no structured prompt-management tool. Frequent revision and explicit management have not advanced together. The survey does not ask how engineers feel about that gap; Yaron leaves that question for a jokingly proposed 2026 follow-up.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The multimodal production gap
Text is well ahead of image, video, and audio in workplace usage. Yaron calls this the multimodal production gap. The gap persists when the chart expands beyond deployments working well to include models already in production but with less traction. Broadening the definition of production adoption does not erase text’s lead.
The next reveal separates nonusers into two groups: those planning eventual adoption and those with no current plan. Audio has the highest adoption intent: 37% of respondents not currently using audio plan to adopt it eventually. That percentage describes intentions among audio nonusers, not current usage across the whole sample. Yaron anticipates an audio wave and expects better, more accessible models to encourage further adoption.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Agents face a larger production hurdle
Before asking about agents, the survey fixes a definition: an AI agent is a system in which an LLM controls the core decision-making or workflow. That avoids collecting adoption figures against hundreds of different interpretations. Under this definition, 80% of respondents say LLMs work well at work, while fewer than 20% say the same about agents. This is respondent-reported workplace success, not benchmark accuracy or a success rate restricted to agent users.
Most respondents are not yet using agents, but most at least plan to; fewer than one in ten say they will never use them. The gap between current success and future interest is substantial.
For agents already in production, permissions are more consequential than the adoption label alone. Most have write access, typically with a human in the loop, and some can act independently. An agent allowed to change a system presents a different operational responsibility from one that only reads it. Yaron identifies tool permissioning as an area she wants to understand better as deployments grow.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Monitoring quality and storing context
Production systems need monitoring, but respondents do not rely on one method. The monitoring question allows multiple selections, and most select multiple approaches. Standard observability is used by 60% of respondents, while more than 50% rely on offline evaluation.
Accuracy and quality evaluation likewise combine several methods, including collecting user data and using benchmarks. Human review remains the most popular. For tracking their own model usage, most respondents rely on internal metrics. These are distinct concerns: observing system behavior, judging output quality, and understanding model usage do not collapse into a single measurement.
Context storage adds another infrastructure choice. Dedicated vector databases are used by 65% of respondents. Yaron interprets that adoption as evidence that specialized databases provide enough value over general-purpose databases with vector extensions for many use cases. She then reports 35% primarily self-hosting and 30% primarily using a third-party provider, introducing those figures as being among the dedicated-vector-database group. The denominator is not explicit enough to convert those hosting figures into respondent counts.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
What respondents expect next
The survey then turns from deployed systems to opinions and forecasts:
- AI disclosure: Most respondents think agents should be required to disclose that they are AI rather than human.
- Inference-time compute: Willingness to pay more wins, but only by a narrow margin.
- Model architecture: Respondents generally expect transformer-based models to remain dominant in 2030, prompting Yaron’s callback that attention may be all we will need.
- Open and closed models: A majority expect open-source and closed-source models to converge.
On a more speculative question, the mean respondent estimate for the share of the US Gen Z population that will have AI girlfriends or boyfriends is 26%. This is a forecast, not measured prevalence; the spoken question gives no forecast year.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Evaluation hurts; learning remains a shared practice
Asked for the single most painful part of AI engineering, respondents put evaluation first. Yaron connects that result to the conference’s emphasis on evals and the preceding talk: the attention is responding to a problem practitioners already feel.
The closing survey question asks which podcasts and newsletters respondents actively learn from at least once a month, allowing multiple selections. The slide presents the top ten in each category. Latent Space heads both lists, and Yaron singles out Swyx for appearing in both. Latent Space was also a survey partner, which may have influenced who responded and the resulting learning-source rankings. The lists offer a view of this community’s learning habits, rather than a population-wide ranking of publications.
At the time of the presentation, Yaron says the full report will appear the following week, adding detail on respondents’ favorite models and tools across the stack. She closes by inviting further discussion of AI trends online and in the conference hallways. The early findings leave evaluation as the leading practical pain point, alongside a community still exchanging the knowledge needed to address it.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
From the talk
The original paper explaining fine-tuning with small trainable matrices while keeping base-model weights frozen.
Combines low-rank adapters with a quantized base model to reduce fine-tuning memory requirements.
Introduces a direct training objective for aligning language models with preference data.
Technical writing and experiments about language models, programming, and AI tools.
An AI engineering publication featured in the talk’s learning-source discussion.
Further reading
Foundational research on combining retrieved knowledge with language generation.
Read the complete timestamped transcript
- 0:00
[upbeat music] [audience applauding]
- 0:28
All right. Hi, everyone. Uh, thank you for having me here, and huge thanks to Ben, to Swyx, to all the organizers who've put so much time and heart into bringing this community together. [audience applauding]
- 0:41
Yeah. All right. So we're here because we care about AI engineering and where this field is headed. So to better understand the current landscape, we launched the twenty-twenty five State of AI Engineering Survey, and I'm excited to share some early findings with you today.
- 1:03
All right. Before we dive into the results, the least interesting slide. Uh, I don't know everyone in this audience, but I'm Barr. I'm an investment partner at Amplify, where I'm lucky to invest in technical founders, including companies built by and for AI engineers.
- 1:19
And, uh, with that, let's get into what you actually care about, which is enough Barr and more bar charts, and there are a lot of bar charts coming up.
- 1:30
Okay, so first, our sample. We had five hundred respondents fill out the survey, including many of you here in the audience today and on the live stream. Thank you for doing that.
- 1:42
And the largest group called themselves engineers, whether software engineers or AI engineers. While this is the AI engineering conference, it's clear from the speakers, from the hallway chats, there's a wide mix of titles and roles.
- 1:57
You even let a VC sneak in. Um, so let's test this with a quick show of hands. Raise your hand if your title is actually AI engineer at the AI engineering conference.
- 2:08
Okay, that is extremely sparse. [laughs] Uh, raise your hand-- Put your hands down. Raise your hand if your title is something else entirely. So that should be almost everyone. Keep it up if you think you're doing the exact same work as many of the AI engineers.
- 2:27
All right, so this sort of tracks. Titles are weird right now, but the community is broad. It's technical. It's growing. We expect that AI engineer label to gain even more ground.
- 2:37
Uh, couldn't help myself. Quick Google Trends search. Term AI engineering barely registered before late twenty-twenty two. Uh, we know what happened. ChatGPT launched, and the moment for AI engineering interest has not slowed since.
- 2:51
Okay, so people had a wide variety of titles, but also a wide variety of experience. Uh, the interesting part here is that many of our most seasoned developers are AI newcomers.
- 3:01
So among software engineers with ten plus years of software experience, nearly half have been working with AI for three years or less, and one in ten started just this past year.
- 3:12
So change right now is the only constant, even for the veterans.
- 3:17
All right, so what are folks actually building? Let's get into the juice. So more than half of the respondents are using LLMs for both internal and external use cases.
- 3:27
Uh, what was striking to me was that three out of the top five models and half of the top ten models that respondents are using for those external cases for the customer-facing products are from OpenAI.
- 3:41
The top use cases that we saw are code generation and code intelligence and writing assistant content generation. Maybe that's not particularly surprising. Uh, but the real story here is heterogeneity.
- 3:51
So ninety-four percent of people who use LLMs are using it for at least two use cases, eighty-two percent using it for at least three. Basically, folks who are using LLMs are using it internally, externally, and across multiple use cases.
- 4:06
All right. So you may ask, how are folks actually interfacing with the models, and how are they customizing their systems to-- for these use cases? Uh, besides few-shot learning, RAG is the most popular way folks are customizing their systems, so seventy percent of respondents said they're using it.
- 4:25
The real surprise for me here, I, uh, I'm, I'm looking to gauge surprise in the audience, was how much fine-tune is hap-- fine-tuning is happening across the board. It was much more than I had expected overall.
- 4:37
Uh, in the sample, we have researchers and we have research engineers who are the ones doing fine-tuning by far the most. We also asked an open-ended question for those who were fine-tuning, what specific techniques are you using?
- 4:50
So here's what the fine-tuners had to say. Uh, forty percent mentioned LoRA or QLoRA, reflecting a strong preference for parameter-efficient methods, and we also saw a bunch of different fine-tuning methods, uh, including DPO, reinforcement fine-tuning, and the most popular core training approach was good old supervised fine-tuning.
- 5:12
Many hybrid approaches were listed as well. Um, moving on top, uh, to up-- on top of updating systems, sometimes it can feel like new models come out every single week.
- 5:26
Just as you finished integrating one, another one drops with better benchmarks and a breaking change. So it turns out more than fifty percent are updating their models at least monthly, seventeen percent weekly.
- 5:40
And folks are updating their prompts much more frequently. So seventy percent of respondents are updating prompts at least monthly, and one in ten are doing it daily. So it sounds like some of you have not stopped typing since GPT-4 dropped.
- 5:54
Um, but I also understand. I have empathy. Uh, seeing one blog post from Simon Willison, and suddenly your trusty prompt just isn't good enough anymore.
- 6:07
Despite all of these prompt changes, a full thirty-one percent of respondents don't have any way of managing their prompts. [laughs] Uh, what I did not ask is how AI engineers feel about not doing anything to manage their prompts.
- 6:21
So we have the twenty twenty-six survey for that.
- 6:25
We also asked folks across the different modalities who is actually using these models at work, and is it actually going well? And we see that image, video, and audio usage all lag text usage by significant margins.
- 6:41
I like to call this the multimodal production gap
- 6:45
'cause I wanted an animation. [laughs] Um, and this gap still pers- persists when we add in folks who have these models in production but have not garnered as much traction.
- 7:00
Okay, what's interesting here is when we add the folks who are not using models at all in this chart, too. So here we can see folks who are not using text, not using image, not using audio, or not using video.
- 7:15
And we have two categories. It's broken down by folks who plan to eventually use these modalities and folks who do not currently plan to.
- 7:24
You can roughly see this ratio of no plan to adopt versus plan to adopt. Audio has the highest intent to adopt, so thirty-seven percent of the folks not using audio today have a plan to eventually adopt audio.
- 7:39
So get ready to see an audio wave. Um, of course, as models get better and more accessible, I imagine some of these adoption numbers will go up even further.
- 7:48
All right, so we have to talk about agents. One question I almost put in the survey was, how do you define an AI agent? But I thought I would still be reading through different responses.
- 7:59
Uh, so for the sake of clarity, we defined an AI agent as a system where an LLM controls the core decision-making or workflow.
- 8:08
Eighty percent of respondents say LLMs are working well at work, but less than twenty percent say the same about agents.
- 8:16
Agents aren't everywhere yet, but they're coming. Uh, the majority of folks, uh, may not be using agents, but most at least plan to. So fewer than one in ten say that they will never use agents.
- 8:28
All to say that people want their agents, and I'm probably, uh, preaching to the choir.
- 8:34
Um, the majority of agents already in production do have write access, uh, typically with a human in the loop, and some can even take actions independently. So, um, excited as more agents are adopted to learn more about the tool permissioning that folks, uh, have access to.
- 8:52
If we want AI in production, of course, we need strong monitoring and observability. So we asked, "Do you manage and monitor your AI systems?" This was a multi-select question, so most folks are using multiple methods to monitor their systems.
- 9:06
Sixty percent are using standard observability. Over fifty percent rely on offline eval. And we asked the same thing for how you evaluate your model and system accuracy and quality.
- 9:18
So folks are using a combination of methods, including data collection from users, benchmarks, et cetera, but the most popular at the, at the end of the day is still human review.
- 9:28
Um, and for monitoring their own model usage, most respondents rely on internal metrics.
- 9:35
So storage is important, too. Where does the context live? How do we get it when we need it? Sixty-five percent of respondents are using a dedicated vector database, and this suggests that for many use cases, specialized vector databases are providing enough value over general purpose databases with vector extensions.
- 9:53
Uh, among that group, thirty-five set-- percent said that they primarily self-host. Thirty percent primarily use a third-party provider.
- 10:02
All right. I think we've been having fun this whole time, but we're entering a section I like to formally call Other Fun Stuff. Uh, I spent hours workshopping the name.
- 10:12
So we asked AI engineers, "Should agents be required to disclose when they're AI and not human?" Most folks think, yes, agents should disclose that they're AI. Uh, we asked folks if they'd pay more for inference time compute, and the answer was yes, but not by a wide margin.
- 10:28
And we asked folks if transformer-based models will be dominant in twenty-thirty, and it seems like people do believe that attention is all we'll need in twenty-thirty.
- 10:38
Uh, the majority of respondents also think open source and closed source models are gonna converge, so I will let you debate that after. Um, no commentary needed here. So, uh, the average or the mean guess for the percentage of US Gen Z population that will have AI girlfriends, boyfriends is twenty-six percent.
- 10:57
Um, I don't really know what to say or expect here, but we'll see. Uh, we'll see what happens, uh, [laughs] in a world where folks don't know if they're being left on read or just facing latency issues, um, or, uh, of course, the dreaded, "It's not you, it's my algorithm."
- 11:16
And finally, we asked folks, "What is the number one most painful thing about AI engineering today?" And evaluation topped that list. Uh, so it's a good thing this conference and the talk before me has been so focused on evals 'cause clearly they're causing some serious pain.
- 11:31
Okay, and now to bring us home, I'm gonna show you what's popular. So we asked folks to pick all the podcasts and newsletters that they actively learn something from at least once a month, and these were the top ten of each.
- 11:43
So if you're looking for new content to follow and to learn from, this is your guide. Uh, many of the creators are in this room, so keep up the great work.
- 11:52
And I'll just shout out that Swyx is listed both on popular newsletter and popular podcast for latent space. Uh, so I will just leave this here. [laughs] [audience laughing]
- 12:05
Um, I think that's enough bar charts and bar time, but if you wanna geek out about AI trends, you can come find me online, in the hallways. Uh, we're gonna be publishing a full report next week.
- 12:14
Uh, I'll let Elon and Musk have Twitter today. But, um, it's gonna include more juicy details, including everyone's favorite models and tools across the stack. Thank you for the time.
- 12:25
Enjoy the afternoon. [upbeat music]