← All AI Engineer talks

AI Engineer World's Fair 2025

The 2025 AI Engineering Report — Barr Yaron, Amplify Partners

About this talk

Amplify Partners investment partner Barr Yaron presents findings from the 2025 State of AI Engineering Survey, covering practitioners' varied job titles and AI experience, widespread internal and customer-facing LLM deployments, and OpenAI model adoption. She reports that 70% of respondents use RAG, discusses LoRA, QLoRA, DPO, and supervised fine-tuning, and notes that 70% update prompts at least monthly while 31% lack structured prompt management. The talk also addresses the multimodal production gap, AI agents, and model-usage monitoring.

Chapters

  1. 0:28Introducing the State of AI Engineering Survey
  2. 1:42Practitioner titles, experience, and the rise of AI engineering
  3. 3:17LLM adoption, RAG, and fine-tuning techniques
  4. 5:40Prompt-management gaps and multimodal production
  5. 7:48AI agents and model-usage monitoring
  6. 10:02Additional survey findings and community resources

Talk transcript

  1. 0:00

    [upbeat music] [audience applauding]

  2. 0:28

    All right. Hi, everyone. Uh, thank you for having me here, and huge thanks to Ben, to Swyx, to all the organizers who've put so much time and heart into bringing this community together. [audience applauding]

  3. 0:41

    Yeah. All right. So we're here because we care about AI engineering and where this field is headed. So to better understand the current landscape, we launched the twenty-twenty five State of AI Engineering Survey, and I'm excited to share some early findings with you today.

  4. 1:03

    All right. Before we dive into the results, the least interesting slide. Uh, I don't know everyone in this audience, but I'm Barr. I'm an investment partner at Amplify, where I'm lucky to invest in technical founders, including companies built by and for AI engineers.

  5. 1:19

    And, uh, with that, let's get into what you actually care about, which is enough Barr and more bar charts, and there are a lot of bar charts coming up.

  6. 1:30

    Okay, so first, our sample. We had five hundred respondents fill out the survey, including many of you here in the audience today and on the live stream. Thank you for doing that.

  7. 1:42

    And the largest group called themselves engineers, whether software engineers or AI engineers. While this is the AI engineering conference, it's clear from the speakers, from the hallway chats, there's a wide mix of titles and roles.

  8. 1:57

    You even let a VC sneak in. Um, so let's test this with a quick show of hands. Raise your hand if your title is actually AI engineer at the AI engineering conference.

  9. 2:08

    Okay, that is extremely sparse. [laughs] Uh, raise your hand-- Put your hands down. Raise your hand if your title is something else entirely. So that should be almost everyone. Keep it up if you think you're doing the exact same work as many of the AI engineers.

  10. 2:27

    All right, so this sort of tracks. Titles are weird right now, but the community is broad. It's technical. It's growing. We expect that AI engineer label to gain even more ground.

  11. 2:37

    Uh, couldn't help myself. Quick Google Trends search. Term AI engineering barely registered before late twenty-twenty two. Uh, we know what happened. ChatGPT launched, and the moment for AI engineering interest has not slowed since.

  12. 2:51

    Okay, so people had a wide variety of titles, but also a wide variety of experience. Uh, the interesting part here is that many of our most seasoned developers are AI newcomers.

  13. 3:01

    So among software engineers with ten plus years of software experience, nearly half have been working with AI for three years or less, and one in ten started just this past year.

  14. 3:12

    So change right now is the only constant, even for the veterans.

  15. 3:17

    All right, so what are folks actually building? Let's get into the juice. So more than half of the respondents are using LLMs for both internal and external use cases.

  16. 3:27

    Uh, what was striking to me was that three out of the top five models and half of the top ten models that respondents are using for those external cases for the customer-facing products are from OpenAI.

  17. 3:41

    The top use cases that we saw are code generation and code intelligence and writing assistant content generation. Maybe that's not particularly surprising. Uh, but the real story here is heterogeneity.

  18. 3:51

    So ninety-four percent of people who use LLMs are using it for at least two use cases, eighty-two percent using it for at least three. Basically, folks who are using LLMs are using it internally, externally, and across multiple use cases.

  19. 4:06

    All right. So you may ask, how are folks actually interfacing with the models, and how are they customizing their systems to-- for these use cases? Uh, besides few-shot learning, RAG is the most popular way folks are customizing their systems, so seventy percent of respondents said they're using it.

  20. 4:25

    The real surprise for me here, I, uh, I'm, I'm looking to gauge surprise in the audience, was how much fine-tune is hap-- fine-tuning is happening across the board. It was much more than I had expected overall.

  21. 4:37

    Uh, in the sample, we have researchers and we have research engineers who are the ones doing fine-tuning by far the most. We also asked an open-ended question for those who were fine-tuning, what specific techniques are you using?

  22. 4:50

    So here's what the fine-tuners had to say. Uh, forty percent mentioned LoRA or QLoRA, reflecting a strong preference for parameter-efficient methods, and we also saw a bunch of different fine-tuning methods, uh, including DPO, reinforcement fine-tuning, and the most popular core training approach was good old supervised fine-tuning.

  23. 5:12

    Many hybrid approaches were listed as well. Um, moving on top, uh, to up-- on top of updating systems, sometimes it can feel like new models come out every single week.

  24. 5:26

    Just as you finished integrating one, another one drops with better benchmarks and a breaking change. So it turns out more than fifty percent are updating their models at least monthly, seventeen percent weekly.

  25. 5:40

    And folks are updating their prompts much more frequently. So seventy percent of respondents are updating prompts at least monthly, and one in ten are doing it daily. So it sounds like some of you have not stopped typing since GPT-4 dropped.

  26. 5:54

    Um, but I also understand. I have empathy. Uh, seeing one blog post from Simon Willison, and suddenly your trusty prompt just isn't good enough anymore.

  27. 6:07

    Despite all of these prompt changes, a full thirty-one percent of respondents don't have any way of managing their prompts. [laughs] Uh, what I did not ask is how AI engineers feel about not doing anything to manage their prompts.

  28. 6:21

    So we have the twenty twenty-six survey for that.

  29. 6:25

    We also asked folks across the different modalities who is actually using these models at work, and is it actually going well? And we see that image, video, and audio usage all lag text usage by significant margins.

  30. 6:41

    I like to call this the multimodal production gap

  31. 6:45

    'cause I wanted an animation. [laughs] Um, and this gap still pers- persists when we add in folks who have these models in production but have not garnered as much traction.

  32. 7:00

    Okay, what's interesting here is when we add the folks who are not using models at all in this chart, too. So here we can see folks who are not using text, not using image, not using audio, or not using video.

  33. 7:15

    And we have two categories. It's broken down by folks who plan to eventually use these modalities and folks who do not currently plan to.

  34. 7:24

    You can roughly see this ratio of no plan to adopt versus plan to adopt. Audio has the highest intent to adopt, so thirty-seven percent of the folks not using audio today have a plan to eventually adopt audio.

  35. 7:39

    So get ready to see an audio wave. Um, of course, as models get better and more accessible, I imagine some of these adoption numbers will go up even further.

  36. 7:48

    All right, so we have to talk about agents. One question I almost put in the survey was, how do you define an AI agent? But I thought I would still be reading through different responses.

  37. 7:59

    Uh, so for the sake of clarity, we defined an AI agent as a system where an LLM controls the core decision-making or workflow.

  38. 8:08

    Eighty percent of respondents say LLMs are working well at work, but less than twenty percent say the same about agents.

  39. 8:16

    Agents aren't everywhere yet, but they're coming. Uh, the majority of folks, uh, may not be using agents, but most at least plan to. So fewer than one in ten say that they will never use agents.

  40. 8:28

    All to say that people want their agents, and I'm probably, uh, preaching to the choir.

  41. 8:34

    Um, the majority of agents already in production do have write access, uh, typically with a human in the loop, and some can even take actions independently. So, um, excited as more agents are adopted to learn more about the tool permissioning that folks, uh, have access to.

  42. 8:52

    If we want AI in production, of course, we need strong monitoring and observability. So we asked, "Do you manage and monitor your AI systems?" This was a multi-select question, so most folks are using multiple methods to monitor their systems.

  43. 9:06

    Sixty percent are using standard observability. Over fifty percent rely on offline eval. And we asked the same thing for how you evaluate your model and system accuracy and quality.

  44. 9:18

    So folks are using a combination of methods, including data collection from users, benchmarks, et cetera, but the most popular at the, at the end of the day is still human review.

  45. 9:28

    Um, and for monitoring their own model usage, most respondents rely on internal metrics.

  46. 9:35

    So storage is important, too. Where does the context live? How do we get it when we need it? Sixty-five percent of respondents are using a dedicated vector database, and this suggests that for many use cases, specialized vector databases are providing enough value over general purpose databases with vector extensions.

  47. 9:53

    Uh, among that group, thirty-five set-- percent said that they primarily self-host. Thirty percent primarily use a third-party provider.

  48. 10:02

    All right. I think we've been having fun this whole time, but we're entering a section I like to formally call Other Fun Stuff. Uh, I spent hours workshopping the name.

  49. 10:12

    So we asked AI engineers, "Should agents be required to disclose when they're AI and not human?" Most folks think, yes, agents should disclose that they're AI. Uh, we asked folks if they'd pay more for inference time compute, and the answer was yes, but not by a wide margin.

  50. 10:28

    And we asked folks if transformer-based models will be dominant in twenty-thirty, and it seems like people do believe that attention is all we'll need in twenty-thirty.

  51. 10:38

    Uh, the majority of respondents also think open source and closed source models are gonna converge, so I will let you debate that after. Um, no commentary needed here. So, uh, the average or the mean guess for the percentage of US Gen Z population that will have AI girlfriends, boyfriends is twenty-six percent.

  52. 10:57

    Um, I don't really know what to say or expect here, but we'll see. Uh, we'll see what happens, uh, [laughs] in a world where folks don't know if they're being left on read or just facing latency issues, um, or, uh, of course, the dreaded, "It's not you, it's my algorithm."

  53. 11:16

    And finally, we asked folks, "What is the number one most painful thing about AI engineering today?" And evaluation topped that list. Uh, so it's a good thing this conference and the talk before me has been so focused on evals 'cause clearly they're causing some serious pain.

  54. 11:31

    Okay, and now to bring us home, I'm gonna show you what's popular. So we asked folks to pick all the podcasts and newsletters that they actively learn something from at least once a month, and these were the top ten of each.

  55. 11:43

    So if you're looking for new content to follow and to learn from, this is your guide. Uh, many of the creators are in this room, so keep up the great work.

  56. 11:52

    And I'll just shout out that Swyx is listed both on popular newsletter and popular podcast for latent space. Uh, so I will just leave this here. [laughs] [audience laughing]

  57. 12:05

    Um, I think that's enough bar charts and bar time, but if you wanna geek out about AI trends, you can come find me online, in the hallways. Uh, we're gonna be publishing a full report next week.

  58. 12:14

    Uh, I'll let Elon and Musk have Twitter today. But, um, it's gonna include more juicy details, including everyone's favorite models and tools across the stack. Thank you for the time.

  59. 12:25

    Enjoy the afternoon. [upbeat music]