AI Engineer World's Fair 2026

The Self-Improving OSS Agent Stack — Marc Klingen, Langfuse

Read the talk

The Self-Improving OSS Agent Stack

Marc Klingen explains how production traces, datasets and evaluations can become an agent-driven improvement loop—and shows a changelog writer learning to stop leaking internal jargon while humans keep control of what gets published.

From a talk by Marc Klingen

At a glance

Ideas worth remembering

  • Connect production traces to representative offline datasets so proposed improvements can be tested against problems users actually encounter.

  • Human review matters most when choosing relevant failures, changing evaluation criteria and checking that proposed fixes serve the intended task.

  • Approvals, edits and change requests can supply improvement signals during ordinary product use, reducing separate labeling work.

  • The changelog demo makes a missing quality requirement testable, changes the writer’s context and compares V1 with V2 while keeping publication under human review.

  • Agents reading historical traces turn observability into a read-heavy workload and make affordable, long-term data access part of the improvement stack.

Better models make a larger improvement loop possible

Early GitHub issue-to-pull-request automation had a surprisingly small working area: a single HTML file. When Langfuse’s founders tried building it around the release of ChatGPT, multi-file edits were too complicated. Improving an agent required people to do much of the surrounding work themselves. Marc Klingen, Langfuse co-founder, opens this talk with the capability change that makes a different workflow plausible: models can now take on more of the work involved in improving other agents.

The shift is toward loops: repeatedly observe behavior, propose a change and check whether it helps. As models handle more of those steps, the developer can operate at a higher level. Klingen’s perspective comes from teams using Langfuse to trace and evaluate agents; the reference stack he describes concerns the relationship between those activities, rather than a requirement to use one particular implementation.

0:121:12
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:12 · section reference included

Connect production behavior to offline experiments

The stack starts with two kinds of work. Online tracing and monitoring capture how real users interact with the deployed agent. Offline datasets, experiments and evaluations let the team test changes before deploying them. Deployment reconnects the two: the changed agent produces new behavior, which supplies the next round of examples and failures.

Each side becomes less useful when it runs alone. An offline benchmark can drift away from what users actually ask, so better scores may reward improvements to an outdated task. Production monitoring can reveal failures without providing a repeatable way to compare fixes. Bringing the two together turns an observed problem into something the team can reproduce, evaluate and revisit after deployment.

What travels from production into the next release? The cycle below shows the connection: traces inform test cases, test cases support experiments, and evaluated changes return to production. The dataset is the bridge between an individual user experience and a repeatable comparison.

Selected presentation frame from The Self-Improving OSS Agent Stack — Marc Klingen, Langfuse at 150 secondsOpen full source frame
A slide diagrams a reference process for building AI agents.

Traditionally, people drive every transition. They read traces, update datasets, invent evaluators, form hypotheses and change the implementation. That is useful work, but it is also tedious. The question behind self-improvement is how much of this recurring labor an agent can perform while the team continues to decide what improvement means.

How it fits togetherProduction supplies the next offline test

Real users interact with the application.

Offline comparisons stay relevant when their examples come from real use; deployment produces the next round of evidence.

1:422:12
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

1:42 · section reference included

Automate fixes while humans define the target

There are loops inside this loop. At the lowest level, a model repeatedly predicts the next token. Higher levels can reason about what to fix, propose a change and verify whether it works. The practical question is where human judgment remains most valuable as more of those steps become possible.

Several decisions still determine the direction of the system:

  • Dataset selection. Review which examples enter the dataset so that it represents how the agent should be used.
  • Evaluation criteria. Decide how to judge whether those examples were handled successfully.
  • Proposed fixes. Inspect what changed, including whether the agent has overfit to the dataset.
  • Failure relevance. Decide whether an observed failure matters to the intended application. A strange production request may be outside its purpose.
Selected presentation frame from The Self-Improving OSS Agent Stack — Marc Klingen, Langfuse at 271 secondsOpen full source frame
A slide shows nested agent-improvement loops with highlighted review points.

With production traces and examples that reproduce relevant issues, an agent can search for better implementations against a defined target. It might try a different model or assemble context differently, then evaluate the result on the dataset. This is the hill-climbing step: propose alternatives, measure them and continue from changes that help. Keeping the target under human control matters because a system can become better at the benchmark without becoming better at the task users need.

3:123:42
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

3:12 · section reference included

The dataset and evaluator need to improve too

The next level of automation changes the test itself. Klingen describes teams beginning to use AI to maintain datasets and evaluation criteria, as well as propose implementation changes. A support agent illustrates why this is necessary. Its initial dataset may contain common support questions, but real users will ask a wider and changing set of questions. The dataset needs to follow the relevant parts of that use.

Recurring errors can also reveal missing quality criteria. An agent can reproduce those errors and propose evaluators for distinct requirements:

  • Content restrictions. Avoid naming competitors where that is part of the intended support policy.
  • Concision. Check whether the response is appropriately brief.
  • Language matching. Answer in the language requested by the user.

These changes deserve review because they alter what future optimization rewards. An agent that freely rewrites its own dataset and evaluator can steer subsequent improvements toward an unintended standard. Human review is especially useful here: it checks the objective before another agent spends time optimizing against it.

The objective was probably incomplete from the start. “Automate customer support” begins as a broad response to slow answers and a costly workload; it does not specify everything that happens in support. Error cases reveal requirements during implementation. New fixes then expose new error classes. The team learns the task while building it—Klingen’s image is assembling the plane while flying it—and uses that learning to revise goals, datasets and evaluators.

Selected presentation frame from The Self-Improving OSS Agent Stack — Marc Klingen, Langfuse at 452 secondsOpen full source frame
A slide reads “The target you give an agent is always incomplete” above a winding path.
5:416:11
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

5:41 · section reference included

Spend human time on direction and collect feedback in the product

A manually maintained loop can produce a high-quality agent, but it needs a team to work through the data every week. Full automation reduces that effort while risking a system that generates plenty of tokens without useful progress. The intended middle ground keeps people involved in consequential decisions and lets AI examine more failures and perform more of the recurring work.

The attraction is both time and coverage. Human attention limits how many traces a team can inspect; an agent can inspect more cases. Klingen presents lower effort and potentially higher quality as the goal of this arrangement, rather than a measured comparison between the three operating modes.

Selected presentation frame from The Self-Improving OSS Agent Stack — Marc Klingen, Langfuse at 512 secondsOpen full source frame
A slide plots agent quality against time investment and marks several approaches.

Feedback need not wait for a developer to label every output. A customer saying the support agent misunderstood provides a signal worth investigating. In an assisted workflow, a person accepting or editing a suggested reply supplies another signal through ordinary work. Those reactions can feed the improvement loop alongside deliberate team feedback, giving the lower-level automation a continuing stream of material to act on.

8:118:41
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

8:11 · section reference included

A changelog writer learns to stop leaking internal jargon

Langfuse’s changelog writer makes the loop concrete. As AI-assisted engineering increases the pace of shipping, documentation and release communication become a bottleneck. Engineers could previously spend a few hours explaining an occasional larger feature; more frequent changes make that process harder to sustain. The writer takes a merged pull request and proposes documentation and changelog updates.

The proposed update is itself a GitHub pull request. A reviewer can approve it or request changes, and the writer revises the draft in response. That review does two jobs: it controls what reaches customers, and its approvals and change requests become feedback for observability and evaluation. A coding agent using Langfuse’s agent skills can then inspect what happened and look for a recurring improvement.

The diagnosed failure is specific. The writer is factually correct, yet clarity is low and humans frequently request edits. Internal comments and implementation vocabulary reach customer-facing text: terms such as “ingestion pipeline” and “batched eval queue” describe the engineering machinery, while users need an explanation in their own domain language. Accuracy alone therefore misses an important reason the drafts require human work.

The improvement agent first proposes dataset and evaluator changes that test for jargon leakage and clearer customer-facing language. It then proposes an implementation change involving the context and skill supplied to the content writer. The sequence matters: the missing requirement becomes part of the test before the new writer is compared with the old one.

Where does the reviewer’s edit turn into a better future draft? The diagram separates the immediate revision loop from the improvement loop. A change request can repair one draft immediately; accumulated review feedback can also change the tests and instructions used for subsequent drafts.

Selected presentation frame from The Self-Improving OSS Agent Stack — Marc Klingen, Langfuse at 723 secondsOpen full source frame
A slide diagrams a changelog draft moving through review and revision.

Back-testing runs the V1 baseline and V2 candidate against the updated dataset. V1 reproduces the user-facing language problem. V2 improves that dimension while retaining the reported format compliance and accuracy. The demonstration reports a positive change without giving numeric scores or a generalization test; it supports the candidate on this comparison, rather than establishing a broader quality gain.

The team accepts a relatively permissive release policy for the improved writer because its output remains a pull request. It cannot publish directly to production. Prompt management releases the new implementation, and later reviews supply the next increment of feedback. The automation can change how drafts are produced while human review continues to control publication—a practical reason this internal application can tolerate some “YOLO.”

How it fits togetherOne review feeds two loops

The shipped change supplies material for release communication.

Review requests revise the current draft and supply evidence for improving the writer that creates later drafts.

10:1110:41
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

10:11 · section reference included

Self-improvement makes trace storage a read-heavy system

The recurring operating pattern is straightforward: collect production signals, attach them to detailed execution traces and run an improvement agent when there is enough useful new data. Klingen describes daily or weekly cron jobs, or runs triggered by a batch of interesting feedback. Inputs can include user reactions, internal labels and in-app annotations.

This changes the demands on the data system. Langfuse historically handled a write-intensive workload: ingesting traces and evaluation results. Agents that repeatedly inspect those records add a much larger reading workload because they can work through more data and issue many queries. Storage for an improvement loop must support investigation as well as ingestion.

Retention becomes part of the learning strategy. A trace from a year ago can still provide useful context, so sampling records away or keeping them only briefly limits what an improvement agent can revisit. Klingen’s closing case for owning the data layer follows from that need: retain accessible history for long periods, avoid sampling where possible and make repeated analysis affordable. The self-improving stack depends on the data remaining available after the original request is over.

Selected presentation frame from The Self-Improving OSS Agent Stack — Marc Klingen, Langfuse at 995 secondsOpen full source frame
Marc Klingen speaks at the AI Engineer World’s Fair podium.
15:1015:40
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

15:10 · section reference included

Resources

Read the complete timestamped transcript
  1. 0:12

    Everyone, super excited to be here. Uh, okay. My laptop's alerting me that this starts now. Uh, hi, everyone. I'm Marc, one of the founders of Langfuse. Super excited to chat about, um, what we have seen, how, how people like self-improve agents now and what kind of, like, open source reference stack, uh, we see emerging. I mean, obviously, my view is based on working with our community of people that use Langfuse to trace and evaluate their, their agents. But I think-- I try to keep the talk mostly to, to, like, more generic takeaways that work with whatever stacks you're

  2. 0:42

    using. Um, but yeah, happy to chat about the different pros and cons, uh, after the talk. High level, uh, I think this is the year, uh, where, uh, teams talk about, like, up-leveling the, the level of, uh, where, where they operate themselves, where increasingly good, good models just abstract everything downstream. I mean, coming from, I don't write prompts, I have loops, uh, Peter writing about it, auto research becoming really popular. So I think this is kind of where things are going on the, like, like, producing, um, application side, but also on the how to

  3. 1:12

    actually, um, build good, good agent applications. And, um, over the years, early models have, have expanded a lot in capabilities. Um, back when, when we started working on Langfuse, this was right when ChatGPT was released, so, uh, where, where, like, even the simplest applications didn't really work. So for example, we built, like, GitHub issue to pull request automation, similar to what Devin or Cursor are building. But back then, we could only manipulate, like, single page applications, like a single HTML file, because multi-file edits were too complicated. Since then, all of these things really accelerated a lot,

  4. 1:42

    and this really made, made loops now possible. And the same time, a more, like, reference stack emerged of how great AI agents are built because building agents is different from building applications. It's like, um, like a link of, like, tracing and monitoring, like, online how users really use an agent application, and then, like, an offline component of how to build datasets, how to experiment, uh, regarding, like, making changes to these agents, and then evaluating offline whether things work, then deploying new to production, then learning again for

  5. 2:12

    real users who use the application. Uh, so it's like this, this kind of loop, where traditionally, I mean, you have, like, observability and the data docs here on the tracing monitoring side, and you have, like, the MLOps toolings like MLflow, Weights & Biases more on the, like, offline side. Where, like, this new category emerged, and that's also why we built Langfuse, because you need to bring the online and the offline together. As, as otherwise, uh, like, you either benchmark on data that's inaccurate, so the datasets are not on loop and, like, uh, not in sync with what's happening in production, or you're monitoring, like, production data, but you don't really benchmark offline, so you need to

  6. 2:42

    bring the, the both together. But this has been, like, a lot of manual labor because you need to look at traces, update these datasets, think about, like, new evaluators, uh, then make these changes, create new hypothesis. Like, it's, it's, like, a lot of work going into this process. And since starting working on Langfuse, we try to educate teams how to, how to do this process well. But it's, it's usually, like, they ask that, like, "Oh, models got better now. How can we take ourselves out of the loop to, uh, to automate it?" Because it's really tedious, uh, to get things going. And that's what I wanna talk about, how we see how people go

  7. 3:12

    from manually driving this loop to using AI to, to drive this loop. Um, w-we use the visualization, um, like, of, for example, like the Loopcraft article or, like, many, many articles in the space that, like, layer different loops on top of each other. How, I mean, at the, at the lowest level, you just have, like, this token loop of what's the next, what's the next token that's produced. And on the very high level, it's, uh, like, like, the most abstract form o-of thinking where you as a human are still involved. So what we see how, um, uh, we, we go through them one by one.

  8. 3:42

    High level, I just wanna highlight these are, again, possible because models improve. So, uh, when, when we started working on Langfuse in twenty twenty-three, w-like, I mean, we had GitHub Copilot, so, like, the lowest level or the, uh, level two. Um, the level three was possible, like, more in twenty twenty-four, and now twenty twenty-five, twenty twenty-six is more like the higher level loops of actually agents reasoning about what, what to even fix, uh, how to propose new fixes, and how to, how to verify whether they actually work. Where humans are currently still involved is, one, you, uh, need to align

  9. 4:11

    datasets of example questions of how the agents are used. This is, like, where, where humans need to review, um, uh, like, like what, what gets amended to the dataset and how to then evaluate whether the agents actually work. These are the, the two blue things in, in loop number three. Then on the proposal of fixes, because if, um, like, your agent proposes how to change an agent, you still wanna see what was changed because you have the risk of, like, over fitting it to, um, to, like, the dataset. And the failure definition, um, so I mean, usually

  10. 4:41

    agents find all sorts of different, uh, failure patterns, but maybe some of them don't really matter that much and you're like, "Huh, this kind of use case for the agent, it doesn't really matter that much. Like, this is, like, out of what the agent should be actually be doing." So we don't need to over fit no application to this random quirk that we found in production, but we wanna, like, be in the loop to, to monitor this. And now coming from this loop pattern, we try to dissect it again into more like a workflow where, where we currently see A-AI used the most is in proposing fixes. So if we have

  11. 5:11

    production traces, and we have, like, even criterion datasets of how to reproduce issues that we find in production, then there's how do we hill climb against the datasets? That's, like, a really cool, uh, way to kind of like loop your way to a success with, with agents, because usually there are like ten different things of what you could be doing. I don't know. Use a different model, like aggregate context in a new different way. Try whatever you find on X, uh, in a given week to, to see whether it can, like, uh, make a dent into improving the agent. But really, like, you maintain the boundary of what you optimize

  12. 5:41

    against, and you use AI mostly for, like, proposing improvements to the agent implementation. Like, over the last, I'd say month, we've seen more and more teams also using, uh, like, like AI and, and their loops to maintain eval criteria and maintain datasets. So, uh, usually what goes in here is you wanna, like, align how a dataset, uh, looks like for an agent with what actual users are doing. So for example, if we have an customer support application, like support agent, then the dataset should be, like, common support

  13. 6:11

    questions. But, like, users do all sorts of things with a support application, so you need to continuously keep it in line with, uh, with what users are doing, and you can use agents for that. And for evals as well, if you see usual error patterns, then you can reproduce the error patterns and propose new eval criteria on these often datasets. So for example, I don't know, I wa-- I don't wanna name competitors, I wanna be concise, I wanna answer in the same language as the user actually requested, um, like they question my customer support application, so you can also propose new evaluators. This is, I think, rather

  14. 6:41

    new-ish over the last months, um, and where we see, like, the AI more auto looping. However, usually teams are now still involved in reviewing these changes because they create the new boundary for how then other agents try to auto improve against it. If you give this completely out of fam, you risk, like, that you just create more slop. Where, uh, we've used this in, like, a blog post of ours, where I'd say the target of an AI application is always changing because you don't really know. Like, usually you start with a very high level task of, I stick with customer

  15. 7:11

    support. You stick with someone says, "Hey, in our company, cu-customers need to wait a long time for support, uh, answers, and we have a lot of people employed to respond to them." AI should be able to do this automatically, or at least aid in customer support. But it's a very vague target because people don't even know what happens in customer support every day. So it's kind of like, "Oh, we wanna automate all of it." However, then, on the way of doing this, you figure things out based on, like, error cases of what you even need to do. And that's, like, this kind of, like, map that you-- Like, you assemble the plane while you're flying it. Um, and, uh, how we then look at

  16. 7:41

    the, uh, at this kind of, like, looping graphic is, you really wanna be, like, involved in setting the, setting, like, the, the goals and how you make changes to datasets, evaluators. You wanna be in that loop involved because you set thereby the direction, and then AI can automate against it. And then off of the newly implemented changes, you will see new error classes and you can then, like, change course again on the most highest level. But thereby you pull yourself, uh, out of a lot of, like, the more manual and tedious work on the lower levels.

  17. 8:11

    Um, this is then the high level view of how, how I would conceptualize, uh, like the different scenarios that you can find yourself in. I'd say if you implement this, uh, like, AI workflow well of tracing online, uh, how your agents are working, implementing evals on how they are working, um, having datasets, uh, of example questions, and then benchmarking against them, you are in this era-- uh, in this class. Like, very high time invest because you need a team to really work with this data every week. But also, quality of the, of the agent is pretty high. I

  18. 8:41

    mean, that's what, uh, teams have been, for example, using Langfuse for, for the last years, and this is how you can get to, I'd say, success with your agent application. Um, if you're fully automated, I would say you don't really need to invest any time because you're just like, I don't know, Codex goal mode your way to success or, like, you add higher level loops on top of it. But you also risk that, like, it produce a lot of tokens but don't really make sense. So I think you wanna be in this kind of, like, promise land of, you don't really need to invest a lot of time, but you're, like, in the loop at the most important steps of, like, setting the

  19. 9:11

    direction of where you wanna go, but you take yourself out of everything that's tedious. Thereby, you're way less time invested than running it manually, but also the quality can be even higher because usually if you need to do everything manually, you're bottlenecked by your own time, your own, like, perseverance, endurance, uh, or, like, patience to look through so much data, and AI can just look so-- through, like, so much more error cases. This agent will be higher performing, and you need less time invest. That's the-- I think that's what we are w-going for. And how this then looks like is,

  20. 9:41

    uh, you involved at the top, but also, I mean, you want ideally your application to feed in interesting signal. So for example, when we talk about the, uh, like, customer support application use case, usually, I mean, customers can swear at the support agent. Customers can be like, "Oh, no, you don't understand. Uh, like, uh, this was not correct." Or, like, you can propose messages to an internal agent who then can accept them or change them when sending them out. So there's, like, a lot of, like, implicit signal that can feed into your process without you as the agent developer needing to do

  21. 10:11

    this manually. Uh, thereby you have, like an, like an in stream of, of signal that your loop can act on. Um, so it's either you, your team, or some signal coming in. Um, and the lower level, uh, loops can be automated. That, that's high level how we, how we think about it. And like a, uh, like we put together, like, a super quick demo application, um, that can, that can show this. Um, so for example, like we as the Langfuse team, we have a big problem now because our, like, engineering team uses AI a whole lot, thus they ship a whole lot,

  22. 10:41

    and somehow we need to update our customers, uh, and user base, uh, what is, what is new every week. Because, uh, like in recent weeks, like, a lot was released, uh, and usually you need to update documentation, update a changelog, post about it on socials. And historically, like, engineers did this themselves because they maybe released, like, a bigger feature, like, once every month. So it's okay to spend a couple of hours. But now if they take themselves out and really ship a lot of, like, products, they wanna ideally, uh, like, automate also that kind of, like, release process, because otherwise we are bottlenecked by how fast we can

  23. 11:11

    communicate about changes. So, um, we thought about, like, putting together, like, a changelog writer that just uploads, like, like, updates a lot of, like, the public assets based on what we've shipped. Um, however, then the question is, i-is, like, how we communicate good? Because if we communicate badly about what we've released, then, I mean, we, we kind of like sabotage our releases if we, like, misrepresent, for example, what has even been released. So, um, setup is we have a merged PR, we have a changelog writer that, uh, suggests a change to

  24. 11:41

    documentation, but then there's like a, just like a code review step as like a GitHub PR where someone can review and either, like, approve on GitHub and get it merged or, like, submit, like, change requests of, like, well, what needs changes in this, um, uh, in this, in this change in documentation and, and the changelog. Then the change writer would act on this again to update the draft until it's finally released. And now both of these kinds of like, either like approvals or, uh, requests for changes can get feeded into like the AI observability and

  25. 12:11

    evals to, to serve as a basis for auto-improvement of that changelog writer agent. So what we've now done is, um, just use a coding agent pointed at the, like, Langfuse agent skills, but generally, I mean, this would also work with other systems, but yeah, Langfuse is like, uh, well positioned for this to, to ask like, "Okay, well what has happened in this application? How can we improve it?" I'll, I'll pause a couple times please, otherwise it, it'll go, it'll go by very fast. So what we, for example, here identified, that the writer, uh, so

  26. 12:40

    the changelog writer is factually, uh, like very correct in how it makes changes to documentation, but clarity is low. The edit ratio of us humans making ch-- like submitting change requests is still high, and that we leak internal jargon to our customers because maybe our code has like internal comments that we leave there for ourselves for the future, but, uh, like our customers don't think about our application as like, I don't know, like an ingestion pipeline or as like, I don't know, some kind of like

  27. 13:10

    batched eval queue. Like they don't know. This is like implementation detail. We should talk about it as like um, user domain language, and that's like, uh, something we have identified here where now the agent suggests changes to the datasets and evaluators to test for leaking jargon and, um, also like increasing clarity in customer domain language, basically. So, um, agent now closes this gap, suggests changes to the datasets and evaluators,

  28. 13:41

    and now, um, has an idea of how to change the agent implementation to basically fix this problem. Because in the end, it's kind of like just like a, a problem of how we provide context, how we like provide the skill that then educates this, um, content writer. And now we can backtest it basically on the updated dataset to see whether the new implementation would do better than the old implementation. And I'll stop here again, where we have basically the baseline. So, um, basically we reproduce the issue of

  29. 14:10

    user-facing language. So let's talk about features in a way of how users would talk about it. So we reproduce this kind of problem with the V1 baseline that's basically existing implementation of the agent. And agent came up with a new implementation where we see the V2 candidate. We are still doing on format compliance and accuracy in the same way as we did before, but we improved on like user-facing language. So we see a positive delta, which sounds like a good change now. Um, so, uh, agent reasons about it, uh, as this, I mean, this is a bit-- I mean

  30. 14:40

    internal application, it's a bit YOLO. So we are like, okay, we are good of, of just merging this and then acting again on the next change because this agent never publishes anything to actual production. It just raises PRs on a repo, so it's okay if it kind of like auto-improves, and then we just provide feedback again to whatever the next increment is. So, um, here we use our like prompt management logic so that the agent can just basically feature release this new, um, implementation. And, uh, now it just summarizes basically, um, how, how we

  31. 15:10

    acted on this. So high level, um, this is what we have seen like most of our best users do already, uh, when they, when they build agents, like collect production signals, uh, track them alongside like detailed execution traces, and then run agents like on a loop, uh, usually like a cron job, like every day, every week, whenever you have like a batch of basically interesting user data again to then act on it. Um, either it's just user feedback, how, how this happens in this case, or internal labeling from like a production system. But it could also be you ask people to annotate

  32. 15:40

    data like in-app. For example, in Langfuse we also allow like for annotations to feed into that, uh, into that queue. And yeah, that's very high level, uh, what we've seen. Happy to give you like a more deep dive, um, demo at our, at our booth or after this talk. I just wanted to, to wrap it up very concisely with how we see this, the space evolves. And, um, and I think what's, what's interesting here is it drives the need for like a very scalable data system because you'll wanna have agents really loop on this data and, uh, like produce lots of queries. This, we see how Langfuse was historically

  33. 16:10

    very write-intensive to like, um, ingest evals and traces. Now it's way more read-intensive because agents can like tr-- like chew through so much more data. Interesting, number one. And two, you wanna really like own the data layer because, uh, like now the traces of like, for example, a year ago are interesting context and you don't wanna want them to be locked up in like, uh, like a more commercial system or something where, uh, you need to sample data or retain them for a long-- short time period. But we wanna like retain them for a long time period and not sample them. And yeah, this really drives the need for like a scalable, cheap

  34. 16:40

    solution and, uh, we've worked with many customers on this, so happy to talk, uh, one-on-about, uh, one-on-one about this if you're interested. Thank you so much for your time.