Jake Broekhuizen develops practical approaches to helping AI agents improve through experience. He joined LangChain as one of its first deployed engineers and spoke as a member of LangChain Labs at the AI Engineer World’s Fair 2026. His work connects three problems: giving agents useful memory, detecting failures in real conversations, and turning successful behavior into training data.
From enterprise software to agent improvement
Broekhuizen studied electrical engineering and computer science at UC Berkeley. After college, he began as a software engineer at ServiceNow, then moved into machine-learning engineering to build a platform that made OpenAI’s APIs usable across business units. That work involved connecting language models to the workflows and data of a large organization, and brought him into contact with LangChain and LangSmith. He subsequently joined LangChain to work closely with customers building enterprise agents.
His published work in 2026 developed several ways to learn from those agents after deployment. In June, he published a memory architecture for converting selected lessons from interactions into reusable context. In August, he co-authored the introduction of LangSmith Tuned Evaluators with Shamik Karkhanis and Vivek Trivedy. In September, he co-authored the launch of LangSmith Fine-Tuning with Ankush Gola and Trivedy. These approaches change different parts of an agent system: the context it reads, the feedback used to assess it, and the model’s behavior through training.
Turning recorded experience into useful memory
Broekhuizen’s approach to agent memory starts with a distinction between recording an interaction and learning from it. Traces provide evidence of what happened. They become useful memory when a selected lesson becomes persistent context that the agent reads and acts on in a later run.
In Giving AI Agents Memory That Learns, he illustrates the problem with a financial-services agent whose tone repeatedly shifted from advisory to pushy. Observing the drift could help a team diagnose it; changing the context used on subsequent runs could help prevent it. The example explains why he emphasizes procedural memory: instructions, workflows, skills, and tool-use rules can directly alter behavior.
Two engineering choices make that distinction practical:
Select the right intervention. His architecture separates recording behavior, diagnosing recurring problems, and updating context. LangSmith Observability stores traces, Engine analyzes them, and Context Hub manages the resulting instructions and skills. Most traces should remain history. A failure may instead call for a code change, a clearer tool description, or an evaluation example. An agent that repeatedly calls tools in the wrong order may need a better procedure rather than more historical examples.
Make updates reach the agent. A revised instruction has no effect if the runtime continues loading a cached version. Broekhuizen treats retrieval and caching as part of the learning mechanism, and recommends human review for consequential procedural changes. Writing a lesson and applying it are separate engineering responsibilities.
Detecting failures and learning from successful runs
Broekhuizen’s collaborative work on evaluation and fine-tuning addresses two other requirements for improvement: recognizing mistakes and preserving enough context to learn from good examples.
Finding failures without ratings.LangSmith Tuned Evaluators addresses conversations where something went wrong without a system error or explicit user score. Its Perceived Error evaluator looks for corrections, repeated requests, contradictions, and unresolved misunderstandings. It attaches an explanation to the interaction so teams can investigate the problem and build tests around it. This extends quality assessment to users who express dissatisfaction through the conversation itself.
Training on the context behind an action. In LangSmith Fine-Tuning, successful agent trajectories become examples for supervised fine-tuning. A trajectory preserves messages, tool calls, and results in sequence, including which tools were available at each step. Exporting only the final conversation can lose information the model relied on when choosing an action. The open-source smithtune CLI connects dataset preparation, managed training, evaluation, and deployment.
The fine-tuning work also illustrates the consequences of poor example selection. The team’s initial, less selective code-review training set reduced performance, prompting an additional data-review stage. Their recommended progression begins with improving tools, instructions, and execution logic, then considers fine-tuning when recurring mistakes remain and good demonstrations are available.
Jake Broekhuizen explains how a financial assistant’s tone failures can become instructions for its next run—and why selecting lessons, refreshing context and reviewing rule changes matter as much as storing them.
A trace becomes memory when a selected lesson is written into persistent context that future runs actually read.