Jess Wang is a software engineer turned technical educator whose work helps developers judge whether AI tools deliver useful results. Her projects and writing connect developer experience with practical questions of reliability: whether a generated command is correct, why an autonomous agent stalls, and what successful work actually costs.
From software engineering to evaluation education
Wang worked as a software engineer at Microsoft and DoorDash before moving into developer relations at Warp. Her developer profile and writing connect that career with her RubberDuckiee identity and educational YouTube channel. At Warp, she combined technical writing, video production, documentation, product launches, and building projects.
Her early AI writing already centered on trust. In a 2022 exploration of natural-language Git commands, she tested Warp’s AI Command Search against practical tasks rather than assuming plausible-looking commands were correct. Generated commands could give experienced developers a useful starting point, but beginners needed enough understanding to judge whether a suggestion would do what they intended. Reliability was a teaching problem as well as a product problem.
With Advait Maybhate, Wang co-authored an explanation of Warp’s Same Line Prompt. Placing a command beside its prompt looked like a small interface change, yet it affected text layout, scrolling, selection, wrapping, and shell behavior. Their account showed how a familiar usability request reached into the terminal’s rendering system.
Building remained part of her educational work. She collaborated on Share-Brewfiles, an open-source project that lets developers publish and update their installed tool collections through a command-line interface. Searchable pages and package leaderboards help people discover tools through other developers’ setups. Her development retrospective examines a consequential architectural tradeoff: as the site became more interactive, its Astro pages filled with React components, reducing the advantages that had motivated the original framework choice.
At the beginning of 2026, Wang moved from Warp to Braintrust, following a growing interest in evaluation education. The move extended her work on helping developers understand their tools into helping them measure AI behavior. Correct-looking output was only the beginning; developers also needed ways to investigate failures and compare the quality and cost of different approaches.
Measuring agent behavior and useful outcomes
Wang’s evaluation work connects results with the behavior that produced them. Her examples make abstract measures concrete by following agents through tasks, interruptions, and unsuccessful attempts.
Agent observability: For an experiment with the Ralph Wiggum development loop, Wang built a habit-tracking calendar and asked an agent to improve it. Permission prompts interrupted unattended execution despite permissive language in the prompt. Traces helped her locate the configuration problem, showing why an agent’s instructions alone do not determine whether it can complete a task. Investigating the stalled work gave her a way to address the problem before repeated attempts consumed more time and money.
Cost per successful outcome: In her analysis of Exgentic agent traces, Wang examines models alongside their harnesses—the software that organizes their tools and execution—and the tasks they attempt. An agent that stops early may produce an inexpensive run while accomplishing nothing. Measuring the cost of completed work changes that comparison. She retains the limits of the analysis, including uneven task coverage and estimated success scores, rather than treating a ranking as a universal answer.
Search quality and context: Wang distinguishes finding relevant code from giving an agent enough context to act on it. Her comparison of agentic and vector search illustrates the gap: retrieved chunks could locate nearby code while omitting connections the agent needed, prompting further searches. In that experiment, equal accuracy came with roughly four times the cost for vector search. The useful lesson is to evaluate retrieval within the agent’s complete workflow, where missing context can turn a promising search result into additional work.
Framework-independent evaluation: Her guide to instrumentation across frameworks and providers explains how model calls, tool use, and application logic can share a trace structure. Consistent instrumentation lets teams compare quality, latency, and cost as they change frameworks or providers, preserving their ability to investigate behavior while the implementation evolves.
Building models around everyday questions
Wang’s projects also extend beyond developer tools. Her pickleball rating predictor uses match histories, engineered features, and a gradient-boosting model to estimate how hypothetical doubles results might change a player’s rating. It approximates an opaque rating system rather than implementing its official algorithm. The project carries her educational approach into a different setting: start with a concrete question, build something readers can inspect, and explain what its behavior reveals about the underlying system.
Jess Wang walks through an end-to-end coding-agent evaluation: turn real bug fixes into tasks, hold the agent harness steady, score with the repository’s tests, and inspect traces to understand why similar accuracy came with very different costs.
An eval combines a dataset, a task, a scorer, and an experiment configuration to answer a specific question about an AI system.
In this small experiment, equal reported accuracy concealed about four times the cost for vector search. Traces linked the extra calls to repeated retrieval that lacked surrounding code relationships.