Hossein Niazmandi’s work connects enterprise data platforms with the practical challenge of improving AI agents in production. In 2026, he led solutions engineering in the US West at Braintrust. His approach to agent quality connects evaluation before release with observability afterward: production failures become test cases, and those tests help teams judge whether their changes improve the behavior users encounter.
From enterprise data to agent quality
At Databricks, Niazmandi worked with enterprise customers and co-authored a March 2025 account of Prisma Cloud’s data-platform evolution with Ram Katakam, Krishnan Narayan, and Walton Stephens. The article examined how Palo Alto Networks approached cloud security as a data problem. Separate security modules produced signals that became more useful when normalized into common objects and correlated across systems, revealing relationships and risks that were harder to identify in isolation.
Maintaining that shared data platform also carried an operational cost. Thousands of processing jobs, bespoke integrations, scattered governance, and infrastructure troubleshooting consumed engineering time. The authors described consolidating those responsibilities on Databricks so teams could devote more effort to developing security capabilities. Niazmandi contributed to the collaborative account; the platform and its results belonged to the teams building and operating it.
Niazmandi left Databricks to build alongside customers at Braintrust, taking responsibility for US West solutions engineering. The move reflected his conviction that production AI quality needed sustained attention as teams put agents to work. Non-deterministic language models require both controlled evaluations and observation of real behavior. Side-by-side experiments help engineers, product managers, and domain experts compare changes; production traces reveal failures their initial tests did not anticipate. Connecting the two creates an improvement loop rather than a static scorecard.
Turning production behavior into better evaluations
Niazmandi and Braintrust product manager Rahil Sondhi’s jointly authored guide to Patterns, Topics, and Loop develops active observability: using production behavior to discover problems and guide the next round of evaluation. Their work explains three related capabilities:
Discovering unanticipated failures: Explicit evaluators measure problems a team already knows to test for. Topics classifies traces, while Patterns investigates behavior and supplies supporting examples. A newly discovered failure can become an evaluator, allowing the team to measure how widely the problem occurs.
Evaluating whole conversations: A plausible final answer can conceal a failed tool call or an unsuitable sequence of actions. Evaluation can target an individual operation, an entire trace, or a multi-turn session. Grouping related traces preserves conversational context so teams can assess the agent’s trajectory as well as its answer.
Measuring whether fixes hold: Once a team identifies a failure, creates an evaluation, and ships a correction, it can monitor whether that specific problem declines in production. This ties testing to customer experience and gives the next improvement a concrete starting point.
Niazmandi also explains why this loop demands more than an evaluation interface. Agent traces can be large and contain nested JSON, while teams need both real-time inspection and longer-running analysis. In his account of building an evaluation platform, these demands explain Braintrust’s development of BTQL. He extends the same approach to coding agents running evaluations and surfacing failures people had not thought to look for, with humans reviewing the results and deciding which changes to accept.
Hossein Niazmandi follows evals from a spreadsheet to a production feedback loop, showing how collaboration, large traces and competing query workloads turn a simple testing interface into a systems problem.
Offline evals establish expectations; production observability tests those expectations against real interactions.