AI agent development, evaluation, and reliability
Haize Labs
Haize Labs builds AI agents for enterprises handling mission-critical work. Its proprietary Reliability Harness combines tailored agent workflows and post-trained models with supervisor models, simulation testing, adversarial testing, runtime guardrails, and annotation tools. These capabilities help teams test behavior before deployment, monitor agents during operation, and refine performance using human feedback. Its Verdict framework lets developers compose multiple LLM judges into evaluation systems that combine reasoning, verification, debate, and aggregation.
Founded in 2023 by Leonard Tang, Richard Liu, and Steve Li, Haize Labs is led by Tang as CEO. Its research addresses how to evaluate and train models when correctness cannot be checked directly. TournO combines comparisons between candidate responses with individual reward scores for reinforcement learning. Responses compete in small tournaments to produce a locally comparative training signal; the method adjusts the balance between the two reward types to reduce overfitting.
Haize Labs works directly with customer teams through engagements spanning data audits, agent development, validation against human experts on unseen data, and knowledge transfer. Ongoing software and services support improvements based on production interactions. This delivery model connects its evaluation technology to the practical work of building, testing, and maintaining agents for specific business tasks.
1 talk
Newest first1 speaker at AIE
Affiliations reflect their AIE appearances, not necessarily current employment.
Messages from the stage
Testing across conversational modalities
The talk describes enterprise banking voice-agent evaluation across text, audio, and multi-turn conversations, alongside Verdict's composable LLM-as-a-judge architecture and the reinforcement-learning-trained j1-micro reward model.
Affiliations reflect each recorded session, not necessarily current employment.
