Real ROI: Lessons from Enterprises that Have already succeeded with LLMs Scale
AI Engineer World's Fair 2024 · 20:01
LLM evaluation, prompt management and observability
Humanloop built an enterprise platform for LLM evaluation, prompt management and observability, used by teams at Gusto, Vanta and Duolingo. Engineers and product teams could version prompts, compare model outputs and evaluate applications using human feedback, code and LLM judges. Its Flows tooling traced multistep applications, supporting debugging and evaluation of both individual components and complete workflows. The Humanloop platform is no longer available: its shutdown date was September 8, 2025.
Founded in 2020, Humanloop’s founders were Raza Habib, its CEO; Jordan Burgess, its CPO; and Peter Hayes, its CTO. Its earlier engineering work applied active learning to select useful training examples. In a 2021 experiment with Black Swan Data across three NLP datasets, the company reported at least 40% less labeling while improving model performance. All three founders and several engineers and researchers joined Anthropic in a 2025 acqui-hire; the arrangement did not transfer Humanloop’s assets or intellectual property.
By December 2024, Humanloop reported processing millions of LLM logs daily across thousands of AI products deployed in production. The company had secured $7.91 million in funding by August 2025, following its origins as a University College London spinout.
AI Engineer World's Fair 2024 · 20:01
Affiliations reflect their AIE appearances, not necessarily current employment.
Habib breaks LLM applications into models, prompts, data selection or retrieval, and optional function calling. He argues against unnecessary orchestration complexity.
Affiliations reflect each recorded session, not necessarily current employment.