AI evaluation consulting and education
Parlance Labs
Parlance Labs provides consulting and education for teams building AI products with large language models. It works with companies that have moved beyond prototypes but need repeatable ways to diagnose failures and decide what to improve next. Its AI Evals for Engineers & PMs course teaches evaluation methods, while its evals skills give coding agents reusable workflows for auditing evaluation pipelines, reviewing traces, generating test inputs and validating automated evaluators.
Founder Hamel Husain runs the business, drawing on machine learning experience at GitHub and Airbnb. Repeated evaluation problems in consulting engagements led him to develop the course with Shreya Shankar. His published Critique Shadowing approach starts with domain experts making pass/fail judgments and explaining their reasoning, then uses those examples to develop and test an LLM judge. The emphasis is on evaluation criteria grounded in a product’s actual tasks and human judgment.
In 2026, the company’s website reported 4,500 engineers and product managers had taken its course. Separately, Husain and Shankar described experience helping 50+ companies in their joint publication on evaluation skills.
3 talks
Newest firstHow to construct domain-specific LLM evaluation systems.
AI Engineer World's Fair 2024 · 18:45
What We Learned From A Year of Building With LLMs
Eugene Yan · Hamel Husain · Jason Liu · Dr Bryan Bischof · Charles Frye · Shreya Shankar
AI Engineer World's Fair 2024 · 35:21
1 speaker at AIE
Affiliations reflect their AIE appearances, not necessarily current employment.
Start here
- How to construct domain-specific LLM evaluation systems.
Start here for a concrete application case showing how evaluation checks fit into existing CI and Metabase workflows, with LangSmith used where useful.
Hamel Husain · Emil SedghAI Engineer World's Fair 2024
- What We Learned From A Year of Building With LLMs
Use this discussion to understand the strategic, operational, and tactical work required to turn an LLM demo into a production application.
Eugene Yan · Hamel Husain · Jason Liu · Dr Bryan Bischof · Charles Frye · Shreya ShankarAI Engineer World's Fair 2024
- How To Build an AI Strategy That Fails
Read the inverted strategy advice to identify where hype, GPU spending, and gimmicky applications can displace concrete customer problems.
Hamel Husain · Greg CeccarelliAI Engineer Summit 2025
Messages from the stage
Build evaluations around observed failures
The Rechat discussion with Emil Sedgh describes starting with inexpensive unit tests and assertions, then expanding coverage with synthetic real-estate-agent prompts and repeated evaluation cycles.
Plan for change beyond the demo
The shared lessons session questions models as a durable competitive advantage and considers provider switching alongside production monitoring, guardrails, and evolving team workflows.
Recognize organizational failure modes
Hamel Husain and Greg Ceccarelli's satirical presentation connects weak AI strategy to executive isolation, vague goals, excessive documentation, and vanity metrics.
Affiliations reflect each recorded session, not necessarily current employment. These appearances are joint sessions with speakers from other organizations or backgrounds.


