The Benchmarks Game: Why It's Rigged and How You Can (Really) Win
AI Engineer World's Fair 2025 · 11:20
AI agent simulation and evaluation
Scorecard builds a simulation and evaluation platform for teams developing AI agents. Developers can define scenarios through a no-code interface or Python and TypeScript SDKs, test live or staged agents, and inspect failures. Scorecard Playground supports comparing models and prompts, while production tracing helps teams turn real failures into reusable tests. Its Autoloop beta, described in 2026, extends this workflow into improvement: simulations run against local environment snapshots, and tool-equipped graders and human annotations inform proposed changes to agent code and grading rubrics for review.
Founder and CEO Darius Emrani previously led simulation work at Uber’s Advanced Technologies Group and product teams building simulation and evaluation infrastructure at Waymo. That experience informs Scorecard’s use of repeatable scenarios to test agent behavior. The company also introduced AgentEval.org in 2025, an open benchmarking initiative initially focused on legal AI. It brings together existing domain-specific datasets, studies and assessment practices to help researchers, legal practitioners and vendors evaluate applications in practical contexts.
By September 2025, the company reported running millions of evaluation tests for customers, including Thomson Reuters, which uses Scorecard to test and deploy its CoCounsel legal AI suite. That year, Scorecard announced $3.75 million in seed funding from investors including Kindred Ventures, Neo, Inception Studio and Tekton Ventures.
AI Engineer World's Fair 2025 · 11:20
Affiliations reflect their AIE appearances, not necessarily current employment.
Emrani identified selective inference configurations, privileged benchmark access, and a preference for polished responses over accuracy as ways benchmark comparisons can become misleading.
Affiliations reflect each recorded session, not necessarily current employment. The published summary corrects the talk's AutoRover reference to AutoCodeRover and qualifies its FrontierMath access claim by noting a separate holdout.