← All organizations

AI agent simulation and evaluation

Scorecard

Scorecard builds a simulation and evaluation platform for teams developing AI agents. Developers can define scenarios through a no-code interface or Python and TypeScript SDKs, test live or staged agents, and inspect failures. Scorecard Playground supports comparing models and prompts, while production tracing helps teams turn real failures into reusable tests. Its Autoloop beta, described in 2026, extends this workflow into improvement: simulations run against local environment snapshots, and tool-equipped graders and human annotations inform proposed changes to agent code and grading rubrics for review.

Founder and CEO Darius Emrani previously led simulation work at Uber’s Advanced Technologies Group and product teams building simulation and evaluation infrastructure at Waymo. That experience informs Scorecard’s use of repeatable scenarios to test agent behavior. The company also introduced AgentEval.org in 2025, an open benchmarking initiative initially focused on legal AI. It brings together existing domain-specific datasets, studies and assessment practices to help researchers, legal practitioners and vendors evaluate applications in practical contexts.

By September 2025, the company reported running millions of evaluation tests for customers, including Thomson Reuters, which uses Scorecard to test and deploy its CoCounsel legal AI suite. That year, Scorecard announced $3.75 million in seed funding from investors including Kindred Ventures, Neo, Inception Studio and Tekton Ventures.

www.scorecard.io

1 talk

Newest first

1 speaker at AIE

Affiliations reflect their AIE appearances, not necessarily current employment.

Messages from the stage

What benchmark comparisons reward

Emrani identified selective inference configurations, privileged benchmark access, and a preference for polished responses over accuracy as ways benchmark comparisons can become misleading.

Affiliations reflect each recorded session, not necessarily current employment. The published summary corrects the talk's AutoRover reference to AutoCodeRover and qualifies its FrontierMath access claim by noting a separate holdout.

Company sources · checked 2026-08-28