AI evaluation and safety research
METR
METR (Model Evaluation and Threat Research) is a nonprofit that evaluates frontier AI systems to help developers and policymakers understand autonomous capabilities and catastrophic risks. Its assessments test whether agents can conduct research, develop applications, carry out cyberattacks or resist shutdown. Its Inspect Hawk platform lets evaluation teams run Inspect AI tasks on AWS, managing isolated execution, credentials, logs and results with a web interface. Hawk supplies operational infrastructure around the UK AI Safety Institute’s evaluation framework.
Founded by Beth Barnes in 2022 as ARC Evals within the Alignment Research Center, the organization adopted the METR name in 2023; Barnes is its current CEO. Its task-completion time horizons measure the human-expert task duration at which an agent is predicted to succeed with a specified probability. Based primarily on software engineering, machine-learning and cybersecurity tasks, these measurements describe task difficulty, rather than uninterrupted agent runtime.
METR helped prototype the Responsible Scaling Policies approach. In August 2026, it reported approximately $71 million in funding commitments raised over the preceding six months to expand its team and research. METR is donation-funded, with a small European AI Office technical-assistance contract. It has not accepted funding from AI companies, although they provide significant free tokens for evaluations, research and engineering.
2 talks
Newest firstWhy Agent Hype can fall short of reality – Joel Becker, METR
AI Engineer Code 2025 · 21:22
1 speaker at AIE
Affiliations reflect their AIE appearances, not necessarily current employment.
Start here
- Why Agent Hype can fall short of reality – Joel Becker, METR
Start here to understand how a randomized field study provides a different test of AI usefulness from human-calibrated task-horizon measurements.
Joel BeckerAI Engineer Code 2025
- Long Tasks and Experienced Open Source Dev Productivity
Use the workshop for its discussion of AI Village, where agents attempt loosely specified real-world goals, alongside software-development evaluation.
Joel BeckerAI Engineer Code 2025
Messages from the stage
Interpreting capability measurements
Becker explains why benchmark scores and autonomous task completion need not predict gains on context-dependent software work. Small-study limitations and uncertainty remain central to interpreting the field evidence.
Evaluating adoption on natural tasks
The workshop examines agent-adoption J-curves, unreliable productivity self-reports and Cursor familiarity as complications in evaluating experienced developers on their everyday work.
Affiliations reflect each recorded session, not necessarily current employment.

