Training Frontier Models to Out-Think Hackers — Uri Rolls, Arithmetic & Thom Wolf, Hugging Face
AI Engineer World's Fair 2026 · 17:28
AI cybersecurity research and evaluation
Arithmetic explores how AI models can become more capable cybersecurity defenders through evaluation, training data, and benchmarks. Its stated research goal is to improve models’ ability to reason about vulnerabilities, addressing the gap between performing reconnaissance and making the logical leap needed to uncover an exploit.
The work presented in 2026 includes an access-control benchmark using blackbox environments built from vulnerabilities discovered by researchers. Deterministic grading makes these environments a way to test model capabilities against concrete security problems.
Arithmetic’s stated direction centers on improving the models themselves as a route to stronger defense. High-quality evaluations and data underpin that ambition: developing defenders capable of outperforming attackers, rather than treating today’s model capabilities as sufficient.
AI Engineer World's Fair 2026 · 17:28
Affiliations reflect their AIE appearances, not necessarily current employment.
The joint discussion pairs deterministic grading with concrete security failures, including an authorization flaw caused by inconsistent name-versus-ID checks.
Affiliations reflect each recorded session, not necessarily current employment. The cybersecurity session is a joint presentation with Hugging Face's Thom Wolf.