← All organizations

AI cybersecurity research and evaluation

Arithmetic

Arithmetic explores how AI models can become more capable cybersecurity defenders through evaluation, training data, and benchmarks. Its stated research goal is to improve models’ ability to reason about vulnerabilities, addressing the gap between performing reconnaissance and making the logical leap needed to uncover an exploit.

The work presented in 2026 includes an access-control benchmark using blackbox environments built from vulnerabilities discovered by researchers. Deterministic grading makes these environments a way to test model capabilities against concrete security problems.

Arithmetic’s stated direction centers on improving the models themselves as a route to stronger defense. High-quality evaluations and data underpin that ambition: developing defenders capable of outperforming attackers, rather than treating today’s model capabilities as sufficient.

1 talk

Newest first

1 speaker at AIE

Affiliations reflect their AIE appearances, not necessarily current employment.

Messages from the stage

Testing authorization with deterministic grading

The joint discussion pairs deterministic grading with concrete security failures, including an authorization flaw caused by inconsistent name-versus-ID checks.

Affiliations reflect each recorded session, not necessarily current employment. The cybersecurity session is a joint presentation with Hugging Face's Thom Wolf.

Company sources · checked 2026-08-28