← All organizations

AI benchmarking and evaluation

Artificial Analysis

Artificial Analysis independently benchmarks AI models, agents, cloud infrastructure and chips to help developers and companies choose how to build and run AI applications. Its evaluations compare capabilities, quality, speed and price across language, image, video, speech and music models. Optima lets users build custom benchmarks from datasets, agent traces or coding workflows, then compare models on quality, cost per task and time per task. The company also offers advisory and custom benchmarking services for organizations adopting AI and technology providers.

Co-founders Micah Hill-Smith, the CEO, and George Cameron, the CPO, began the project in 2023, prompted by model-selection challenges while Hill-Smith built a legal AI assistant; it launched publicly in 2024. Its performance methodology measures customer-experienced inference, rather than theoretical hardware limits. Its research includes AA-Omniscience, which tests factual reliability across 6,000 questions spanning 42 topics in six domains. The benchmark rewards correct answers, penalizes hallucinations and leaves abstention unpenalized, distinguishing knowledge reliability from willingness to answer.

As of August 2026, the company reported benchmarking 500+ models, 100+ inference providers and 1,000+ endpoints, with more than one trillion evaluation tokens. By early 2026, Artificial Analysis had raised seed funding from Nat Friedman and Daniel Gross.

artificialanalysis.ai

1 talk

Newest first

2 speakers at AIE

Affiliations reflect their AIE appearances, not necessarily current employment.

Messages from the stage

Reasoning has a waiting-time cost

Cameron contrasts GPT-4.1’s 4.7-second median response with more than 40 seconds for o4-mini-high, making response time a concrete consideration alongside benchmark performance.

Affiliations reflect each recorded session, not necessarily current employment.

Company sources · checked 2026-08-28