AI benchmarking and evaluation
Artificial Analysis
Artificial Analysis independently benchmarks AI models, agents, cloud infrastructure and chips to help developers and companies choose how to build and run AI applications. Its evaluations compare capabilities, quality, speed and price across language, image, video, speech and music models. Optima lets users build custom benchmarks from datasets, agent traces or coding workflows, then compare models on quality, cost per task and time per task. The company also offers advisory and custom benchmarking services for organizations adopting AI and technology providers.
Co-founders Micah Hill-Smith, the CEO, and George Cameron, the CPO, began the project in 2023, prompted by model-selection challenges while Hill-Smith built a legal AI assistant; it launched publicly in 2024. Its performance methodology measures customer-experienced inference, rather than theoretical hardware limits. Its research includes AA-Omniscience, which tests factual reliability across 6,000 questions spanning 42 topics in six domains. The benchmark rewards correct answers, penalizes hallucinations and leaves abstention unpenalized, distinguishing knowledge reliability from willingness to answer.
As of August 2026, the company reported benchmarking 500+ models, 100+ inference providers and 1,000+ endpoints, with more than one trillion evaluation tokens. By early 2026, Artificial Analysis had raised seed funding from Nat Friedman and Daniel Gross.
1 talk
Newest first2 speakers at AIE
Affiliations reflect their AIE appearances, not necessarily current employment.
Messages from the stage
Reasoning has a waiting-time cost
Cameron contrasts GPT-4.1’s 4.7-second median response with more than 40 seconds for o4-mini-high, making response time a concrete consideration alongside benchmark performance.
Affiliations reflect each recorded session, not necessarily current employment.
Company sources · checked 2026-08-28
- About | Artificial Analysis
- Articles | Artificial Analysis
- Artificial Analysis: Independent LLM Evals as a Service — with George Cameron and Micah-Hill Smith
- Artificial Analysis Benchmarking Methodology
- AA-Omniscience: Knowledge and Hallucination Benchmark
- Announcing Optima: create a custom benchmark for your use case
- Advisory & Custom Benchmarking Services | Artificial Analysis
