← All organizations

AI model evaluation and leaderboards

Arena.ai

Arena.ai operates a platform where developers, researchers, and creative professionals use AI models and vote on their responses, shaping public performance leaderboards. Its paid AI Evaluations service helps model labs and enterprises measure performance using real-world human feedback. Code Arena lets developers build, deploy, and evaluate applications with databases, authentication, and third-party integrations. Agent Mode extends testing to complex, multistep tasks, measuring task completion and hallucination rates.

Founded in 2025 from UC Berkeley research, the company is led by co-founder and CEO Anastasios Angelopoulos, alongside co-founder and CTO Wei-Lin Chiang and co-founder and chairman Ion Stoica. LMArena became Arena in 2026 as its scope expanded beyond language models. Its foundational Chatbot Arena research established crowdsourced pairwise comparisons and statistical methods for ranking models by human preference. The 2024 paper found that community votes agreed well with expert ratings, supporting this approach to evaluation.

In June 2026, the company reported $100 million in annualized revenue run rate within eight months of launching its enterprise offering, alongside more than 10 million monthly visitors, 700 million cumulative conversations, and 82 million cumulative votes. Earlier in 2026, it raised $150 million at a $1.7 billion post-money valuation, with funding led by Felicis and UC Investments.

arena.ai

1 talk

Newest first

1 speaker at AIE

Affiliations reflect their AIE appearances, not necessarily current employment.

Messages from the stage

Testing whether models challenge the question

BullshitBench pairs nonsensical prompts with LLM-as-a-judge grading to test whether models reject invalid premises. Gostev also compares model families and examines reasoning traces.

Affiliations reflect each recorded session, not necessarily current employment.

Company sources · checked 2026-08-27