AI training data and evaluation
Surge AI
Surge AI supplies training data, human evaluation and reinforcement-learning environments to AI developers. Its dataset customers include Google, Anthropic and OpenAI. Its RLHF data captures human preferences, while supervised fine-tuning demonstrations teach models skills such as computer use, web navigation and reasoning. Developers can use its rubrics and verifiers to score model behavior, commission expert and multimodal data, or use ready-made datasets. Its language work covers more than 70 languages, incorporating cultural context alongside grammar and idiom.
Founded in 2020 by Edwin Chen, who remains CEO, Surge combines human expertise with research on model training and evaluation. It built OpenAI’s GSM8K dataset of 8,500 math problems to train and measure mathematical reasoning. Its CoreCraft environment, part of EnterpriseBench, simulates a customer-support organization for agent training. In its 2026 research, training in this environment improved performance on held-out tasks and external benchmarks, providing evidence that skills learned in realistic workflows can transfer beyond the training setting.
Surge generated more than $1 billion in revenue in 2024. In 2025, Chen estimated that it worked with more than one million contractors. The company remains independent and, Chen says, has taken no outside funding.
1 talk
Newest first1 speaker at AIE
Affiliations reflect their AIE appearances, not necessarily current employment.
Messages from the stage
Align prompts and verifiers in both directions
Heiner advocates complete two-way alignment between prompts and verifiers. Contradictory instructions and Unicode-based reward hacking illustrate the need for evaluation designs that withstand adversarial behavior.
Affiliations reflect each recorded session, not necessarily current employment.
