DeepSWE: A Contamination-Resistant Coding Benchmark — James Shi, Datacurve
AI Engineer World's Fair 2026 · 17:34
AI training data and evaluation
Datacurve supplies expert-created coding data and evaluation infrastructure to foundation-model labs and AI developer-tool companies. Its offerings include reinforcement-learning environments, prebuilt datasets, supervised fine-tuning demonstrations and benchmarks. Customers use its data to train models for code editing, debugging, completion and design-to-code tasks. Its agent trajectories capture tool calls, checks and recoveries, while long-horizon tasks cover work stretching across hours or days—preserving the process of solving a problem as well as the answer.
Founded in 2024 by Serena Ge, its CEO, and Charley Lee, Datacurve participated in Y Combinator’s Winter 2024 batch. Its research includes DeepSWE, a benchmark introduced in 2026 for evaluating coding agents on extended software-engineering tasks. The August 2026 benchmark comprised 113 tasks across 91 repositories and five programming languages, with tasks written from scratch and handwritten verifiers that test software behavior rather than implementation details.
Datacurve collects data through Shipd, a platform where skilled software engineers compete in paid coding challenges. This contributor system supports its commercial work supplying complex coding datasets to AI labs. By October 2025, Datacurve had paid more than $1 million in contributor bounties and signed multimillion-dollar contracts with AI labs, establishing a business around expert engineering work as training material.
AI Engineer World's Fair 2026 · 17:34
Affiliations reflect their AIE appearances, not necessarily current employment.
Shi contrasts authored tasks, broader repository coverage, Git-history safeguards, and behavior-focused verifiers with weaknesses in benchmarks derived from pull requests.
Affiliations reflect each recorded session, not necessarily current employment.