AI research and reinforcement learning infrastructure
General Reasoning
General Reasoning builds infrastructure for training and evaluating language model agents. Its OpenReward platform lets researchers and developers discover community reinforcement learning environments, serve their own through API endpoints, and run them on managed infrastructure. Environment code stays in developers’ GitHub repositories, while training frameworks and sandbox providers remain independently selectable. At its 2026 public beta launch, the company reported a catalog of 330+ environments, including software engineering and command-line tasks.
Its co-founders include CEO Ross Taylor, president Chengxi Taylor and Kip Parker. Ross previously led Meta AI’s reasoning team, working on Llama 2 and Llama 3. General Reasoning’s Open Reward Standard connects language models to environments through familiar function calling, adding reward signals, episode termination, task definitions and reproducible task splits. These environments can run locally, on users’ infrastructure or through OpenReward’s optional managed hosting.
The company also offers BackSearch, a preview service for searching and retrieving historical web content from frozen news and SEC filing archives. It gives agent developers a repeatable information horizon for forecasting evaluations, financial backtesting and reinforcement learning. Users retrieve archived page text bounded by a chosen cutoff rather than today’s version. BackSearch charges for successful searches and fetches through a prepaid OpenReward balance, without a subscription.
1 talk
Newest first2 speakers at AIE
Affiliations reflect their AIE appearances, not necessarily current employment.
Messages from the stage
Simulation and optimization constraints
Chengxi Taylor outlines remaining challenges in realistic simulation, inference-time and training trade-offs, GPU utilization, and value-model bias.
Affiliations reflect each recorded session, not necessarily current employment.
