← All organizations

Independent publication on AI training data

Independent / State of Data

State of Data is Sean Cai’s independent publication and recurring analysis series about AI training-data markets, post-training and reinforcement-learning environments. It combines market commentary with practical writing for human-data companies, including a guide to building a human-data startup. Readers can explore how data collection, quality assurance and evaluation design affect the usefulness of training data.

Cai is the publication’s author and editorial voice. His analysis distinguishes state-based data, such as static records and work outputs, from process-based data: trajectories, rubrics and descriptions of reasoning workflows. He argues that realistic professional workflows and verifiable tasks matter more than benchmark difficulty alone when constructing environments for model improvement.

The publication is reader-supported through free and paid subscriptions on Substack. Subscribers can access its newsletter and archives, receive posts by email and participate in comments, with app-based reading also available. Its business coverage examines how buyer requirements, vendor quality and specialized datasets shape AI data markets, connecting technical production choices with the economics of supplying data.

seanzcai.substack.com

1 talk

Newest first

1 speaker at AIE

Affiliations reflect their AIE appearances, not necessarily current employment.

Messages from the stage

Testing capability beyond benchmark scores

Cai argues that benchmark gaming and differences in evaluation harnesses obscure real model capability. Finance-task examples illustrate the need for robust rubrics and deterministic verification.

Affiliations reflect each recorded session, not necessarily current employment.

Company sources · checked 2026-08-28