Independent publication on AI training data
Independent / State of Data
State of Data is Sean Cai’s independent publication and recurring analysis series about AI training-data markets, post-training and reinforcement-learning environments. It combines market commentary with practical writing for human-data companies, including a guide to building a human-data startup. Readers can explore how data collection, quality assurance and evaluation design affect the usefulness of training data.
Cai is the publication’s author and editorial voice. His analysis distinguishes state-based data, such as static records and work outputs, from process-based data: trajectories, rubrics and descriptions of reasoning workflows. He argues that realistic professional workflows and verifiable tasks matter more than benchmark difficulty alone when constructing environments for model improvement.
The publication is reader-supported through free and paid subscriptions on Substack. Subscribers can access its newsletter and archives, receive posts by email and participate in comments, with app-based reading also available. Its business coverage examines how buyer requirements, vendor quality and specialized datasets shape AI data markets, connecting technical production choices with the economics of supplying data.
1 talk
Newest first1 speaker at AIE
Affiliations reflect their AIE appearances, not necessarily current employment.
Messages from the stage
Testing capability beyond benchmark scores
Cai argues that benchmark gaming and differences in evaluation harnesses obscure real model capability. Finance-task examples illustrate the need for robust rubrics and deterministic verification.
Affiliations reflect each recorded session, not necessarily current employment.
