Bogdan Gaza is the co-founder and CTO of DatologyAI, where he applies his experience in large-scale infrastructure to building better AI training data. His work connects data research with the engineering needed to put it into use: selecting useful material, removing redundancy, generating synthetic text, and delivering datasets to model-training pipelines.
From search infrastructure to founding companies
Gaza’s earlier career included software engineering at Amazon and engineering management at Twitter. His earlier projects included automation and monitoring for PressLabs’ publishing infrastructure, along with a game-server prototype and AWS backend for BlastBusters.
At Twitter, Gaza was one of three core engineers behind SuperRoot, a service that unified access to the company’s separate search indexes. Recent tweets, archived tweets, and accounts lived in different systems; product teams previously had to combine results, remove duplicates, and handle pagination themselves. SuperRoot moved that work into shared infrastructure, simplifying product development while allowing the search team to change the underlying indexes. Gaza also contributed to the Omnisearch index-format work that enabled Twitter to store larger documents in its real-time indexes.
He subsequently co-founded Moonsense with Andrei Savu and served as CTO. The company worked on fraud detection and risk data, emphasizing understanding the person behind a device. Moonsense later concluded its operations.
Gaza initially had little appetite for immediately founding another startup. In 2023, he agreed to meet Ari Morcos to share his experience building natural-language and search infrastructure at Twitter. Within a few meetings, he decided to join Morcos and Matthew Leavitt in building DatologyAI. His data and backend infrastructure experience complemented their research backgrounds: data-curation methods needed to work beyond individual experiments.
Turning data research into training infrastructure
Gaza’s work at DatologyAI spans engineering leadership and collaborative research:
Scalable data curation and generation:DatologyAI manages the path from data in object storage to datasets consumed by training code. Gaza’s infrastructure work uses AWS services including S3, EMR, EC2, and Kubernetes through EKS to support product iteration and scale. In his account of generating roughly 12 trillion synthetic tokens across web, math, and code, he describes the team’s move from separate Slurm and Kubernetes systems to a unified Ray, KubeRay, and vLLM pipeline. At that scale, apparently minor overhead becomes consequential: batching S3 metadata fetches reduced a task taking roughly 9–11 days to about two hours. Partition sizing and checkpointing made GPU failures recoverable, while inference tuning yielded about 40% more throughput.
BeyondWeb: Gaza contributed leadership to BeyondWeb, DatologyAI’s synthetic pretraining-data framework. Its central mechanism is targeted document rephrasing: start with high-quality source material and generate varied, information-dense training text. The research found that diversity helps sustain learning over longer training runs, while simply extending documents offers limited gains. Pratyush Maini and Vineeth Dorna led the core research; Gaza contributed to the leadership supporting the project. The work connects a data-quality question—what makes synthetic text useful for learning—with the infrastructure needed to generate it at scale.
Specialized pretraining: Gaza provided research guidance and draft feedback for work on introducing domain data earlier in training, led by Christina Baek. Experiments with chemistry, music, and mathematical-proof datasets found that mixing specialized material into pretraining could improve subsequent domain performance and reduce forgetting during fine-tuning. This extends DatologyAI’s focus from choosing useful data to deciding when a model should encounter it.
Bogdan Gaza explains how DatologyAI seeds synthetic data with curated documents, then keeps trillion-token generation moving through metadata batching, recoverable jobs, coordinated scheduling and inference tuning.
BeyondWeb seeds generation with curated documents, then rephrases or restructures them to expand useful training material.
Metadata discovery can dominate startup at trillion-token scale. Batched S3 listing reduced DatologyAI’s estimated 9–11-day approach to about two hours of API processing.
Cross-cluster scheduling must coordinate a job’s CPU head and GPU workers, because separate pools can each have capacity without satisfying the whole job.
Inference parameter sweeps produced about 40% more throughput in the reported workload, while domain-specific data recipes remain an active experimental question.