The Hallway Track
Engineering Insights

Lessons from Generating 12 Trillion Synthetic Tokens — Bogdan Gaza, DatologyAI

AI Engineer · Oct 02, 2026 · Engineering Insights

DatologyAI generated 12 trillion synthetic tokens for pre-training across web, math, and code domains.

“we have hit a "data wall." You need to spend exponentially more computing power and data to get models that only get linearly better.”

DatologyAI's CTO Bogdan Gaza shares engineering lessons from generating 12 trillion synthetic tokens for LLM pre-training, arguing the industry has hit a data wall where internet data is insufficient for frontier model capabilities. Their 'Beyond Web' synthetic data recipe demonstrates that smaller models trained on high-quality synthetic data can match larger models on benchmarks, with a 3B parameter model matching an 8B on NeMo TransSent. This is a high-signal engineering talk for anyone building or curating training datasets at scale.

synthetic-data pre-training data-curation scaling DatologyAI

Watch / read the original source →