Lessons from Generating 12 Trillion Synthetic Tokens — Bogdan Gaza, DatologyAI
DatologyAI generated 12 trillion synthetic tokens for pre-training across web, math, and code domains.
“we have hit a "data wall." You need to spend exponentially more computing power and data to get models that only get linearly better.”
DatologyAI's CTO Bogdan Gaza shares engineering lessons from generating 12 trillion synthetic tokens for LLM pre-training, arguing the industry has hit a data wall where internet data is insufficient for frontier model capabilities. Their 'Beyond Web' synthetic data recipe demonstrates that smaller models trained on high-quality synthetic data can match larger models on benchmarks, with a 3B parameter model matching an 8B on NeMo TransSent. This is a high-signal engineering talk for anyone building or curating training datasets at scale.