The Messy Reality of Scale: Synthetic Data and Pre-Training — Marah Abdin & Robert McHardy, poolside
Poolside releases open-weight Laguna M and Laguna XS models, shifting from enterprise-only distribution.
“at least at Pulsar we don't see it as a way to replace organic data”
Poolside (appearing as 'Pulsar' in captions) announced a strategic shift from enterprise-only to public open-weight model releases, dropping Laguna M and Laguna XS on Hugging Face. The talk details their synthetic data strategy: using it not as a replacement for organic data but as a complement to surface implicit rationale, planning, and structure, settling on 13% synthetic mix for pre-training. They also introduced 'auto mixer' for cheaper data blend sweeps before expensive training runs, offering a practical playbook for teams scaling pre-training data pipelines.