Hugging Face Journal Club: Scaling Laws for Pre-training & RL
New research uses chess as a testbed to derive scaling laws for post-training compute allocation
“In pre-training we had the chinchilla scaling laws uh which is roughly like uh uh you would allocate 20x the tokens per parameters of u of your model. But in a we didn't know like how much you would allocate.”
A paper reviewed at Hugging Face Journal Club investigates how pre-training compute allocation affects downstream post-training performance across SFT and RL stages, addressing an open question the field currently handles with heuristics. Using chess as a controlled three-stage training harness, the researchers attempt to derive principled scaling laws for post-training analogous to Chinchilla for pre-training. This is a meaningful research direction given industry-wide reliance on guesswork for post-training compute budgets, though the chess-domain constraint limits direct applicability.