How to Create an LLM Dataset | FineWeb Overview
FineWeb-Edu's 1.3T distilled tokens outperform the full 15T-token FineWeb dataset.
“even if a company that creates large language models publishes these models open source, they rarely publish their pre-training data as well because there is just very little incentive to do that”
A walkthrough of Hugging Face's FineWeb project documents every step of building open LLM pretraining datasets from Common Crawl, including the 15T-token FineWeb and the distilled 1.3T-token FineWeb-Edu. It matters because pretraining data recipes are rarely published, making this a useful open reference, though the dataset is a couple years old and the content is educational rather than a fresh announcement.