The Hallway Track
Engineering Insights

How to Create an LLM Dataset | FineWeb Overview

Hugging Face · Jun 02, 2026 · Engineering Insights

FineWeb-Edu's 1.3T distilled tokens outperform the full 15T-token FineWeb dataset.

“even if a company that creates large language models publishes these models open source, they rarely publish their pre-training data as well because there is just very little incentive to do that”

A walkthrough of Hugging Face's FineWeb project documents every step of building open LLM pretraining datasets from Common Crawl, including the 15T-token FineWeb and the distilled 1.3T-token FineWeb-Edu. It matters because pretraining data recipes are rarely published, making this a useful open reference, though the dataset is a couple years old and the content is educational rather than a fresh announcement.

datasets pretraining huggingface fineweb common-crawl

Watch / read the original source →