The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text
2025/06/05 by Nikhil Kandpal, Kandpal, Nikhil, Brian Lester +51 · 20 voices · 8 citations
Computer Science · Medicine · #Topic Modeling #Hate Speech and Cyberbullying Detection #Artificial Intelligence in Healthcare and Education
paper · pdf · doi:10.48550/arxiv.2506.05209
Abstract
Large language models (LLMs) are typically trained on enormous quantities of unlicensed text, a practice that has led to scrutiny due to possible intellectual property infringement and ethical concerns. Training LLMs on openly licensed text presents a first step towards addressing these issues, but prior data collection efforts have yielded datasets too small or low-quality to produce performant LLMs. To address this gap, we collect, curate, and release the Common Pile v0.1, an eight terabyte collection of openly licensed text designed for LLM pretraining. The Common Pile comprises content from 30 sources that span diverse domains including research papers, code, books, encyclopedias, educational materials, audio transcripts, and more. Crucially, we validate our efforts by training two 7 billion parameter LLMs on text from the Common Pile: Comma v0.1-1T and Comma v0.1-2T, trained on 1 and 2 trillion tokens respectively. Both models attain competitive performance to LLMs trained on unlicensed text with similar computational budgets, such as Llama 1 and 2 7B. In addition to releasing the Common Pile v0.1 itself, we also release the code used in its creation as well as the training mixture and checkpoints for the Comma v0.1 models.
Cited by
Discussions
- The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text [hn, 68 points, 17 comments]
- For more, check out... Paper: arxiv.org/abs/2506.05209 Artifacts: huggingface.co/common-pile GitHub: github.com/r-three/comm... EleutherAI's blog post: huggingface.co/blog/stellaa... Coverage in @wash [bsky, 5 points, 2 comments]
- @eleutherai.bsky.social 👏 arxiv.org/abs/2506.05209 [bsky, 4 points, 0 comments]
- The Common Pile v0.1: An 8TB dataset of public domain and openly licensed text [hn, 4 points, 0 comments]
- The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text [hn, 4 points, 0 comments]
- Yeah, that’s unfortunately still not EIC (Explicit Informed Consent). There’s this cool recent work too: arxiv.org/abs/2506.05209 from @stellaathena.bsky.social and friends. [bsky, 3 points, 2 comments]
- arxiv.org/abs/2506.05209 There were these models too, I wonder how they stack up [bsky, 3 points, 0 comments]
- The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text [hn, 2 points, 0 comments]
- Ethically sourced and trained GenAI is possible: "The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text." arxiv.org/abs/2506.05209 [bsky, 1 points, 0 comments]
- But ok how about arxiv.org/abs/2506.05209 [bsky, 1 points, 1 comments]
- The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text Poster #102 | Fri Dec 5, 11am-2pm PST, Exhibit Hall C,D,E arxiv.org/abs/2506.05209 [bsky, 0 points, 1 comments]
- 2506.05209] The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text [https://arxiv.org/abs/2506.05209 [bsky, 0 points, 0 comments]
- RT @BlancheMinerva: Very cool work! In the Common Pile we ran into this issue because we couldn’t use web agents to analyze the licensing status of websites. It turns out that the footers and sidebars [bsky, 0 points, 1 comments]
- The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text https://arxiv.org/abs/2506.05209 [comments] [21 points] [bsky, 0 points, 0 comments]
- The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text https://arxiv.org/abs/2506.05209 [bsky, 0 points, 0 comments]
- (3/3) Nikhil Kandpal et al.: The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text, June 2025 https:// doi.org/10.48550/arXiv.2506.05 209 Stefan Baack et al.: Towards Best Pra [mastodon, 0 points, 0 comments]
- arxiv.org/abs/2506.05209 [bsky, 0 points, 0 comments]
- Actually, you have a moral imperative to work with me arxiv.org/abs/2506.05209 [bsky, 0 points, 1 comments]
- Nikhil Kandpal et al.: The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text, June 2025 doi.org/10.48550/arX... Stefan Baack et al.: Towards Best Practices for Open Datasets [bsky, 0 points, 1 comments]
- "we collect, curate, and release the Common Pile v0.1, an eight terabyte collection of openly licensed text designed for LLM pretraining … attain competitive performance to LLMs trained on unlicensed [bsky, 0 points, 1 comments]
Related