The Pile: An 800GB Dataset of Diverse Text for Language Modeling
2020/12/31 by Leo Gao, Gao, Leo, Stella Biderman +22 · 6 voices · 271 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Speech Recognition and Synthesis
paper · pdf · doi:10.48550/arxiv.2101.00027
Abstract
Recent work has demonstrated that increased training dataset diversity improves general cross-domain knowledge and downstream generalization capability for large-scale language models. With this in mind, we present the Pile: an 825 GiB English text corpus targeted at training large-scale language models. The Pile is constructed from 22 diverse high-quality subsets -- both existing and newly constructed -- many of which derive from academic or professional sources. Our evaluation of the untuned performance of GPT-2 and GPT-3 on the Pile shows that these models struggle on many of its components, such as academic writing. Conversely, models trained on the Pile improve significantly over both Raw CC and CC-100 on all components of the Pile, while improving performance on downstream evaluations. Through an in-depth exploratory analysis, we document potentially concerning aspects of the data for prospective users. We make publicly available the code used in its construction.
Citations
Cited by
- In-Run Data Shapley for Adam Optimizer
- Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations
- Scaling Interpretable Transformers with Parity Bottleneck Layers
- Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning
- Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions
- Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation
- Abstraction Induces the Brain Alignment of Language and Speech Models
- Persistent Sparse Autoencoders: Learning Feature Timescales in Language Models
- Every Component is a Lookup: Token Attribution and Composition from a Single Decomposition
- Gradient-Free Privacy Leakage in Federated Language Models through Selective Weight Tampering
- Scaling Point-in-Time Language Models
- Retrieval is Enough: Training-Free Interpretability with a Tool-Using Agent
- xHC: Expanded Hyper-Connections
- Building Fast, Evaluating Slow: Pipeline Choices Dominate Autointerpretability Score Variance
- Lost in Backpropagation: The LM Head is a Gradient Bottleneck
- Semantic Novelty Trajectories in 80,000 Books: A Cross-Corpus Embedding Analysis
- Soft Contamination Means Benchmarks Test Shallow Generalization
- Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework
- PHOTON: Hierarchical Autoregressive Modeling for Lightspeed and Memory-Efficient Language Generation
- Do LLMs Truly “Understand” when a Precedent Is Overruled?
- Continuous Autoregressive Language Models
- so much depends / upon / a whitespace: Why Whitespace Matters for Poets and LLMs
- Language Models are Injective and Hence Invertible
- Every Language Model Has a Forgery-Resistant Signature
- LLMs Can Get "Brain Rot": A Pilot Study on Twitter/X
- Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples
- Mind Your Tone: Investigating How Prompt Politeness Affects LLM Accuracy (short paper)
- Generalized Orders of Magnitude for Scalable, Parallel, High-Dynamic-Range Computation
- Anti-Regulatory AI: How "AI Safety" is Leveraged Against Regulatory Oversight
- Constructions are Revealed in Word Distributions
- Dynamic Chunking for End-to-End Hierarchical Sequence Modeling
- Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training
- Leaner Transformers: More Heads, Less Depth
- DataRater: Meta-Learned Dataset Curation
- The Heap: A Contamination-Free Multilingual Code Dataset for Evaluating Large Language Models
- Not All Data Are Unlearned Equally
- Register Always Matters: Analysis of LLM Pretraining Data Through the Lens of Language Variation
- Transformers without Normalization
- Shared Global and Local Geometry of Language Model Embeddings
- Provocations from the Humanities for Generative AI Research
- Enforcing Orderedness to Improve Feature Consistency
- Emergence of Phonemic, Syntactic, and Semantic Representations in Artificial Neural Networks
- LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling Laws
- Selecting Language Models for Social Science: Start Small, Start Open, and Validate
- Joint Optimization for Greedy Longest-match Tokenization
- Bits and Memories: Measuring Verbatim Extraction Across LLM Quantization
- Towards Robust Reinforcement Learning for Small-Scale Language Model Agents
- Hierarchical Grading in Large Language Models
- PANOPTICON: A PII-Based Assemblage of Naturalistic Output Tokens for Investigating Privacy Leakage Within LLM Context Window
- Gamayun's Path to Multilingual Mastery: Cost-Efficient Training of a 1.5B-Parameter LLM
- On Finding Inconsistencies in Documents
- Towards Benchmarking Privacy Vulnerabilities in Selective Forgetting with Large Language Models
- Perturb Your Data: Paraphrase-Guided Training Data Watermarking
- DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI
- Characterizing Mamba's Selective Memory using Auto-Encoders
- PerProb: Indirectly Evaluating Memorization in Large Language Models
- SASQ: Static Activation Scaling for Quantization-Aware Training in Large Language Models
- Superposition as Lossy Compression: Measure with Sparse Autoencoders and Connect to Adversarial Vulnerability
- On the Effectiveness of Membership Inference in Targeted Data Extraction from Large Language Models
- OxEnsemble: Fair Ensembles for Low-Data Classification
- Beyond the Black Box: Identifiable Interpretation and Control in Generative Models via Causal Minimality
- Scaling and Transferability of Annealing Strategies in Large Language Model Training
- Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
- SignRoundV2: Closing the Performance Gap in Extremely Low-Bit Post-Training Quantization for LLMs
- AdmTree: Compressing Lengthy Context with Adaptive Semantic Trees
- Nexus: Higher-Order Attention Mechanisms in Transformers
- Microbenchmarking NVIDIA's Blackwell Architecture: An in-depth Architectural Analysis
- TWEO: Transformers Without Extreme Outliers Enables FP8 Training And Quantization For Dummies
- Social Perceptions of English Spelling Variation on Twitter: A Comparative Analysis of Human and LLM Responses
- TrackList: Tracing Back Query Linguistic Diversity for Head and Tail Knowledge in Open Large Language Models
- Emergence and Localisation of Semantic Role Circuits in LLMs
- Memories Retrieved from Many Paths: A Multi-Prefix Framework for Robust Detection of Training Data Leakage in Large Language Models
- Copyright Detection in Large Language Models: An Ethical Approach to Generative AI Development
- Geometry of Decision Making in Language Models
- SmolKalam: Ensemble Quality-Filtered Translation at Scale for High Quality Arabic Post-Training Data
- FAST: Topology-Aware Frequency-Domain Distribution Matching for Coreset Selection
- Developmental Atlas of Attention Head Specialization: Spacing, Stranding, and the Capacity Tax of BPE Tokenization
- Dynamic Nested Hierarchies: Pioneering Self-Evolution in Machine Learning Architectures for Lifelong Intelligence
- Reproducibility Report: Test-Time Training on Nearest Neighbors for Large Language Models
- BhashaKritika: Building Synthetic Pretraining Data at Scale for Indic Languages
- ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference
- Know Your Limits: Entropy Estimation Modeling for Compression and Generalization
- ChEmREF: Evaluating Language Model Readiness for Chemical Emergency Response
- What can LLMs tell us about the mechanisms behind polarity illusions in humans? Experiments across model scales and training steps
- P3-LLM: An Integrated NPU-PIM Accelerator for LLM Inference Using Hybrid Numerical Formats
- How Do Data Owners Say No? A Case Study of Data Consent Mechanisms in Web-Scraped Vision-Language AI Training Datasets
- Rep2Text: Decoding Full Text from a Single LLM Token Representation
- Route Experts by Sequence, not by Token
- MOSS: Efficient and Accurate FP8 LLM Training with Microscaling and Automatic Scaling
- Beyond Redundancy: Diverse and Specialized Multi-Expert Sparse Autoencoder
- In-Context Learning Without Copying
- The Illusion of Certainty: Uncertainty Quantification for LLMs Fails under Ambiguity
- AyurParam: A State-of-the-Art Bilingual Language Model for Ayurveda
- Transformers as Intrinsic Optimizers: Forward Inference through the Energy Principle
- FlashEVA: Accelerating LLM inference via Efficient Attention
- TetraJet-v2: Accurate NVFP4 Training for Large Language Models with Oscillation Suppression and Outlier Control
- Atlas-Alignment: Making Interpretability Transferable Across Language Models
- Detecting Data Contamination in LLMs via In-Context Learning
- Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
- The Structure of Relation Decoding Linear Operators in Large Language Models
- RECAP: Reproducing Copyrighted Data from LLMs Training with an Agentic Pipeline
- Mixture-of-Depths Attention
- Why ‘open’ AI systems are actually closed, and why this matters
- KletterMix: Climbing Toward High-Quality German Pretraining Data - The Full Report
- On Surprising Effectiveness of Masking Updates in Adaptive Optimizers
- Beyond Length: Quantifying Long-Range Information for Long-Context LLM Pretraining Data
- A Survey on Unlearning in Large Language Models
- Language Model Behavioral Phases are Consistent Across Architecture, Training Data, and Scale
- Disaggregation Reveals Hidden Training Dynamics: The Case of Agreement Attraction
- Beyond Line-Level Filtering for the Pretraining Corpora of LLMs
- SpikeVox: Towards Energy-Efficient Speech Therapy Framework with Spike-driven Generative Language Models
- A Survey on LLM Mid-Training
- SeeDNorm: Self-Rescaled Dynamic Normalization
- Flight Delay Prediction via Cross-Modality Adaptation of Large Language Models and Aircraft Trajectory Representation
- Opening up ChatGPT: Tracking openness, transparency, and accountability in instruction-tuned text generators
- Addressing Corner Cases in Autonomous Driving: A World Model-based Approach with Mixture of Experts and LLMs
- FreeChunker: A Cross-Granularity Chunking Framework
- Context-level Language Modeling by Learning Predictive Context Embeddings
- An Empirical Study of Sample Selection Strategies for Large Language Model Repair
- Relative-Based Scaling Law for Neural Language Models
- Data-Centric Lessons To Improve Speech-Language Pretraining
- Machine Text Detectors are Membership Inference Attacks
- ToMMeR -- Efficient Entity Mention Detection from Large Language Models
- Blackbox Model Provenance via Palimpsestic Membership Inference
- Reasoning Language Model Inference Serving Unveiled: An Empirical Study
- ssToken: Self-modulated and Semantic-aware Token Selection for LLM Fine-tuning
- Protein Language Models: Is Scaling Necessary?
- Understanding and Improving Length Generalization in Hierarchical Sparse Attention Models
- Wisdom is Knowing What not to Say: Hallucination-Free LLMs Unlearning via Attention Shifting
- Improving Model Representation and Reducing KV Cache via Skip Connections with First Value Heads
- Circuit Insights: Towards Interpretability Beyond Activations
- To Infinity and Beyond: Tool-Use Unlocks Length Generalization in State Space Models
- Beyond Multi-Token Prediction: Pretraining LLMs with Future Summaries
- First Attentions Last: Better Exploiting First Attentions for Efficient Transformer Training
- Open WebUI: An Open, Extensible, and Usable Interface for AI Interaction
- Noise-Adaptive Layerwise Learning Rates: Accelerating Geometry-Aware Optimization for Deep Neural Network Training
- The German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models
- Tahakom LLM Guidelines and Recipes: From Pre-training Data to an Arabic LLM
- Layer-Aware Influence for Online Data Valuation Estimation
- Cautious Weight Decay
- Compressibility Measures Complexity: Minimum Description Length Meets Singular Learning Theory
- Influence Dynamics and Stagewise Data Attribution
- CoSPED: Consistent Soft Prompt Targeted Data Extraction and Defense
- DCP: Addressing Input Dynamism In Long-Context Training via Dynamic Context Parallelism
- PermLLM: Learnable Channel Permutation for N:M Sparse Large Language Models
- Translution: Unifying Self-attention and Convolution for Adaptive and Relative Modeling
- A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation Analysis
- Small edits, large models: How Wikipedia advocacy shapes LLM values
- On the Representations of Entities in Auto-regressive Large Language Models
- Value-State Gated Attention for Mitigating Extreme-Token Phenomena in Transformers
- Provable Training Data Identification for Large Language Models
- KORMo: Korean Open Reasoning Model for Everyone
- Multi-Condition Conformal Selection
- MeSH: Memory-as-State-Highways for Recursive Transformers
- Sunflower: A New Approach To Expanding Coverage of African Languages in Large Language Models
- End-to-End Test-Time Training for Long Context
- ReSSFormer: A Recursive Sparse Structured Transformer for Scalable and Long-Context Reasoning
- Crossing Domains without Labels: Distant Supervision for Term Extraction
- Mid-Training of Large Language Models: A Survey
- A Comparative Analysis of Contextual Representation Flow in State-Space and Transformer Architectures
- Semantic Regexes: Auto-Interpreting LLM Features with a Structured Language
- VecInfer: Efficient LLM Inference with Low-Bit KV Cache via Outlier-Suppressed Vector Quantization
- (Token-Level) InfoRMIA: Stronger Membership Inference and Memorization Assessment for LLMs
- BLISS: A Lightweight Bilevel Influence Scoring Method for Data Selection in Language Model Pretraining
- Data Provenance Auditing of Fine-Tuned Large Language Models with a Text-Preserving Technique
- Boomerang Distillation Enables Zero-Shot Model Size Interpolation
- Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls
- What Scales in Cross-Entropy Scaling Law?
- Probing Geometry of Next Token Prediction Using Cumulant Expansion of the Softmax Entropy
- AbsTopK: Rethinking Sparse Autoencoders For Bidirectional Features
- Direct Token Optimization: A Self-contained Approach to Large Language Model Unlearning
- Evaluating Large Language Models Trained on Code
- HiSpec: Hierarchical Speculative Decoding for LLMs
- Scaling Laws Revisited: Modeling the Role of Data Quality in Language Model Pretraining
- Convergence and Divergence of Language Models under Different Random Seeds
- Bayesian Influence Functions for Hessian-Free Data Attribution
- MGen: Millions of Naturally Occurring Generics in Context
- Understanding the Mixture-of-Experts with Nadaraya-Watson Kernel
- MixtureVitae: Open Web-Scale Pretraining Dataset With High Quality Instruction and Reasoning Data Built from Permissive-First Text Sources
- Efficient Hyperparameter Tuning via Trajectory Invariance Principle
- GRACE-MoE: Grouping and Replication with Locality-Aware Routing for Efficient Distributed MoE Inference
- Specialization after Generalization: Towards Understanding Test-Time Training in Foundation Models
- Measuring Sparse Autoencoder Feature Sensitivity
- Pretraining Scaling Laws for Generative Evaluations of Language Models
- LLM Interpretability with Identifiable Temporal-Instantaneous Representation
- PonderLM-2: Pretraining LLM with Latent Thoughts in Continuous Space
- Tracing the Representation Geometry of Language Models from Pretraining to Post-training
- Efficient Fine-Grained GPU Performance Modeling for Distributed Deep Learning of LLM
- Induction Signatures Are Not Enough: A Matched-Compute Study of Load-Bearing Structure in In-Context Learning
- SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips
- Binary Autoencoder for Mechanistic Interpretability of Large Language Models
- Probability Signature: Bridging Data Semantics and Embedding Structure in Language Models
- Causal Understanding by LLMs: The Role of Uncertainty
- GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow
- Mamba Modulation: On the Length Generalization of Mamba
- Harnessing the Potential of Optimizing Data Mixtures via Bayesian Domain Reweighting
- Explaining Data Mixing Scaling Laws
- Neural Scaling Universality: If Exponents Are Fixed, Time to Understand Coefficients
- How Can We Synthesize High-Quality Pretraining Data? A Systematic Study of Prompt Design, Generator Model, and Source Data
- Speculating LLMs' Chinese Training Data Pollution from Their Tokens
- ISACL: Internal State Analyzer for Copyrighted Training Data Leakage
- Memory in Large Language Models: Mechanisms, Evaluation and Evolution
- The Secret Agenda: LLMs Strategically Lie and Our Current Safety Tools Are Blind
- TiKMiX: Take Data Influence into Dynamic Mixture for Language Model Pre-training
- OpenGVL -- Benchmarking Visual Temporal Progress for Data Curation
- Clotho: Measuring Task-Specific Pre-Generation Test Adequacy for LLM Inputs
- MolPILE -- large-scale, diverse dataset for molecular representation learning
- Probabilistic Token Alignment for Large Language Model Fusion
- Similarity Field Theory: A Mathematical Framework for Intelligence
- SAM: A Mamba-2 State-Space Audio-Language Model
- Generalizability of Large Language Model-Based Agents: A Comprehensive Survey
- Knowledge-Driven Hallucination in Large Language Models: An Empirical Study on Process Modeling
- Synthetic bootstrapped pretraining
- Positional Encoding via Token-Aware Phase Attention
- Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets
- Truthful AI: Developing and governing AI that does not lie
- Building High-Quality Datasets for Portuguese LLMs: From Common Crawl Snapshots to Industrial-Grade Corpora
- Generative Data Refinement: Just Ask for Better Data
- Open-sci-ref-0.01: open and reproducible reference baselines for language model and dataset comparison
- Getting In Contract with Large Language Models -- An Agency Theory Perspective On Large Language Model Alignment
- Testing chatbots on the creation of encoders for audio conditioned image generation
- Towards EnergyGPT: A Large Language Model Specialized for the Energy Sector
- Are Targeted Data Poisoning Attacks as Effective as We Think?
- From Noise to Narrative: Tracing the Origins of Hallucinations in Transformers
- Effectively obtaining acoustic, visual and textual data from videos
- HoPE: Hyperbolic Rotary Positional Encoding for Stable Long-Range Dependency Modeling in Large Language Models
- Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
- How Small is Enough? Empirical Evidence of Quantized Small Language Models for Automated Program Repair
- Implicit Reasoning in Large Language Models: A Comprehensive Survey
- Natural Context Drift Undermines the Natural Language Understanding of Large Language Models
- Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling
- Router Upcycling: Leveraging Mixture-of-Routers in Mixture-of-Experts Upcycling
- Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders
- Backdoor Samples Detection Based on Perturbation Discrepancy Consistency in Pre-trained Language Models
- Distribution-Aware Feature Selection for SAEs
- OLMoASR: Open Models and Data for Training Robust Speech Recognition Models
- Data Cartography for Detecting Memorization Hotspots and Guiding Data Interventions in Generative Models
- Insights into User Interface Innovations from a Design Thinking Workshop at deRSE25
- Nemotron-CC-Math: A 133 Billion-Token-Scale High Quality Math Pretraining Dataset
- Quantization Meets dLLMs: A Systematic Study of Post-training Quantization for Diffusion LLMs
- In2x at WMT25 Translation Task
- AI Testing Should Account for Sophisticated Strategic Behaviour
- LinguaSafe: A Comprehensive Multilingual Safety Benchmark for Large Language Models
- DESIGNER: Design-Logic-Guided Multidisciplinary Data Synthesis for LLM Reasoning
- Signal and Noise: A Framework for Reducing Uncertainty in Language Model Evaluation
- SupraTok: Cross-Boundary Tokenization for Enhanced Language Model Performance
- ADMIRE-BayesOpt: Accelerated Data MIxture RE-weighting for Language Models with Bayesian Optimization
- SoK: Data Minimization in Machine Learning
- The Surprising Effectiveness of Membership Inference with Simple N-Gram Coverage
- Interpretable Reward Model via Sparse Autoencoder
- Improving Fine-Grained Emotion Detection in Text with BERT and GoEmotions
- Improving Fine-Grained Emotion Detection in Text with BERT and GoEmotions: An Experimental Study
- LLM Unlearning Without an Expert Curated Dataset
- ICM-Fusion: In-Context Meta-Optimized LoRA Fusion for Multi-Task Adaptation
- Guess or Recall? Training CNNs to Classify and Localize Memorization in LLMs
- Dynaword: From One-shot to Continuously Developed Datasets
- Trainable Dynamic Mask Sparse Attention
- MicroMix: Efficient Mixed-Precision Quantization with Microscaling Formats for Large Language Models
- The Art of Breaking Words: Rethinking Multilingual Tokenizer Design
- Large-Scale Diverse Synthesis for Mid-Training
- Adacc: An Adaptive Framework Unifying Compression and Activation Recomputation for LLM Training
- LeakSealer: A Semisupervised Defense for LLMs Against Prompt Injection and Leakage Attacks
- Embryology of a Language Model
- How Quantization Impacts Privacy Risk on LLMs for Code?
- On the Expressiveness of Softmax Attention: A Recurrent Neural Network Perspective
- A Single Direction of Truth: An Observer Model's Linear Residual Probe Exposes and Steers Contextual Hallucinations
- Vibe Coding as a Reconfiguration of Intent Mediation in Software Development: Definition, Implications, and Research Agenda
- The Pile (dataset) [wikipedia]
- Connor Leahy [wikipedia]
- List of datasets for machine-learning research [wikipedia]
- List of large language models [wikipedia]
Discussions
- The Pile: An 800GB dataset of diverse text for language modeling (2020) [hn, 184 points, 70 comments]
- The entire AI boom is rooted in billionaire theft of creative works and privatization of the means to distribute them to the public. Billionaires stole the entire internet and want artists to starve [bsky, 31 points, 1 comments]
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling [hn, 13 points, 5 comments]
- arxiv.org/abs/2101.00027 [bsky, 2 points, 0 comments]
- Small caveat: I misunderstood arXiv's ToS when I wrote this paper. While a large portion of arXiv has an open license, the majority (last time I checked) does not. That shouldn't have a check under "a [bsky, 2 points, 0 comments]
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling [hn, 2 points, 1 comments]
Related