The Pile: An 800GB Dataset of Diverse Text for Language Modeling
2020/12/31 by Leo Gao, Gao, Leo, Stella Biderman +22 · 6 voices · 495 citations
Computer Science · #Algorithm #Artificial intelligence #Computer science #Natural Language Processing Techniques #Natural language processing #Pile #Speech Recognition and Synthesis #Topic Modeling #cs.CL
paper · pdf · doi:10.48550/arxiv.2101.00027
published in arXiv (Cornell University) (Cornell University)
arxiv created 2020/12/31 · openalex publication_date 2020/12/31 · arxiv updated 2021/01/05 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/30
Abstract
Recent work has demonstrated that increased training dataset diversity improves general cross-domain knowledge and downstream generalization capability for large-scale language models. With this in mind, we present the Pile: an 825 GiB English text corpus targeted at training large-scale language models. The Pile is constructed from 22 diverse high-quality subsets -- both existing and newly constructed -- many of which derive from academic or professional sources. Our evaluation of the untuned performance of GPT-2 and GPT-3 on the Pile shows that these models struggle on many of its components, such as academic writing. Conversely, models trained on the Pile improve significantly over both Raw CC and CC-100 on all components of the Pile, while improving performance on downstream evaluations. Through an in-depth exploratory analysis, we document potentially concerning aspects of the data for prospective users. We make publicly available the code used in its construction.
Citations
Cited by
- In-Run Data Shapley for Adam Optimizer
- Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations
- Scaling Interpretable Transformers with Parity Bottleneck Layers
- Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning
- Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions
- Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation
- Abstraction Induces the Brain Alignment of Language and Speech Models
- Persistent Sparse Autoencoders: Learning Feature Timescales in Language Models
- Every Component is a Lookup: Token Attribution and Composition from a Single Decomposition
- Gradient-Free Privacy Leakage in Federated Language Models through Selective Weight Tampering
- Scaling Point-in-Time Language Models
- Retrieval is Enough: Training-Free Interpretability with a Tool-Using Agent
- xHC: Expanded Hyper-Connections
- Building Fast, Evaluating Slow: Pipeline Choices Dominate Autointerpretability Score Variance
- Lost in Backpropagation: The LM Head is a Gradient Bottleneck
- Semantic Novelty Trajectories in 80,000 Books: A Cross-Corpus Embedding Analysis
- Soft Contamination Means Benchmarks Test Shallow Generalization
- Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework
- PHOTON: Hierarchical Autoregressive Modeling for Lightspeed and Memory-Efficient Language Generation
- Do LLMs Truly “Understand” when a Precedent Is Overruled?
- Continuous Autoregressive Language Models
- so much depends / upon / a whitespace: Why Whitespace Matters for Poets and LLMs
- Language Models are Injective and Hence Invertible
- Every Language Model Has a Forgery-Resistant Signature
- LLMs Can Get "Brain Rot": A Pilot Study on Twitter/X
- Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples
- Mind Your Tone: Investigating How Prompt Politeness Affects LLM Accuracy (short paper)
- Generalized Orders of Magnitude for Scalable, Parallel, High-Dynamic-Range Computation
- Anti-Regulatory AI: How "AI Safety" is Leveraged Against Regulatory Oversight
- Constructions are Revealed in Word Distributions
- Dynamic Chunking for End-to-End Hierarchical Sequence Modeling
- Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training
- Leaner Transformers: More Heads, Less Depth
- DataRater: Meta-Learned Dataset Curation
- The Heap: A Contamination-Free Multilingual Code Dataset for Evaluating Large Language Models
- Not All Data Are Unlearned Equally
- Register Always Matters: Analysis of LLM Pretraining Data Through the Lens of Language Variation
- Transformers without Normalization
- Shared Global and Local Geometry of Language Model Embeddings
- Provocations from the Humanities for Generative AI Research
- Enforcing Orderedness to Improve Feature Consistency
- Emergence of Phonemic, Syntactic, and Semantic Representations in Artificial Neural Networks
- LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling Laws
- Selecting Language Models for Social Science: Start Small, Start Open, and Validate
- Joint Optimization for Greedy Longest-match Tokenization
- Bits and Memories: Measuring Verbatim Extraction Across LLM Quantization
- Towards Robust Reinforcement Learning for Small-Scale Language Model Agents
- Hierarchical Grading in Large Language Models
- PANOPTICON: A PII-Based Assemblage of Naturalistic Output Tokens for Investigating Privacy Leakage Within LLM Context Window
- Gamayun's Path to Multilingual Mastery: Cost-Efficient Training of a 1.5B-Parameter LLM
- On Finding Inconsistencies in Documents
- Towards Benchmarking Privacy Vulnerabilities in Selective Forgetting with Large Language Models
- Perturb Your Data: Paraphrase-Guided Training Data Watermarking
- DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI
- Characterizing Mamba's Selective Memory using Auto-Encoders
- PerProb: Indirectly Evaluating Memorization in Large Language Models
- SASQ: Static Activation Scaling for Quantization-Aware Training in Large Language Models
- Superposition as Lossy Compression: Measure with Sparse Autoencoders and Connect to Adversarial Vulnerability
- On the Effectiveness of Membership Inference in Targeted Data Extraction from Large Language Models
- OxEnsemble: Fair Ensembles for Low-Data Classification
- Beyond the Black Box: Identifiable Interpretation and Control in Generative Models via Causal Minimality
- Scaling and Transferability of Annealing Strategies in Large Language Model Training
- Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
- SignRoundV2: Closing the Performance Gap in Extremely Low-Bit Post-Training Quantization for LLMs
- AdmTree: Compressing Lengthy Context with Adaptive Semantic Trees
- Nexus: Higher-Order Attention Mechanisms in Transformers
- Microbenchmarking NVIDIA's Blackwell Architecture: An in-depth Architectural Analysis
- TWEO: Transformers Without Extreme Outliers Enables FP8 Training And Quantization For Dummies
- Social Perceptions of English Spelling Variation on Twitter: A Comparative Analysis of Human and LLM Responses
- TrackList: Tracing Back Query Linguistic Diversity for Head and Tail Knowledge in Open Large Language Models
- Emergence and Localisation of Semantic Role Circuits in LLMs
- Memories Retrieved from Many Paths: A Multi-Prefix Framework for Robust Detection of Training Data Leakage in Large Language Models
- Copyright Detection in Large Language Models: An Ethical Approach to Generative AI Development
- Geometry of Decision Making in Language Models
- SmolKalam: Ensemble Quality-Filtered Translation at Scale for High Quality Arabic Post-Training Data
- FAST: Topology-Aware Frequency-Domain Distribution Matching for Coreset Selection
- Developmental Atlas of Attention Head Specialization: Spacing, Stranding, and the Capacity Tax of BPE Tokenization
- Dynamic Nested Hierarchies: Pioneering Self-Evolution in Machine Learning Architectures for Lifelong Intelligence
- Reproducibility Report: Test-Time Training on Nearest Neighbors for Large Language Models
- BhashaKritika: Building Synthetic Pretraining Data at Scale for Indic Languages
- ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference
- Know Your Limits: Entropy Estimation Modeling for Compression and Generalization
- ChEmREF: Evaluating Language Model Readiness for Chemical Emergency Response
- What can LLMs tell us about the mechanisms behind polarity illusions in humans? Experiments across model scales and training steps
- P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats
- How Do Data Owners Say No? A Case Study of Data Consent Mechanisms in Web-Scraped Vision-Language AI Training Datasets
- Rep2Text: Decoding Full Text from a Single LLM Token Representation
- Route Experts by Sequence, not by Token
- MOSS: Efficient and Accurate FP8 LLM Training with Microscaling and Automatic Scaling
- Beyond Redundancy: Diverse and Specialized Multi-Expert Sparse Autoencoder
- In-Context Learning Without Copying
- The Illusion of Certainty: Uncertainty Quantification for LLMs Fails under Ambiguity
- AyurParam: A State-of-the-Art Bilingual Language Model for Ayurveda
- Transformers as Intrinsic Optimizers: Forward Inference through the Energy Principle
- FlashEVA: Accelerating LLM inference via Efficient Attention
- TetraJet-v2: Accurate NVFP4 Training for Large Language Models with Oscillation Suppression and Outlier Control
- Atlas-Alignment: Making Interpretability Transferable Across Language Models
- Detecting Data Contamination in LLMs via In-Context Learning
- Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
- The Structure of Relation Decoding Linear Operators in Large Language Models
- RECAP: Reproducing Copyrighted Data from LLMs Training with an Agentic Pipeline
- Mixture-of-Depths Attention
- Why ‘open’ AI systems are actually closed, and why this matters
- KletterMix: Climbing Toward High-Quality German Pretraining Data - The Full Report
- On Surprising Effectiveness of Masking Updates in Adaptive Optimizers
- Beyond Length: Quantifying Long-Range Information for Long-Context LLM Pretraining Data
- A Survey on Unlearning in Large Language Models
- Language Model Behavioral Phases are Consistent Across Architecture, Training Data, and Scale
- Disaggregation Reveals Hidden Training Dynamics: The Case of Agreement Attraction
- Beyond Line-Level Filtering for the Pretraining Corpora of LLMs
- SpikeVox: Towards Energy-Efficient Speech Therapy Framework with Spike-driven Generative Language Models
- A Survey on LLM Mid-Training
- SeeDNorm: Self-Rescaled Dynamic Normalization
- Flight Delay Prediction via Cross-Modality Adaptation of Large Language Models and Aircraft Trajectory Representation
- Opening up ChatGPT: Tracking openness, transparency, and accountability in instruction-tuned text generators
- Addressing Corner Cases in Autonomous Driving: A World Model-based Approach with Mixture of Experts and LLMs
- FreeChunker: A Cross-Granularity Chunking Framework
- Context-level Language Modeling by Learning Predictive Context Embeddings
- An Empirical Study of Sample Selection Strategies for Large Language Model Repair
- Relative-Based Scaling Law for Neural Language Models
- Data-Centric Lessons To Improve Speech-Language Pretraining
- Machine Text Detectors are Membership Inference Attacks
- ToMMeR -- Efficient Entity Mention Detection from Large Language Models
- Blackbox Model Provenance via Palimpsestic Membership Inference
- Reasoning Language Model Inference Serving Unveiled: An Empirical Study
- ssToken: Self-modulated and Semantic-aware Token Selection for LLM Fine-tuning
- Protein Language Models: Is Scaling Necessary?
- Understanding and Improving Length Generalization in Hierarchical Sparse Attention Models
- Wisdom is Knowing What not to Say: Hallucination-Free LLMs Unlearning via Attention Shifting
- Improving Model Representation and Reducing KV Cache via Skip Connections with First Value Heads
- Circuit Insights: Towards Interpretability Beyond Activations
- To Infinity and Beyond: Tool-Use Unlocks Length Generalization in State Space Models
- Beyond Multi-Token Prediction: Pretraining LLMs with Future Summaries
- First Attentions Last: Better Exploiting First Attentions for Efficient Transformer Training
- Open WebUI: An Open, Extensible, and Usable Interface for AI Interaction
- Noise-Adaptive Layerwise Learning Rates: Accelerating Geometry-Aware Optimization for Deep Neural Network Training
- The German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models
- Tahakom LLM Guidelines and Recipes: From Pre-training Data to an Arabic LLM
- Layer-Aware Influence for Online Data Valuation Estimation
- Cautious Weight Decay
- Compressibility Measures Complexity: Minimum Description Length Meets Singular Learning Theory
- Influence Dynamics and Stagewise Data Attribution
- CoSPED: Consistent Soft Prompt Targeted Data Extraction and Defense
- DCP: Addressing Input Dynamism In Long-Context Training via Dynamic Context Parallelism
- PermLLM: Learnable Channel Permutation for N:M Sparse Large Language Models
- Translution: Unifying Self-attention and Convolution for Adaptive and Relative Modeling
- A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation Analysis
- Small edits, large models: How Wikipedia advocacy shapes LLM values
- On the Representations of Entities in Auto-regressive Large Language Models
- Value-State Gated Attention for Mitigating Extreme-Token Phenomena in Transformers
- Provable Training Data Identification for Large Language Models
- KORMo: Korean Open Reasoning Model for Everyone
- Multi-Condition Conformal Selection
- MeSH: Memory-as-State-Highways for Recursive Transformers
- Sunflower: A New Approach To Expanding Coverage of African Languages in Large Language Models
- End-to-End Test-Time Training for Long Context
- ReSSFormer: A Recursive Sparse Structured Transformer for Scalable and Long-Context Reasoning
- Crossing Domains without Labels: Distant Supervision for Term Extraction
- Mid-Training of Large Language Models: A Survey
- A Comparative Analysis of Contextual Representation Flow in State-Space and Transformer Architectures
- Memorization in Language Models through the Lens of Intrinsic Dimension
- Semantic Regexes: Auto-Interpreting LLM Features with a Structured Language
- VecInfer: Efficient LLM Inference with Low-Bit KV Cache via Outlier-Suppressed Vector Quantization
- (Token-Level) InfoRMIA: Stronger Membership Inference and Memorization Assessment for LLMs
- BLISS: A Lightweight Bilevel Influence Scoring Method for Data Selection in Language Model Pretraining
- Data Provenance Auditing of Fine-Tuned Large Language Models with a Text-Preserving Technique
- Boomerang Distillation Enables Zero-Shot Model Size Interpolation
- Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls
- What Scales in Cross-Entropy Scaling Law?
- Probing Geometry of Next Token Prediction Using Cumulant Expansion of the Softmax Entropy
- AbsTopK: Rethinking Sparse Autoencoders For Bidirectional Features
- Direct Token Optimization: A Self-contained Approach to Large Language Model Unlearning
- Evaluating Large Language Models Trained on Code
- HiSpec: Hierarchical Speculative Decoding for LLMs
- Scaling Laws Revisited: Modeling the Role of Data Quality in Language Model Pretraining
- Convergence and Divergence of Language Models under Different Random Seeds
- Bayesian Influence Functions for Hessian-Free Data Attribution
- MGen: Millions of Naturally Occurring Generics in Context
- Understanding the Mixture-of-Experts with Nadaraya-Watson Kernel
- MixtureVitae: Open Web-Scale Pretraining Dataset With High Quality Instruction and Reasoning Data Built from Permissive-First Text Sources
- Efficient Hyperparameter Tuning via Trajectory Invariance Principle
- GRACE-MoE: Grouping and Replication with Locality-Aware Routing for Efficient Distributed MoE Inference
- Specialization after Generalization: Towards Understanding Test-Time Training in Foundation Models
- Measuring Sparse Autoencoder Feature Sensitivity
- Pretraining Scaling Laws for Generative Evaluations of Language Models
- LLM Interpretability with Identifiable Temporal-Instantaneous Representation
- PonderLM-2: Pretraining LLM with Latent Thoughts in Continuous Space
- Tracing the Representation Geometry of Language Models from Pretraining to Post-training
- Efficient Fine-Grained GPU Performance Modeling for Distributed Deep Learning of LLM
- Induction Signatures Are Not Enough: A Matched-Compute Study of Load-Bearing Structure in In-Context Learning
- SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips
- Binary Autoencoder for Mechanistic Interpretability of Large Language Models
- Probability Signature: Bridging Data Semantics and Embedding Structure in Language Models
- Causal Understanding by LLMs: The Role of Uncertainty
- Measuring Coding Challenge Competence With APPS
- Mamba Modulation: On the Length Generalization of Mamba
- Harnessing the Potential of Optimizing Data Mixtures via Bayesian Domain Reweighting
- Explaining Data Mixing Scaling Laws
- Neural Scaling Universality: If Exponents Are Fixed, Time to Understand Coefficients
- How Can We Synthesize High-Quality Pretraining Data? A Systematic Study of Prompt Design, Generator Model, and Source Data
- Speculating LLMs' Chinese Training Data Pollution from Their Tokens
- ISACL: Internal State Analyzer for Copyrighted Training Data Leakage
- Memory in Large Language Models: Mechanisms, Evaluation and Evolution
- The Secret Agenda: LLMs Strategically Lie and Our Current Safety Tools Are Blind
- TiKMiX: Take Data Influence into Dynamic Mixture for Language Model Pre-training
- OpenGVL -- Benchmarking Visual Temporal Progress for Data Curation
- Clotho: Measuring Task-Specific Pre-Generation Test Adequacy for LLM Inputs
- MolPILE -- large-scale, diverse dataset for molecular representation learning
- Probabilistic Token Alignment for Large Language Model Fusion
- Similarity Field Theory: A Mathematical Framework for Intelligence
- SAM: A Mamba-2 State-Space Audio-Language Model
- Generalizability of Large Language Model-Based Agents: A Comprehensive Survey
- Knowledge-Driven Hallucination in Large Language Models: An Empirical Study on Process Modeling
- Synthetic bootstrapped pretraining
- Positional Encoding via Token-Aware Phase Attention
- Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets
- Truthful AI: Developing and governing AI that does not lie
- Building High-Quality Datasets for Portuguese LLMs: From Common Crawl Snapshots to Industrial-Grade Corpora
- Generative Data Refinement: Just Ask for Better Data
- Open-sci-ref-0.01: open and reproducible reference baselines for language model and dataset comparison
- Getting In Contract with Large Language Models -- An Agency Theory Perspective On Large Language Model Alignment
- Testing chatbots on the creation of encoders for audio conditioned image generation
- Towards EnergyGPT: A Large Language Model Specialized for the Energy Sector
- Are Targeted Data Poisoning Attacks as Effective as We Think?
- From Noise to Narrative: Tracing the Origins of Hallucinations in Transformers
- Effectively obtaining acoustic, visual and textual data from videos
- HoPE: Hyperbolic Rotary Positional Encoding for Stable Long-Range Dependency Modeling in Large Language Models
- Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
- How Small is Enough? Empirical Evidence of Quantized Small Language Models for Automated Program Repair
- Implicit Reasoning in Large Language Models: A Comprehensive Survey
- Natural Context Drift Undermines the Natural Language Understanding of Large Language Models
- Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling
- Router Upcycling: Leveraging Mixture-of-Routers in Mixture-of-Experts Upcycling
- Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders
- Backdoor Samples Detection Based on Perturbation Discrepancy Consistency in Pre-trained Language Models
- Distribution-Aware Feature Selection for SAEs
- OLMoASR: Open Models and Data for Training Robust Speech Recognition Models
- Data Cartography for Detecting Memorization Hotspots and Guiding Data Interventions in Generative Models
- Insights into User Interface Innovations from a Design Thinking Workshop at deRSE25
- Nemotron-CC-Math: A 133 Billion-Token-Scale High Quality Math Pretraining Dataset
- Quantization Meets dLLMs: A Systematic Study of Post-training Quantization for Diffusion LLMs
- In2x at WMT25 Translation Task
- AI Testing Should Account for Sophisticated Strategic Behaviour
- LinguaSafe: A Comprehensive Multilingual Safety Benchmark for Large Language Models
- DESIGNER: Design-Logic-Guided Multidisciplinary Data Synthesis for LLM Reasoning
- Signal and Noise: A Framework for Reducing Uncertainty in Language Model Evaluation
- SupraTok: Cross-Boundary Tokenization for Enhanced Language Model Performance
- ADMIRE-BayesOpt: Accelerated Data MIxture RE-weighting for Language Models with Bayesian Optimization
- SoK: Data Minimization in Machine Learning
- The Surprising Effectiveness of Membership Inference with Simple N-Gram Coverage
- Interpretable Reward Model via Sparse Autoencoder
- Improving Fine-Grained Emotion Detection in Text with BERT and GoEmotions
- Improving Fine-Grained Emotion Detection in Text with BERT and GoEmotions: An Experimental Study
- LLM Unlearning Without an Expert Curated Dataset
- ICM-Fusion: In-Context Meta-Optimized LoRA Fusion for Multi-Task Adaptation
- Guess or Recall? Training CNNs to Classify and Localize Memorization in LLMs
- Dynaword: From One-shot to Continuously Developed Datasets
- Trainable Dynamic Mask Sparse Attention
- MicroMix: Efficient Mixed-Precision Quantization with Microscaling Formats for Large Language Models
- The Art of Breaking Words: Rethinking Multilingual Tokenizer Design
- Large-Scale Diverse Synthesis for Mid-Training
- Adacc: An Adaptive Framework Unifying Compression and Activation Recomputation for LLM Training
- LeakSealer: A Semisupervised Defense for LLMs Against Prompt Injection and Leakage Attacks
- Embryology of a Language Model
- How Quantization Impacts Privacy Risk on LLMs for Code?
- On the Expressiveness of Softmax Attention: A Recurrent Neural Network Perspective
- A Single Direction of Truth: An Observer Model's Linear Residual Probe Exposes and Steers Contextual Hallucinations
- Data Mixing Agent: Learning to Re-weight Domains for Continual Pre-training
- Vibe Coding as a Reconfiguration of Intent Mediation in Software Development: Definition, Implications, and Research Agenda
- From Global to Local: A Scalable Benchmark for Local Posterior Sampling
- Cut the CARP: Fishing for zero-shot story evaluation
- The Blessing and Curse of Dimensionality in Safety Alignment
- A Survey on Generative Model Unlearning: Fundamentals, Taxonomy, Evaluation, and Future Direction
- AnalogFed: Privacy-Preserving Discovery of Analog Circuits at Scale with Federated Generative AI
- Rethinking Memorization Measures and their Implications in Large Language Models
- UniPool: A Globally Shared Expert Pool for Mixture-of-Experts
- Mangosteen: An Open Thai Corpus for Language Model Pretraining
- A Unifying Scheme for Extractive Content Selection Tasks
- Language Models Improve When Pretraining Data Matches Target Tasks
- GigaChat Family: Efficient Russian Language Modeling Through Mixture of Experts Architecture
- Dataset Ownership Verification for Pre-trained Masked Models
- On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention
- The value of books in the age of generative AI training data
- KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding
- Towards Secure and Private Language Models for Nuclear Power Plants
- Through the River: Understanding the Benefit of Schedule-Free Methods for Language Model Training
- Sub-Scaling Laws: On the Role of Data Density and Training Strategies in LLMs
- Scaling Laws for Optimal Data Mixtures
- If open source is to win, it must go public
- BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity
- Anthropomimetic Uncertainty: What Verbalized Uncertainty in Language Models is Missing
- Elite Polarization in European Parliamentary Speeches: a Novel Measurement Approach Using Large Language Models
- The Landscape of Memorization in LLMs: Mechanisms, Measurement, and Mitigation
- Differential Mamba
- Data Compressibility Quantifies LLM Memorization
- Meta-Learning Transformers to Improve In-Context Generalization
- A Survey of Pun Generation: Datasets, Evaluations and Methodologies
- Evaluating Federated Pre-Training: On the Reliability of Downstream Fine-Tuning and Intrinsic Evaluation
- A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation
- From Query to Explanation: Uni-RAG for Multi-Modal Retrieval-Augmented Learning in STEM
- Sign Spotting Disambiguation using Large Language Models
- Emergent Inabilities? Inverse Scaling Over the Course of Pretraining
- Read Quietly, Think Aloud: Decoupling Comprehension and Reasoning in LLMs
- Few-Shot Bot: Prompt-Based Learning for Dialogue Systems
- Understanding and Improving Length Generalization in Recurrent Models
- Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability
- PII Jailbreaking in LLMs via Activation Steering Reveals Personal Information Leakage
- CodableLLM: Automating Decompiled and Source Code Mapping for LLM Dataset Generation
- Low-Perplexity LLM-Generated Sequences and Where To Find Them
- SAFER: Probing Safety in Reward Models with Sparse Autoencoder
- Overcoming Long-Context Limitations of State-Space Models via Context-Dependent Sparse Attention
- A Hierarchical Neural Framework for Classification and its Explanation in Large Unstructured Legal Documents
- Private Memorization Editing: Turning Memorization into a Defense to Strengthen Data Privacy in Large Language Models
- Residual Matrix Transformers: Scaling the Size of the Residual Stream
- Language Modeling using LMUs: 10x Better Data Efficiency or Improved Scaling Compared to Transformers
- Grokking in LLM Pretraining? Monitor Memorization-to-Generalization without Test
- Is There a Case for Conversation Optimized Tokenizers in Large Language Models?
- Memba: Membrane-driven Parameter-Efficient Fine-Tuning for Mamba
- MEMOIR: Lifelong Model Editing with Minimal Overwrite and Informed Retention for LLMs
- CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models
- FaithfulSAE: Towards Capturing Faithful Features with Sparse Autoencoders without External Dataset Dependencies
- Reviving Your MNEME: Predicting The Side Effects of LLM Unlearning and Fine-Tuning via Sparse Model Diffing
- Long-Context Generalization with Sparse Attention
- Unlocking Post-hoc Dataset Inference with Synthetic Data
- LexiMark: Robust Watermarking via Lexical Substitutions to Enhance Membership Verification of an LLM's Textual Training Data
- Sampling from Your Language Model One Byte at a Time
- ROSAQ: Rotation-based Saliency-Aware Weight Quantization for Efficiently Compressing Large Language Models
- Beyond Frequency: The Role of Redundancy in Large Language Model Memorization
- Robust LLM Unlearning with MUDMAN: Meta-Unlearning with Disruption Masking And Normalization
- Overview of the NLPCC 2025 Shared Task: Gender Bias Mitigation Challenge
- An Efficient Compression of Deep Neural Network Checkpoints Based on Prediction and Context Modeling
- Infini-gram mini: Exact n-gram Search at the Internet Scale with FM-Index
- SOFT: Selective Data Obfuscation for Protecting LLM Fine-tuning against Membership Inference Attacks
- AWP: Activation-Aware Weight Pruning and Quantization with Projected Gradient Descent
- Towards Open Foundation Language Model and Corpus for Macedonian: A Low-Resource Language
- Hey, That's My Data! Label-Only Dataset Inference in Large Language Models
- Ultra-FineWeb: Efficient Data Filtering and Verification for High-Quality LLM Training Data
- Numerical Investigation of Sequence Modeling Theory using Controllable Memory Functions
- Sparse Autoencoders, Again?
- Line of Sight: On Linear Representations in VLLMs
- Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning
- From Real to Synthetic: Synthesizing Millions of Diversified and Complicated User Instructions with Attributed Grounding
- Rethinking LLM Advancement: Compute-Dependent and Independent Paths to Progress
- TokAlign: Efficient Vocabulary Adaptation via Token Alignment
- Intersectional Bias in Causal Language Models
- IndoSafety: Culturally Grounded Safety for LLMs in Indonesian Languages
- Causal Estimation of Tokenisation Bias
- IF-GUIDE: Influence Function-Guided Detoxification of LLMs
- The Inverse Scaling Effect of Pre-Trained Language Model Surprisal Is Not Due to Data Leakage
- zip2zip: Inference-Time Adaptive Tokenization via Online Compression
- Automatic Calibration for Membership Inference Attack on Large Language Models
- Not Every Token Needs Forgetting: Selective Unlearning to Limit Change in Utility in Large Language Model Unlearning
- Probing Neural Topology of Large Language Models
- Linear Representation Transferability Hypothesis: Leveraging Small Models to Steer Large Models
- Drop Dropout on Single-Epoch Language Model Pretraining
- Advancing Compositional Awareness in CLIP with Efficient Fine-Tuning
- PLAID: A Unified Data Model for Machine Learning on Heterogeneous Physics Simulations
- Emergent Abilities of Large Language Models under Continued Pretraining for Language Adaptation
- Chameleon: A Flexible Data-mixing Framework for Language Model Pretraining and Finetuning
- AC-ODM: Actor--Critic Online Data Mixing for Sample-Efficient LLM Pretraining
- Patterns and Mechanisms of Contrastive Activation Engineering
- Test-Time Training Done Right
- Critical Batch Size Revisited: A Simple Empirical Approach to Large-Batch Language Model Training
- Navigating the Latent Space Dynamics of Neural Models
- Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition
- Pre-Training Curriculum for Multi-Token Prediction in Language Models
- Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives
- Pretraining Language Models to Ponder in Continuous Space
- Sparsified State-Space Models are Efficient Highway Networks
- Efficient Large Language Model Inference with Neural Block Linearization
- Hardware-Efficient Attention for Fast Decoding
- Memorization or Interpolation ? Detecting LLM Memorization through Input Perturbation Analysis
- Incentivizing Inclusive Contributions in Model Sharing Markets
- Position: Foundation Models for Tabular Data within Systemic Contexts Need Grounding
- Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs
- Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression
- FinLoRA: Benchmarking LoRA Methods for Fine-Tuning LLMs on Financial Datasets
- GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining
- Efficient Data Selection at Scale via Influence Distillation
- Shifting AI Efficiency From Model-Centric to Data-Centric Compression
- Rethinking the Understanding Ability across LLMs through Mutual Information
- DB-KSVD: Scalable Alternating Optimization for Disentangling High-Dimensional Embedding Spaces
- An Outlook on the Opportunities and Challenges of Multi-Agent AI Systems
- Data Mixing Can Induce Phase Transitions in Knowledge Acquisition
- A Coreset Selection of Coreset Selection Literature: Introduction and Recent Advances
- Inference-Time Decomposition of Activations (ITDA): A Scalable Approach to Interpreting Large Language Models
- Learning What to Remember: Test-Time Training via Context Distillation
- The Rise of Parameter Specialization for Knowledge Storage in Large Language Models
- Mamba with Hierarchical Memory: Solving Representation Bottleneck in Long Sequence Modeling
- Small-to-Large Generalization: Data Influences Models Consistently Across Scale
- AdamS: Momentum Itself Can Be A Normalizer for LLM Pretraining and Post-training
- NQKV: A KV Cache Quantization Scheme Based on Normal Distribution Characteristics
- Understanding Differential Transformer Unchains Pretrained Self-Attentions
- Ensembling Sparse Autoencoders
- DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling
- Revealing Language Model Trajectories via Kullback-Leibler Divergence
- SUS backprop: linear backpropagation algorithm for long inputs in transformers
- Diagnosing our datasets: How does my language model learn clinical information?
- Likelihood Variance as Text Importance for Resampling Texts to Map Language Models
- Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference
- Studying the Role of Input-Neighbor Overlap in Retrieval-Augmented Language Models Training Efficiency
- Enhancing LLMs via High-Knowledge Data Selection
- Soft Prompts for Evaluation: Measuring Conditional Distance of Capabilities
- Unraveling Interwoven Roles of Large Language Models in Authorship Privacy: Obfuscation, Mimicking, and Verification
- Emergent Specialization: Rare Token Neurons in Language Models
- Enhancing Transformers Through Conditioned Embedded Tokens
- ChemPile: A 250GB Diverse and Curated Dataset for Chemical Foundation Models
- Competition between AI foundation models: dynamics and policy recommendations
- PANORAMA: A synthetic PII-laced dataset for studying sensitive data memorization in LLMs
- On Membership Inference Attacks in Knowledge Distillation
- Movable Antenna Enhanced Federated Fine-Tuning of Large Language Models via Hybrid Client Selection Optimization
- Feature Hedging: Correlated Features Break Narrow Sparse Autoencoders
- Optimal Control for Transformer Architectures: Enhancing Generalization, Robustness and Efficiency
- HessFormer: Hessians at Foundation Scale
- Superposition Yields Robust Neural Scaling
- Parallel Scaling Law for Language Models
- Mining Hidden Thoughts from Texts: Evaluating Continual Pretraining with Synthetic Data for LLM Reasoning
- Language Agents Mirror Human Causal Reasoning Biases. How Can We Help Them Think Like Scientists?
- Probability Consistency in Large Language Models: Theoretical Foundations Meet Empirical Discrepancies
- Large Language Models Meet Stance Detection: A Survey of Tasks, Methods, Applications, Challenges and Future Directions
- The Failure of Plagiarism Detection in Competitive Programming
- Efficient Unstructured Pruning of Mamba State-Space Models for Resource-Constrained Environments
- Large Language Models for Computer-Aided Design: A Survey
- Relative Overfitting and Accept-Reject Framework
- Learning Dynamics in Continual Pre-Training for Large Language Models
- Circuit Partitioning Using Large Language Models for Quantum Compilation and Simulations
- LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
- Sign-Aware Gated Sparse Autoencoders: Modeling Anticorrelated Features with Bi-Jump-ReLU Activations
- Language Modelling with Pixels
- (How) Learning Rates Regulate Catastrophic Overtraining
- Demystifying optimized prompts in language models
- Nothing from Something: Can a Language Model Discover 0?
- What Makes a Good Response? An Empirical Analysis of Quality in Qualitative Interviews
- Do Sparse Autoencoders Capture Concept Manifolds?
- Self-Ablating Transformers: More Interpretability, Less Sparsity
- Adaptive Loops and Memory in Transformers: Think Harder or Know More?
- In-Place Test-Time Training
- The Newton-Muon Optimizer
- The Last Fingerprint: How Markdown Training Shapes LLM Prose
- The Design Space of Tri-Modal Masked Diffusion Models
- Features have life history. And we should care
- ReCIT: Reconstructing Full Private Data from Gradient in Parameter-Efficient Fine-Tuning of Large Language Models
- Multimodal Large Language Models for Medicine: A Comprehensive Survey
- Disentangling MLP Neuron Weights in Vocabulary Space
- The State-Prediction Separation Hypothesis
- Where Does the Signal Live? A Web Data Recipe for Medical Encoder Pretraining
- Neuron Populations Exhibit Divergent Selectivity with Scale
- Loop, Think, & Generalize: Implicit Reasoning in Recurrent-Depth Transformers
- Learning is Forgetting: LLM Training As Lossy Compression
- PLDR-LLMs Reason At Self-Organized Criticality
- Information-Theoretic Storage Cost in Sentence Comprehension
- Procedural Pretraining: Warming Up Language Models with Abstract Data
- Llama-3.1-FoundationAI-SecurityLLM-Base-8B Technical Report
- From Evidence to Belief: A Bayesian Epistemology Approach to Language Models
- Modes of Sequence Models and Learning Coefficients
- Reverse-Engineering Model Editing on Language Models
- Scalable Frequency- and Length-Aware Subdocument Deduplication for Large Language Model Pretraining
- HMI: Hierarchical Knowledge Management for Efficient Multi-Tenant Inference in Pretrained Language Models
- When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars
- Datasheet for the Pile
- Safety Pretraining: Toward the Next Generation of Safe AI
- AdaParse: An Adaptive Parallel PDF Parsing and Resource Scaling Engine
- QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining
- Hardware-aligned Hierarchical Sparse Attention for Efficient Long-term Memory Access
- A Survey of Foundation Model-Powered Recommender Systems: From Feature-Based, Generative to Agentic Paradigms
- LongMamba: Enhancing Mamba's Long Context Capabilities via Training-Free Receptive Field Enlargement
- Compute-Optimal LLMs Provably Generalize Better With Scale
- Aria-MIDI: A Dataset of Piano MIDI Files for Symbolic Music Modeling
- RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents
- Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens
- ParaPO: Aligning Language Models to Reduce Verbatim Reproduction of Pre-training Data
- Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator
- Hypothetical Documents or Knowledge Leakage? Rethinking LLM-based Query Expansion
- A mean teacher algorithm for unlearning of language models
- Future Lens: Anticipating Subsequent Tokens from a Single Hidden State
- Nemotron-CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training
- Towards Understanding and Improving Refusal in Compressed Models via Mechanistic Interpretability
- On Linear Representations and Pretraining Data Frequency in Language Models
- Can Pre-training Indicators Reliably Predict Fine-tuning Outcomes of LLMs?
- A Perplexity and Menger Curvature-Based Approach for Similarity Evaluation of Large Language Models
- Nemotron-CrossThink: Scaling Self-Learning beyond Math Reasoning
- HELIOS: Adaptive Model And Early-Exit Selection for Efficient LLM Inference Serving
- Measuring LLM Novelty As The Frontier Of Original And High-Quality Output
- Exploration of Plan-Guided Summarization for Narrative Texts: the Case of Small Language Models
- Millions of States: Designing a Scalable MoE Architecture with RWKV-7 Meta-learner
- A Survey of Reasoning with Foundation Models: Concepts, Methodologies, and Outlook
- The KL3M Data Project: Copyright-Clean Training Resources for Large Language Models
- SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement
- Knowledge-Instruct: Effective Continual Pre-training from Limited Data using Instructions
- Lattice: Learning to Efficiently Compress the Memory
- Quantization Hurts Reasoning? An Empirical Study on Quantized Reasoning Models
- The Pile (dataset) [wikipedia]
- Connor Leahy [wikipedia]
- List of datasets for machine-learning research [wikipedia]
- List of large language models [wikipedia]
Discussions
- The Pile: An 800GB dataset of diverse text for language modeling (2020) [hn, 184 points, 70 comments]
- The entire AI boom is rooted in billionaire theft of creative works and privatization of the means to distribute them to the public. Billionaires stole the entire internet and want artists to starve [bsky, 31 points, 1 comments]
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling [hn, 13 points, 5 comments]
- arxiv.org/abs/2101.00027 [bsky, 2 points, 0 comments]
- Small caveat: I misunderstood arXiv's ToS when I wrote this paper. While a large portion of arXiv has an open license, the majority (last time I checked) does not. That shouldn't have a check under "a [bsky, 2 points, 0 comments]
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling [hn, 2 points, 1 comments]
Related