Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes
2023/05/03 by Cheng-Yu Hsieh, Chunliang Li, Hsieh, Cheng-Yu +18 · 5 voices · 193 citations
Computer Science · #Artificial intelligence #Benchmark (surveying) #Code (set theory) #Computer science #Distillation #Language model #Machine learning #Natural Language Processing Techniques #Programming language #Task (project management) #Text Readability and Simplification #Topic Modeling #Training set #cs.AI #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.2305.02301
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2023/05/03 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Deploying large language models (LLMs) is challenging because they are memory inefficient and compute-intensive for practical applications. In reaction, researchers train smaller task-specific models by either finetuning with human labels or distilling using LLM-generated labels. However, finetuning and distillation require large amounts of training data to achieve comparable performance to LLMs. We introduce Distilling step-by-step, a new mechanism that (a) trains smaller models that outperform LLMs, and (b) achieves so by leveraging less training data needed by finetuning or distillation. Our method extracts LLM rationales as additional supervision for training small models within a multi-task framework. We present three findings across 4 NLP benchmarks: First, compared to both finetuning and distillation, our mechanism achieves better performance with much fewer labeled/unlabeled training examples. Second, compared to few-shot prompted LLMs, we achieve better performance using substantially smaller model sizes. Third, we reduce both the model size and the amount of data required to outperform LLMs; our finetuned 770M T5 model outperforms the few-shot prompted 540B PaLM model using only 80% of available data on a benchmark, whereas standard finetuning the same T5 model struggles to match even by using 100% of the dataset. We release the code at: https://github.com/google-research/distilling-step-by-step .
Cited by
- From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search
- Benchmarking Fine-tuning and Retrieval Strategies for a Multimodal Language Model on the NRC Reactor Operator Licensing Examination
- Cross-Tokenizer On-Policy Distillation via Byte-Prefix Marginalization
- Quantize with Confidence? An Empirical Study of Quantization for Code Generation
- SOPD-SocialNav: Selective On-Policy Distillation for Vision-Language Social Navigation
- Contrastive On-Policy Distillation
- Rationale-Guided Knowledge Distillation for Cross-Lingual Stance Detection
- DAIS: Dependency-Aware Intermediate QA Supervision for Complex Reasoning
- When a Name Is Not a Name: A Benchmark Dataset and Distilled Reasoning for Culturally Entangled Bangla Homographs in Low-Resource LLMs
- Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models
- Training Continuous Chain of Thought Models: A Tale of Two Regimes
- Multi-Turn On-Policy Distillation with Prefix Replay
- Gold-Guided Programmatic Distillation for Financial Reasoning over Hybrid Tables and Text
- Program-as-Weights: A Programming Paradigm for Fuzzy Functions
- Procedural Knowledge Is Not Low-Rank: Why LoRA Fails to Internalize Multi-Step Procedures
- Answer-then-Edit: Reasoning Skeleton Editing for Anti-Distillation with Preserved Utility
- ShriNep@EEUCA 2026: RAKSHAK - Multi-Task DeBERTa with Rationale Distillation and Jigsaw-Augmented Training for Toxic Intent Classification
- It Takes 8 Tokens: Weak-to-Strong Off-Policy RL via Auxiliary Branches
- Embarrassingly Simple Self-Distillation Improves Code Generation
- Thinking Without Words: Efficient Latent Reasoning with Abstract Chain-of-Thought
- EditLord: Learning Code Transformation Rules for Code Editing
- Beyond Semantics: The Unreasonable Effectiveness of Reasonless Intermediate Tokens
- Position: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!
- Reanalysis-based Global Radiative Response to Sea Surface Temperature Patterns: Evaluating the Ai2 Climate Emulator
- OmniOPD: Logit-Free On-Policy Distillation via Speculative Verification
- Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
- CONSISTRE: A Unified Consistency-Aware Framework for Document-Level Relation Extraction with Large Language Models
- WCM: World-Cognition Model for Generalizable Human-Robot Interaction
- CoTinyVLA: Chain-of-Thought Distillation for a Sub-Billion-Parameter Vision-Language-Action Model
- Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices
- Distilling the Essence: Efficient Reasoning Distillation via Sequence Truncation
- Where Did This Sentence Come From? Tracing Provenance in LLM Reasoning Distillation
- BRIDGE: Budget-aware Reasoning via Intermediate Distillation with Guided Examples
- Reason2Decide: Rationale-Driven Multi-Task Learning
- Rethinking Knowledge Distillation in Collaborative Machine Learning: Memory, Knowledge, and Their Interactions
- Conscious Data Contribution via Community-Driven Chain-of-Thought Distillation
- Knowledge Distillation with Structured Chain-of-Thought for Text-to-SQL
- Explainable Ethical Assessment on Human Behaviors by Generating Conflicting Social Norms
- Made-in China, Thinking in America:U.S. Values Persist in Chinese LLMs
- Instruction-Tuning Open-Weight Language Models for BPMN Model Generation
- Reverse Thinking Enhances Missing Information Detection in Large Language Models
- Network of Theseus (like the ship)
- Inference-Time Distillation: Cost-Efficient Agents Without Fine-Tuning or Manual Prompt Engineering
- Towards Edge General Intelligence: Knowledge Distillation for Mobile Agentic AI
- E3-Pruner: Towards Efficient, Economical, and Effective Layer Pruning for Large Language Models
- Self-Correction Distillation for Structured Data Question Answering
- Selecting Auxiliary Data via Neural Tangent Kernels for Low-Resource Domains
- CoT-X: An Adaptive Framework for Cross-Model Chain-of-Thought Transfer and Optimization
- Iterative Layer-wise Distillation for Efficient Compression of Large Language Models
- Batch Prompting Suppresses Overthinking Reasoning Under Constraint: How Batch Prompting Suppresses Overthinking in Reasoning Models
- ASARL: Autonomous Social-Aware Relevance Learning for QQ Search
- The Scaling Properties of Implicit Deductive Reasoning in Transformers
- GroupRAG: Cognitively Inspired Group-Aware Retrieval and Reasoning via Knowledge-Driven Problem Structuring
- The Kinetics of Reasoning: How Chain-of-Thought Shapes Learning in Transformers?
- Revisiting Knowledge Distillation: The Hidden Role of Dataset Size
- Test-time Verification via Optimal Transport: Coverage, ROC, & Sub-optimality
- Online In-Context Distillation for Low-Resource Vision Language Models
- Leave It to the Experts: Detecting Knowledge Distillation via MoE Expert Signatures
- Position: Require Frontier AI Labs To Release Small "Analog" Models
- Big Reasoning with Small Models: Instruction Retrieval at Inference Time
- Putting on the Thinking Hats: A Survey on Chain of Thought Fine-tuning from the Perspective of Human Reasoning Mechanism
- Stratos: An End-to-End Distillation Pipeline for Customized LLMs under Distributed Cloud Environments
- CPR: Mitigating Large Language Model Hallucinations with Curative Prompt Refinement
- Multi-stage Prompt Refinement for Mitigating Hallucinations in Large Language Models
- LLM-Oriented Token-Adaptive Knowledge Distillation
- STEPER: Step-wise Knowledge Distillation for Enhancing Reasoning Ability in Multi-Step Retrieval-Augmented Language Models
- Self-Filtered Distillation with LLMs-generated Trust Indicators for Reliable Patent Classification
- Distilling Reasoning into Student LLMs: Local Naturalness for Selecting Teacher Data
- Knowledge Graph-Guided Multi-Agent Distillation for Reliable Industrial Question Answering with Datasets
- Fine-tuning with RAG for Improving LLM Learning of New Skills
- Rearchitecting Datacenter Lifecycle for AI: A TCO-Driven Framework
- ORPO-Distill: Mixed-Policy Preference Optimization for Cross-Architecture LLM Distillation
- The Hidden Costs of Translation Accuracy: Distillation, Quantization, and Environmental Impact
- Evaluating Program Semantics Reasoning with Type Inference in System F
- Towards Efficient CoT Distillation: Self-Guided Rationale Selector for Better Performance with Fewer Rationales
- When Does Reasoning Matter? A Controlled Study of Reasoning's Contribution to Model Performance
- Rethinking Large Language Model Distillation: A Constrained Markov Decision Process Perspective
- Interactive Recommendation Agent with Active User Commands
- LogReasoner: Empowering LLMs with Expert-like Coarse-to-Fine Reasoning for Automated Log Analysis
- Teaching Audio Models to Reason: A Unified Framework for Source- and Layer-wise Distillation
- Brittleness and Promise: Knowledge Graph Based Reward Modeling for Diagnostic Reasoning
- When Big Models Train Small Ones: Label-Free Model Parity Alignment for Efficient Visual Question Answering using Small VLMs
- FastMTP: Accelerating LLM Inference with Enhanced Multi-Token Prediction
- Tree-OPO: Off-policy Monte Carlo Tree-Guided Advantage Optimization for Multistep Reasoning
- Learning from Diverse Reasoning Paths with Routing and Collaboration
- Foundation Models for Autonomous Driving Perception: A Survey Through Core Capabilities
- Transparency of medical artificial intelligence systems
- Performative Thinking? The Brittle Correlation Between CoT Length and Problem Complexity
- GeoAnalystBench: A GeoAI benchmark for assessing large language models for spatial analysis workflow and code generation
- Mitigating Spurious Correlations Between Question and Answer via Chain-of-Thought Correctness Perception Distillation
- Less is More Tokens: Efficient Math Reasoning via Difficulty-Aware Chain-of-Thought Distillation
- A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
- Prompting Strategies for Language Model-Based Item Generation in K-12 Education: Bridging the Gap Between Small and Large Language Models
- Utilizing Training Data to Improve LLM Reasoning for Tabular Understanding
- LLMind 2.0: Distributed IoT Automation with Natural Language M2M Communication and Lightweight LLM Agents
- ReaLM: Reflection-Enhanced Autonomous Reasoning with Small Language Models
- Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information
- Towards Efficient and Practical GPU Multitasking in the Era of LLM
- TeamMedAgents: Pareto-Efficient Multi-Agent Medical Reasoning Through Teamwork Theory
- Arce: Augmented Roberta with Contextualized Elucidations for Ner in Automated Rule Checking
- EvolvR: Self-Evolving Pairwise Reasoning for Story Evaluation to Enhance Generation
- Decision-Making with Deliberation: Meta-reviewing as a Document-grounded Dialogue
- FedCoT: Communication-Efficient Federated Reasoning Enhancement for Large Language Models
- A Rose by Any Other Name Would Smell as Sweet: Categorical Homotopy Theory for Large Language Models
- Collaborative Inference and Learning between Edge SLMs and Cloud LLMs: A Survey of Algorithms, Execution, and Open Challenges
- Boosting Visual Knowledge-Intensive Training for LVLMs Through Causality-Driven Visual Object Completion
- VFLAIR-LLM: A Comprehensive Framework and Benchmark for Split Learning of LLMs
- Reasoning as a Resource: Optimizing Fast and Slow Thinking in Code Generation Models
- Fine-Tuning Code Language Models to Detect Cross-Language Bugs
- Basic Reading Distillation
- Auto: The AGI Compiler
- Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought
- BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity
- AbbIE: Autoregressive Block-Based Iterative Encoder for Efficient Sequence Modeling
- Agentic-R1: Distilled Dual-Strategy Reasoning
- MLlm-DR: Towards Explainable Depression Recognition with MultiModal Large Language Models
- TELLER: Dual-Path Iterative Preference Optimization for Table Entity Linking
- Lyria: A Genetic Algorithm-Driven Neuro-Symbolic Reasoning Framework for LLMs
- EarlyDx: An Admission-Anchored Benchmark for Open-Ended Generation of Evidence-Supported ED-Encounter Diagnoses
- Multimodal Mathematical Reasoning with Diverse Solving Perspective
- Information-Theoretic Framework for Understanding Modern Machine-Learning
- Complexity-aware fine-tuning
- A Survey on Model Extraction Attacks and Defenses for Large Language Models
- GPTailor: Large Language Model Pruning Through Layer Cutting and Stitching
- Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation
- Differentiation-Based Extraction of Proprietary Data from Fine-Tuned LLMs
- GenRecal: Generation after Recalibration from Large to Small Vision-Language Models
- MultiFinBen: Benchmarking Large Language Models for Multilingual and Multimodal Financial Application
- Being Strong Progressively! Enhancing Knowledge Distillation of Large Language Models through a Curriculum Learning Framework
- SIGMA: Refining Large Language Model Reasoning via Sibling-Guided Monte Carlo Augmentation
- Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning
- KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
- When LLMs Team Up: The Emergence of Collaborative Affective Computing
- Foresight: Adaptive Layer Reuse for Accelerated and High-Quality Text-to-Video Generation
- Common Inpainted Objects In-N-Out of Context
- Inter-Passage Verification for Multi-evidence Multi-answer QA
- DefenderBench: A Toolkit for Evaluating Language Agents in Cybersecurity Environments
- Enhancing Long-Chain Reasoning Distillation through Error-Aware Self-Reflection
- MemOS: An Operating System for Memory-Augmented Generation (MAG) in Large Language Models
- Research on Driving Scenario Technology Based on Multimodal Large Lauguage Model Optimization
- EvolveSearch: An Iterative Self-Evolving Search Agent
- Optimizing LLMs for Resource-Constrained Environments: A Survey of Model Compression Techniques
- Causal Distillation: Transferring Structured Explanations from Large to Compact Language Models
- Does Rationale Quality Matter? Enhancing Mental Disorder Detection via Selective Reasoning Distillation
- Online Knowledge Distillation with Reward Guidance
- Skip-Thinking: Chunk-wise Chain-of-Thought Distillation Enable Smaller Language Models to Reason Better and Faster
- μ-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts
- The Quest for Efficient Reasoning: A Data-Centric Benchmark to CoT Distillation
- LatentLLM: Attention-Aware Joint Tensor Compression
- Distilling LLM Agent into Small Models with Retrieval and Code Tools
- Two-way Evidence self-Alignment based Dual-Gated Reasoning Enhancement
- LLM-Powered AI Agent Systems and Their Applications in Industry
- Enhancing Large Language Models for Detecting Mental Manipulation via Annotation-Free Data Augmentation and Anti-Curriculum Distillation
- Polar Sparsity: High Throughput Batched LLM Inferencing with Scalable Contextual Sparsity
- ReEx-SQL: Reasoning with Execution-Aware Reinforcement Learning for Text-to-SQL
- A Token is Worth over 1,000 Tokens: Efficient Knowledge Distillation through Low-Rank Clone
- ExpertSteer: Intervening in LLMs through Expert Knowledge
- Recursive Question Understanding for Complex Question Answering over Heterogeneous Personal Data
- MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems
- AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction
- Distilling Reasoning Traces into Advisory Prompts for Software Engineering Tasks
- Relative Overfitting and Accept-Reject Framework
- CORE: Collaborative Reasoning via Cross Teaching
- Efficient Fine-Tuning of Quantized Models via Adaptive Rank and Bitwidth
- CAT-Q: Cost-efficient and Accurate Ternary Quantization for LLMs
- Decoupling Search from Reasoning: A Vendor-Agnostic Grounding Architecture for LLM Agents
- Promoting Critical Thinking With Domain-Specific Generative AI Provocations
- DOPD: Dual On-policy Distillation
- Scaling Laws for Task-Specific LLM Distillation
- AdaptMI: Adaptive Skill-based In-context Math Instruction for Small Language Models
- Think First, Diffuse Fast: Improving Diffusion Language Model Reasoning via Autoregressive Plan Conditioning
- CR-Seg: Attention-Guided and CoT-Enhanced Coarse-to-Refined Reasoning Segmentation
- Hierarchical Prompt-Domain Control and Learning for Resource-Constrained Agentic Language Models
- When Less is Enough: Efficient Inference via Collaborative Reasoning
- ReasonIR: Training Retrievers for Reasoning Tasks
- Memorization Dynamics in Knowledge Distillation for Language Models
- ToolTok: Tool Tokenization for Efficient and Generalizable GUI Agents
- The Frontier LLM Trap in Network Automation
- CausalOPD: First-Wrong-Step Supervision for Distilling Causal Chain Reasoning
- Not Every Divergence Should Be Suppressed: Counterfactual Recoverability in On-Policy Distillation
- Target Concrete Score Matching: A Holistic Framework for Discrete Diffusion
- Honey, I Shrunk the Language Model: Impact of Knowledge Distillation Methods on Performance and Explainability
- LLM-based Semantic Augmentation for Harmful Content Detection
- DistilQwen2.5: Industrial Practices of Training Distilled Open Lightweight Language Models
- Knowledge Distillation and Dataset Distillation of Large Language Models: Emerging Trends, Challenges, and Future Directions
- Collaborative Learning of On-Device Small Model and Cloud-Based Large Model: Advances and Future Directions
- StepReflect: Structured UI Transition Reflection for Mobile GUI Agents
- DynaPix: Can Vision-Language Models Identify the Exact Future?
- Woodpecker Distillation: Weak Models Diagnose Reasoning Bugs in Strong Models
- Mean-Field Dynamics of Chain-of-Thought Reasoning in Large Language Models
- Enhancing Reasoning Abilities of Small LLMs with Cognitive Alignment
- A Survey of Reasoning with Foundation Models: Concepts, Methodologies, and Outlook
- SD2: Self-Distilled Sparse Drafters
Discussions
- Distilling Step-by-Step Outperforming Larger Language Models with Less Training [hn, 153 points, 34 comments]
- “our 770M T5 model outperforms the 540B PaLM model using only 80% of available data” (h/t AK) https://arxiv.org/abs/2305.02301 [bsky, 6 points, 0 comments]
- Excited to introduce Distilling Step-by-Step! ⚗️🪜 A simple mechanism to train small task-specific models to outperform [lemmy, 3 points, 0 comments]
- From: https://arxiv.org/abs/2305.02301 [bsky, 2 points, 0 comments]
- Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes arxiv.org/abs/2305.02301 [bsky, 0 points, 0 comments]
Related