Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes
2023/05/03 by Cheng-Yu Hsieh, Chunliang Li, Hsieh, Cheng-Yu +18 · 5 voices · 106 citations
Computer Science · #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling #cs.AI #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.2305.02301
openalex publication_date 2023/05/03 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Deploying large language models (LLMs) is challenging because they are memory inefficient and compute-intensive for practical applications. In reaction, researchers train smaller task-specific models by either finetuning with human labels or distilling using LLM-generated labels. However, finetuning and distillation require large amounts of training data to achieve comparable performance to LLMs. We introduce Distilling step-by-step, a new mechanism that (a) trains smaller models that outperform LLMs, and (b) achieves so by leveraging less training data needed by finetuning or distillation. Our method extracts LLM rationales as additional supervision for training small models within a multi-task framework. We present three findings across 4 NLP benchmarks: First, compared to both finetuning and distillation, our mechanism achieves better performance with much fewer labeled/unlabeled training examples. Second, compared to few-shot prompted LLMs, we achieve better performance using substantially smaller model sizes. Third, we reduce both the model size and the amount of data required to outperform LLMs; our finetuned 770M T5 model outperforms the few-shot prompted 540B PaLM model using only 80% of available data on a benchmark, whereas standard finetuning the same T5 model struggles to match even by using 100% of the dataset. We release the code at: https://github.com/google-research/distilling-step-by-step .
Cited by
- From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search
- Benchmarking Fine-tuning and Retrieval Strategies for a Multimodal Language Model on the NRC Reactor Operator Licensing Examination
- Cross-Tokenizer On-Policy Distillation via Byte-Prefix Marginalization
- Quantize with Confidence? An Empirical Study of Quantization for Code Generation
- SOPD-SocialNav: Selective On-Policy Distillation for Vision-Language Social Navigation
- Contrastive On-Policy Distillation
- Rationale-Guided Knowledge Distillation for Cross-Lingual Stance Detection
- DAIS: Dependency-Aware Intermediate QA Supervision for Complex Reasoning
- When a Name Is Not a Name: A Benchmark Dataset and Distilled Reasoning for Culturally Entangled Bangla Homographs in Low-Resource LLMs
- Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models
- Training Continuous Chain of Thought Models: A Tale of Two Regimes
- Multi-Turn On-Policy Distillation with Prefix Replay
- Gold-Guided Programmatic Distillation for Financial Reasoning over Hybrid Tables and Text
- Program-as-Weights: A Programming Paradigm for Fuzzy Functions
- Procedural Knowledge Is Not Low-Rank: Why LoRA Fails to Internalize Multi-Step Procedures
- Answer-then-Edit: Reasoning Skeleton Editing for Anti-Distillation with Preserved Utility
- ShriNep@EEUCA 2026: RAKSHAK - Multi-Task DeBERTa with Rationale Distillation and Jigsaw-Augmented Training for Toxic Intent Classification
- It Takes 8 Tokens: Weak-to-Strong Off-Policy RL via Auxiliary Branches
- Embarrassingly Simple Self-Distillation Improves Code Generation
- Thinking Without Words: Efficient Latent Reasoning with Abstract Chain-of-Thought
- EditLord: Learning Code Transformation Rules for Code Editing
- Beyond Semantics: The Unreasonable Effectiveness of Reasonless Intermediate Tokens
- Position: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!
- Reanalysis-based Global Radiative Response to Sea Surface Temperature Patterns: Evaluating the Ai2 Climate Emulator
- OmniOPD: Logit-Free On-Policy Distillation via Speculative Verification
- Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
- CONSISTRE: A Unified Consistency-Aware Framework for Document-Level Relation Extraction with Large Language Models
- WCM: World-Cognition Model for Generalizable Human-Robot Interaction
- CoTinyVLA: Chain-of-Thought Distillation for a Sub-Billion-Parameter Vision-Language-Action Model
- LIME:Accelerating Collaborative Lossless LLM Inference on Memory-Constrained Edge Devices
- Distilling the Essence: Efficient Reasoning Distillation via Sequence Truncation
- Where Did This Sentence Come From? Tracing Provenance in LLM Reasoning Distillation
- BRIDGE: Budget-aware Reasoning via Intermediate Distillation with Guided Examples
- Reason2Decide: Rationale-Driven Multi-Task Learning
- Rethinking Knowledge Distillation in Collaborative Machine Learning: Memory, Knowledge, and Their Interactions
- Conscious Data Contribution via Community-Driven Chain-of-Thought Distillation
- Knowledge Distillation with Structured Chain-of-Thought for Text-to-SQL
- Explainable Ethical Assessment on Human Behaviors by Generating Conflicting Social Norms
- Made-in China, Thinking in America:U.S. Values Persist in Chinese LLMs
- Instruction-Tuning Open-Weight Language Models for BPMN Model Generation
- Reverse Thinking Enhances Missing Information Detection in Large Language Models
- Network of Theseus (like the ship)
- In-Context Distillation with Self-Consistency Cascades: A Simple, Training-Free Way to Reduce LLM Agent Costs
- Towards Edge General Intelligence: Knowledge Distillation for Mobile Agentic AI
- E3-Pruner: Towards Efficient, Economical, and Effective Layer Pruning for Large Language Models
- Self-Correction Distillation for Structured Data Question Answering
- Selecting Auxiliary Data via Neural Tangent Kernels for Low-Resource Domains
- CoT-X: An Adaptive Framework for Cross-Model Chain-of-Thought Transfer and Optimization
- Iterative Layer-wise Distillation for Efficient Compression of Large Language Models
- Batch Prompting Suppresses Overthinking Reasoning Under Constraint: How Batch Prompting Suppresses Overthinking in Reasoning Models
- ASARL: Autonomous Social-Aware Relevance Learning for QQ Search
- The Scaling Properties of Implicit Deductive Reasoning in Transformers
- GroupRAG: Cognitively Inspired Group-Aware Retrieval and Reasoning via Knowledge-Driven Problem Structuring
- The Kinetics of Reasoning: How Chain-of-Thought Shapes Learning in Transformers?
- Revisiting Knowledge Distillation: The Hidden Role of Dataset Size
- Test-time Verification via Optimal Transport: Coverage, ROC, & Sub-optimality
- Online In-Context Distillation for Low-Resource Vision Language Models
- Leave It to the Experts: Detecting Knowledge Distillation via MoE Expert Signatures
- Position: Require Frontier AI Labs To Release Small "Analog" Models
- Big Reasoning with Small Models: Instruction Retrieval at Inference Time
- Putting on the Thinking Hats: A Survey on Chain of Thought Fine-tuning from the Perspective of Human Reasoning Mechanism
- Stratos: An End-to-End Distillation Pipeline for Customized LLMs under Distributed Cloud Environments
- CPR: Mitigating Large Language Model Hallucinations with Curative Prompt Refinement
- Multi-stage Prompt Refinement for Mitigating Hallucinations in Large Language Models
- LLM-Oriented Token-Adaptive Knowledge Distillation
- STEPER: Step-wise Knowledge Distillation for Enhancing Reasoning Ability in Multi-Step Retrieval-Augmented Language Models
- Self-Filtered Distillation with LLMs-generated Trust Indicators for Reliable Patent Classification
- Distilling Reasoning into Student LLMs: Local Naturalness for Selecting Teacher Data
- Knowledge Graph-Guided Multi-Agent Distillation for Reliable Industrial Question Answering with Datasets
- Fine-tuning with RAG for Improving LLM Learning of New Skills
- Rearchitecting Datacenter Lifecycle for AI: A TCO-Driven Framework
- ORPO-Distill: Mixed-Policy Preference Optimization for Cross-Architecture LLM Distillation
- The Hidden Costs of Translation Accuracy: Distillation, Quantization, and Environmental Impact
- Evaluating Program Semantics Reasoning with Type Inference in System F
- Towards Efficient CoT Distillation: Self-Guided Rationale Selector for Better Performance with Fewer Rationales
- When Does Reasoning Matter? A Controlled Study of Reasoning's Contribution to Model Performance
- Rethinking Large Language Model Distillation: A Constrained Markov Decision Process Perspective
- Interactive Recommendation Agent with Active User Commands
- LogReasoner: Empowering LLMs with Expert-like Coarse-to-Fine Reasoning for Automated Log Analysis
- Teaching Audio Models to Reason: A Unified Framework for Source- and Layer-wise Distillation
- Brittleness and Promise: Knowledge Graph Based Reward Modeling for Diagnostic Reasoning
- When Big Models Train Small Ones: Label-Free Model Parity Alignment for Efficient Visual Question Answering using Small VLMs
- FastMTP: Accelerating LLM Inference with Enhanced Multi-Token Prediction
- Tree-OPO: Off-policy Monte Carlo Tree-Guided Advantage Optimization for Multistep Reasoning
- Learning from Diverse Reasoning Paths with Routing and Collaboration
- Foundation Models for Autonomous Driving Perception: A Survey Through Core Capabilities
- Transparency of medical artificial intelligence systems
- Performative Thinking? The Brittle Correlation Between CoT Length and Problem Complexity
- GeoAnalystBench: A GeoAI benchmark for assessing large language models for spatial analysis workflow and code generation
- Mitigating Spurious Correlations Between Question and Answer via Chain-of-Thought Correctness Perception Distillation
- Less is More Tokens: Efficient Math Reasoning via Difficulty-Aware Chain-of-Thought Distillation
- A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
- Prompting Strategies for Language Model-Based Item Generation in K-12 Education: Bridging the Gap Between Small and Large Language Models
- Utilizing Training Data to Improve LLM Reasoning for Tabular Understanding
- LLMind 2.0: Distributed IoT Automation with Natural Language M2M Communication and Lightweight LLM Agents
- ReaLM: Reflection-Enhanced Autonomous Reasoning with Small Language Models
- Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information
- Towards Efficient and Practical GPU Multitasking in the Era of LLM
- TeamMedAgents: Pareto-Efficient Multi-Agent Medical Reasoning Through Teamwork Theory
- Arce: Augmented Roberta with Contextualized Elucidations for Ner in Automated Rule Checking
- EvolvR: Self-Evolving Pairwise Reasoning for Story Evaluation to Enhance Generation
- Decision-Making with Deliberation: Meta-reviewing as a Document-grounded Dialogue
- FedCoT: Communication-Efficient Federated Reasoning Enhancement for Large Language Models
- A Rose by Any Other Name Would Smell as Sweet: Categorical Homotopy Theory for Large Language Models
- Boosting Visual Knowledge-Intensive Training for LVLMs Through Causality-Driven Visual Object Completion
- VFLAIR-LLM: A Comprehensive Framework and Benchmark for Split Learning of LLMs
Discussions
- Distilling Step-by-Step Outperforming Larger Language Models with Less Training [hn, 153 points, 34 comments]
- “our 770M T5 model outperforms the 540B PaLM model using only 80% of available data” (h/t AK) https://arxiv.org/abs/2305.02301 [bsky, 6 points, 0 comments]
- Excited to introduce Distilling Step-by-Step! ⚗️🪜 A simple mechanism to train small task-specific models to outperform [lemmy, 3 points, 0 comments]
- From: https://arxiv.org/abs/2305.02301 [bsky, 2 points, 0 comments]
- Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes arxiv.org/abs/2305.02301 [bsky, 0 points, 0 comments]
Related