Towards Alignment-Centric Paradigm: A Survey of Instruction Tuning in Large Language Models
2025/08/24 by Han, Xudong, Yang, Junjie, Wang, Tianyang +4 · 4 citations
#Computation and Language (cs.CL) #FOS: Computer and information sciences #I.2.6 #I.2.7
paper · doi:10.48550/arxiv.2508.17184
Abstract
Instruction tuning is a pivotal technique for aligning large language models (LLMs) with human intentions, safety constraints, and domain-specific requirements. This survey provides a comprehensive overview of the full pipeline, encompassing (i) data collection methodologies, (ii) full-parameter and parameter-efficient fine-tuning strategies, and (iii) evaluation protocols. We categorized data construction into three major paradigms: expert annotation, distillation from larger models, and self-improvement mechanisms, each offering distinct trade-offs between quality, scalability, and resource cost. Fine-tuning techniques range from conventional supervised training to lightweight approaches, such as low-rank adaptation (LoRA) and prefix tuning, with a focus on computational efficiency and model reusability. We further examine the challenges of evaluating faithfulness, utility, and safety across multilingual and multimodal scenarios, highlighting the emergence of domain-specific benchmarks in healthcare, legal, and financial applications. Finally, we discuss promising directions for automated data generation, adaptive optimization, and robust evaluation frameworks, arguing that a closer integration of data, algorithms, and human feedback is essential for advancing instruction-tuned LLMs. This survey aims to serve as a practical reference for researchers and practitioners seeking to design LLMs that are both effective and reliably aligned with human intentions.
Citations
- M-IFEval: Multilingual Instruction-Following Evaluation
- Fast Solvers for Discrete Diffusion Models: Theory and Applications of High-Order Algorithms
- Unifying Two Types of Scaling Laws from the Perspective of Conditional Kolmogorov Complexity
- Loops with involution and the Cayley-Dickson doubling process
- Right vs. Right: Can LLMs Make Tough Choices?
- Unveiling the Secret Recipe: A Guide For Supervised Fine-Tuning Small LLMs
- Sloth: scaling laws for LLM skills to predict multi-benchmark performance across families
- Densing Law of LLMs
- A Primer on Large Language Models and their Limitations
- Scaling Law for Language Models Training Considering Batch Size
- MDCure: A Scalable Pipeline for Multi-Document Instruction-Following
- Beware of Calibration Data for Pruning Large Language Models
- Compute-Constrained Data Selection
- Ethics Whitepaper: Whitepaper on Ethical Research into Large Language Models
- Optimizing Instruction Synthesis: Effective Exploration of Evolutionary Space with Tree Search
- Rethinking Conventional Wisdom in Machine Learning: From Generalization to Scaling
- Fine-tuning large language models for domain adaptation: Exploration of training strategies, scaling, model merging and synergistic capabilities
- Efficient Multi-task Prompt Tuning for Recommendation
- Awes, Laws, and Flaws From Today's LLM Research
- The Ultimate Guide to Fine-Tuning LLMs from Basics to Breakthroughs: An Exhaustive Review of Technologies, Research, Best Practices, Applied Research Challenges and Opportunities
- Risks, Causes, and Mitigations of Widespread Deployments of Large Language Models (LLMs): A Survey
- Neurosymbolic AI for Enhancing Instructability in Generative AI
- CMR Scaling Law: Predicting Critical Mixture Ratios for Continual Pre-training of Language Models
- LLMBox: A Comprehensive Library for Large Language Models
- Deconstructing The Ethics of Large Language Models from Long-standing Issues to New-emerging Dilemmas: A Survey
- Empirical Guidelines for Deploying LLMs onto Resource-constrained Edge Devices
- A Survey of Useful LLM Evaluation
- Stacking Your Transformers: A Closer Look at Model Growth for Efficient LLM Pre-Training
- Pruning as a Domain-specific LLM Extractor
- Quantifying the Capabilities of LLMs across Scale and Precision
- Apprentices to Research Assistants: Advancing Research with Large Language Models
- LLaMA-Excitor: General Instruction Tuning via Indirect Feature Interaction
- Large Language Models for Education: A Survey and Outlook
- Comprehensive Reassessment of Large-Scale Evaluation Outcomes in LLMs: A Multifaceted Statistical Approach
- Everybody Prune Now: Structured Pruning of LLMs with only Forward Passes
- A Closer Look at the Limitations of Instruction Tuning
- Faster and Lighter LLMs: A Survey on Current Challenges and Way Forward
- MM-LLMs: Recent Advances in MultiModal Large Language Models
- From Understanding to Utilization: A Survey on Explainability for Large Language Models
- L-TUNING: Synchronized Label Tuning for Prompt and Prefix in LLMs
- Demystifying Instruction Mixing for Fine-tuning Large Language Models
- A Rank Stabilization Scaling Factor for Fine-Tuning with LoRA
- Reflection-Tuning: Data Recycling Improves LLM Instruction-Tuning
- Ethical Reasoning over Moral Alignment: A Case and Framework for In-Context Ethical Policies in LLMs
- Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning
- Compresso: Structured Pruning with Collaborative Prompting Learns Compact Large Language Models
- ClusT3: Information Invariant Test-Time Training
- From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs after Instruction Tuning
- From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning
- LMTuner: An user-friendly and highly-integrable Training Framework for fine-tuning Large Language Models
- A Survey on Model Compression for Large Language Models
- A Comprehensive Overview of Large Language Models
- Friend or Foe? Exploring the Implications of Large Language Models on the Science System
- Inverse Scaling: When Bigger Isn't Better
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- Auto-GPT for Online Decision Making: Benchmarks and Additional Opinions
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Multi-Task Instruction Tuning of LLaMa for Specific Scenarios: A Preliminary Study on Writing Assistance
- Maybe Only 0.5% Data is Needed: A Preliminary Exploration of Low Training Data Instruction Tuning
- TinyStories: How Small Can Language Models Be and Still Speak Coherent English?
- Harnessing the Power of LLMs in Practice: A Survey on ChatGPT and Beyond
- A Bibliometric Review of Large Language Models Research from 2017 to 2023
- GPT-4 Technical Report
- Self-Instruct: Aligning Language Models with Self-Generated Instructions
- Affective Coherence Monitoring for Transformer-Based Language Models
- Understanding Scaling Laws for Recommendation Models
- Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
- Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context Learning
- Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Training Compute-Optimal Large Language Models
- Training language models to follow instructions with human feedback
- Red Teaming Language Models with Language Models
- Multitask Prompted Training Enables Zero-Shot Task Generalization
- TruthfulQA: Measuring How Models Mimic Human Falsehoods
- Finetuned Language Models Are Zero-Shot Learners
- BitFit: Simple Parameter-efficient Fine-tuning for Transformer-based Masked Language-models
- LoRA: Low-Rank Adaptation of Large Language Models
- Compacter: Efficient Low-Rank Hypercomplex Adapter Layers
- The Power of Scale for Parameter-Efficient Prompt Tuning
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
- Measuring Massive Multitask Language Understanding
- Meta-learning for Few-shot Natural Language Processing: A Survey
- AdapterHub: A Framework for Adapting Transformers
- Language Models are Few-Shot Learners
- XCOPA: A Multilingual Dataset for Causal Commonsense Reasoning
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
- Are Sixteen Heads Really Better than One?
- BERTScore: Evaluating Text Generation with BERT
- Parameter-Efficient Transfer Learning for NLP
- Deep reinforcement learning from human preferences
- SQuAD: 100,000+ Questions for Machine Comprehension of Text
- The interplay between domain specialization and model size
- RankAdaptor: Hierarchical Rank Allocation for Efficient Fine-Tuning Pruned LLMs via Performance Model
- WizardLM: Empowering large pre-trained language models to follow complex instructions
- Prefix-Tuning: Optimizing Continuous Prompts for Generation
- Prefix-Tuning: Optimizing Continuous Prompts for Generation
Cited by
Related