Large Language Model Alignment: A Survey
2023/09/26 by Tianhao Shen, Shen, Tianhao, Renren Jin +15 · 37 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Software Engineering Research #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2309.15025
openalex publication_date 2023/09/26 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Recent years have witnessed remarkable progress made in large language models (LLMs). Such advancements, while garnering significant attention, have concurrently elicited various concerns. The potential of these models is undeniably vast; however, they may yield texts that are imprecise, misleading, or even detrimental. Consequently, it becomes paramount to employ alignment techniques to ensure these models to exhibit behaviors consistent with human values. This survey endeavors to furnish an extensive exploration of alignment methodologies designed for LLMs, in conjunction with the extant capability research in this domain. Adopting the lens of AI alignment, we categorize the prevailing methods and emergent proposals for the alignment of LLMs into outer and inner alignment. We also probe into salient issues including the models' interpretability, and potential vulnerabilities to adversarial attacks. To assess LLM alignment, we present a wide variety of benchmarks and evaluation methodologies. After discussing the state of alignment research for LLMs, we finally cast a vision toward the future, contemplating the promising avenues of research that lie ahead. Our aspiration for this survey extends beyond merely spurring research interests in this realm. We also envision bridging the gap between the AI alignment research community and the researchers engrossed in the capability exploration of LLMs for both capable and safe LLMs.
Cited by
- Is Chain-of-Thought Really Not Explainability? Chain-of-Thought Can Be Faithful without Hint Verbalization
- Breaking Minds, Breaking Systems: Jailbreaking Large Language Models via Human-like Psychological Manipulation
- Towards Proactive Personalization through Profile Customization for Individual Users in Dialogues
- Invasive Context Engineering to Control Large Language Models
- Factors That Support Grounded Responses in LLM Conversations: A Rapid Review
- LLM Reinforcement in Context
- When Data is the Algorithm: A Systematic Study and Curation of Preference Optimization Datasets
- Convergence and Stability Analysis of Self-Consuming Generative Models with Heterogeneous Human Curation
- Revisiting Entropy in Reinforcement Learning for Large Reasoning Models
- Control Barrier Function for Aligning Large Language Models
- Inter-Agent Trust Models: A Comparative Study of Brief, Claim, Proof, Stake, Reputation and Constraint in Agentic Web Protocol Design-A2A, AP2, ERC-8004, and Beyond
- A Criminology of Machines
- Multi-Agent Debate Strategies: Survey, Taxonomy, and Challenges
- A word association network methodology for evaluating implicit biases in LLMs compared to humans
- Temporal Blindness in Multi-Turn LLM Agents: Misaligned Tool Use vs. Human Time Perception
- Counterfactual Reasoning for Steerable Pluralistic Value Alignment of Large Language Models
- Stop Reducing Responsibility in LLM-Powered Multi-Agent Systems to Local Alignment
- In-Distribution Steering: Balancing Control and Coherence in Language Model Generation
- On the Role of Preference Variance in Preference Optimization
- Contrastive Weak-to-strong Generalization
- RLRF: Competitive Search Agent Design via Reinforcement Learning from Ranker Feedback
- Toward Preference-aligned Large Language Models via Residual-based Model Steering
- Interactivity under generative AI: a systems-theoretical analysis of media transformation
- DecipherGuard: Understanding and Deciphering Jailbreak Prompts for a Safer Deployment of Intelligent Software Systems
- Can an Individual Manipulate the Collective Decisions of Multi-Agents?
- SMARTER: A Data-efficient Framework to Improve Toxicity Detection with Explanation via Self-augmenting Large Language Models
- TurnBack: A Geospatial Route Cognition Benchmark for Large Language Models through Reverse Route
- Getting In Contract with Large Language Models -- An Agency Theory Perspective On Large Language Model Alignment
- EPT Benchmark: Evaluation of Persian Trustworthiness in Large Language Models
- MEUV: Achieving Fine-Grained Capability Activation in Large Language Models via Mutually Exclusive Unlock Vectors
- Better Language Model-Based Judging Reward Modeling through Scaling Comprehension Boundaries
- A Comprehensive Evaluation framework of Alignment Techniques for LLMs
- Steering Towards Fairness: Mitigating Political Bias in LLMs
- A Survey on Training-free Alignment of Large Language Models
- Securing Educational LLMs: A Generalised Taxonomy of Attacks on LLMs and DREAD Risk Assessment
- Alleviating Attention Hacking in Discriminative Reward Modeling through Interaction Distillation
- Distributed AI Agents for Cognitive Underwater Robot Autonomy
Related