Towards Integrated Alignment
2025/08/08 by Reis, Ben Y., La Cava, William
#Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2508.06592
Abstract
As AI adoption expands across human society, the problem of aligning AI models to match human preferences remains a grand challenge. Currently, the AI alignment field is deeply divided between behavioral and representational approaches, resulting in narrowly aligned models that are more vulnerable to increasingly deceptive misalignment threats. In the face of this fragmentation, we propose an integrated vision for the future of the field. Drawing on related lessons from immunology and cybersecurity, we lay out a set of design principles for the development of Integrated Alignment frameworks that combine the complementary strengths of diverse alignment approaches through deep integration and adaptive coevolution. We highlight the importance of strategic diversity - deploying orthogonal alignment and misalignment detection approaches to avoid homogeneous pipelines that may be "doomed to success". We also recommend steps for greater unification of the AI alignment research field itself, through cross-collaboration, open model weights and shared community resources.
Citations
- Behavioural vs. Representational Systematicity in End-to-End Models: An Opinionated Survey
- Mitigating Deceptive Alignment via Self-Monitoring
- Auditing language models for hidden objectives
- Training large language models on narrow tasks can lead to broad misalignment
- Zero Trust Architecture: A Systematic Literature Review
- PRISM: Perspective Reasoning for Integrated Synthesis and Mediation as a Multi-Perspective Framework for AI Alignment
- Towards Unified Attribution in Explainable AI, Data-Centric AI, and Mechanistic Interpretability
- DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning
- Medical large language models are vulnerable to data-poisoning attacks
- Alignment faking in large language models
- Paradigms of AI Evaluation: Mapping Goals, Methodologies and Culture
- Analyzing the Generalization and Reliability of Steering Vectors
- Robust Reinforcement Learning from Corrupted Human Feedback
- Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
- Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
- Scaling and evaluating sparse autoencoders
- Not All Language Model Features Are One-Dimensionally Linear
- Cross-Care: Assessing the Healthcare Implications of Pre-training Data on Language Model Bias
- Mechanistic Interpretability for AI Safety -- A Review
- RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations
- Integrating Graceful Degradation and Recovery through Requirement-driven Adaptation
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
- A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
- Alignment for Honesty
- Steering Llama 2 via Contrastive Activation Addition
- Eliciting Latent Knowledge from Quirky Language Models
- In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering
- AI Alignment: A Comprehensive Survey
- AI Supported Degradation of the Self Concept: A Theoretical Framework Grounded in Established Cognitive and Computational Mechanisms
- Getting aligned on representational alignment
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
- Language Models Represent Space and Time
- Representation Engineering: A Top-Down Approach to AI Transparency
- Benchmarks for Detecting Measurement Tampering
- AI Deception: A Survey of Examples, Risks, and Potential Solutions
- MITRE ATT&CK: State of the Art and Way Forward
- Deceptive Alignment Monitoring
- Towards Automated Circuit Discovery for Mechanistic Interpretability
- A Multi-Level Framework for the AI Alignment Problem
- Engineering Monosemanticity in Toy Models
- Goal Misgeneralization: Why Correct Specifications Aren't Enough For Correct Goals
- Polysemanticity and Capacity in Neural Networks
- Governance Architecture for Neural Network Superposition: A Structural Solution to Hallucination via Routing and Interference Filtering
- The Alignment Problem from a Deep Learning Perspective
- Training language models to follow instructions with human feedback
- Toll-Like Receptor Signaling and Its Role in Cell-Mediated Immunity
- Goal Misgeneralization in Deep Reinforcement Learning
- Probing Classifiers: Promises, Shortcomings, and Advances
- Probing Classifiers: Promises, Shortcomings, and Advances
- An overview of 11 proposals for building safe advanced AI
- Synthesize, Execute and Debug: Learning to Repair for Neural Program Synthesis
- Probing the Probing Paradigm: Does Probing Accuracy Entail Task Relevance?
- Defining trained immunity and its role in health and disease
- Roles of repertoire diversity in robustness of humoral immune response
- Scalable agent alignment via reward modeling: a research direction
- Supervising strong learners by amplifying weak experts
- AGI Safety Literature Review
- AI safety via debate
- The Surprising Creativity of Digital Evolution: A Collection of Anecdotes from the Evolutionary Computation and Artificial Life Research Communities
- The Surprising Creativity of Digital Evolution: A Collection of Anecdotes from the Evolutionary Computation and Artificial Life Research Communities
- Understanding intermediate layers using linear classifier probes
- Cooperative Inverse Reinforcement Learning
- Position: Towards Bidirectional Human-AI Alignment
Related