A Roadmap to Pluralistic Alignment
2024/02/07 by Taylor Sorensen, Sorensen, Taylor, Jared Moore +21 · 1 voice · 70 citations
Computer Science · Medicine · Social Sciences · #Artificial Intelligence in Healthcare and Education #Computer science #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #Geography #Political science #Regional science
paper · pdf · doi:10.48550/arxiv.2402.05070
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2024/02/07 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
With increased power and prevalence of AI systems, it is ever more critical that AI systems are designed to serve all, i.e., people with diverse values and perspectives. However, aligning models to serve pluralistic human values remains an open research question. In this piece, we propose a roadmap to pluralistic alignment, specifically using language models as a test bed. We identify and formalize three possible ways to define and operationalize pluralism in AI systems: 1) Overton pluralistic models that present a spectrum of reasonable responses; 2) Steerably pluralistic models that can steer to reflect certain perspectives; and 3) Distributionally pluralistic models that are well-calibrated to a given population in distribution. We also formalize and discuss three possible classes of pluralistic benchmarks: 1) Multi-objective benchmarks, 2) Trade-off steerable benchmarks, which incentivize models to steer to arbitrary trade-offs, and 3) Jury-pluralistic benchmarks which explicitly model diverse human ratings. We use this framework to argue that current alignment techniques may be fundamentally limited for pluralistic AI; indeed, we highlight empirical evidence, both from our own experiments and from other work, that standard alignment procedures might reduce distributional pluralism in models, motivating the need for further research on pluralistic alignment.
Cited by
- AI Value Alignment for Evolving Social Norms
- ToolAlignBench: Investigating Alignment Conflicts in Tool-Calling Enabled LLMs
- The Algorithmic Gaze of Image Quality Assessment: An Audit and Trace Ethnography of the LAION-Aesthetics Predictor
- TALES: A Taxonomy and Analysis of Cultural Representations in LLM-generated Stories
- What's In My Human Feedback? Learning Interpretable Descriptions of Preference Data
- The Homogenizing Effect of Large Language Models on Human Expression and Thought
- Human-Curated Data Authoring with LLMs: A Small-Data Approach to Domain Adaptation
- Stop treating `AGI' as the north-star goal of AI research
- AI as Entertainment
- Alignment Is Not Enough: A Relational Framework for Moral Standing in Human-AI Interaction
- Critically Engaged Pragmatism: Scientific Norm and Social, Pragmatist Epistemology for AI Science Evaluation Tools
- SUMFORU: An LLM-Based Review Summarization Framework for Personalized Purchase Decision Support
- A Systematic Evaluation of Preference Aggregation in Federated RLHF for Pluralistic Alignment of LLMs
- ClinTutor-R1: Advancing Scalable and Robust One-to-Many Alignment in Clinical Socratic Education
- Personalized Reward Modeling for Text-to-Image Generation
- Designing Beyond Language: Sociotechnical Barriers in AI Health Technologies for Limited English Proficiency
- Pluralistic Behavior Suite: Stress-Testing Multi-Turn Adherence to Custom Behavioral Policies
- Human-AI Collaboration with Misaligned Preferences
- CREATE: Testing LLMs for Associative Creativity
- Towards Low-Resource Alignment to Diverse Perspectives with Sparse Feedback
- MENLO: From Preferences to Proficiency -- Evaluating and Modeling Native-like Quality Across 47 Languages
- Botender: Supporting Communities in Collaboratively Designing AI Agents through Case-Based Provocations
- Not My Agent, Not My Boundary? Elicitation of Personal Privacy Boundaries in AI-Delegated Information Sharing
- A Model of Multi-turn Human Persuadability Using Probabilistic Belief Tracing
- Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization
- The Alignment Bottleneck
- Interaction Context Often Increases Sycophancy in LLMs
- Decoding Alignment: A Critical Survey of LLM Development Initiatives through Value-setting and Data-centric Lens
- Heterogeneous preferences and asymmetric insights for AI use among welfare claimants and non-claimants
- CUPID: Evaluating Personalized and Contextualized Alignment of LLMs from Interactions
- Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities
- Learning to summarize user information for personalized reinforcement learning from human feedback
- CALMA: A Process for Deriving Context-aligned Axes for Language Model Alignment
- ALIGN: Prompt-based Attribute Alignment for Reliable, Responsible, and Personalized LLM-based Decision-Making
- On the Impossibility of Separating Intelligence from Judgment: The Computational Intractability of Filtering for AI Alignment
- Scaling Human Judgment in Community Notes with LLMs
- The Singapore Consensus on Global AI Safety Research Priorities
- Perspectives in Play: A Multi-Perspective Approach for More Inclusive NLP Systems
- Aggregated Individual Reporting for Post-Deployment Evaluation
- Reward Model Interpretability via Optimal and Pessimal Tokens
- Because we have LLMs, we Can and Should Pursue Agentic Interpretability
- Beyond RLHF and NLHF: Population-Proportional Alignment under an Axiomatic Framework
- QQSUM: A Novel Task and Model of Quantitative Query-Focused Summarization for Review-based Product Question Answering
- The Coming Crisis of Multi-Agent Misalignment: AI Alignment Must Be a Dynamic and Social Process
- Aligning VLM Assistants with Personalized Situated Cognition
- Fluent but Foreign: Even Regional LLMs Lack Cultural Alignment
- OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models
- Embracing Contradiction: Theoretical Inconsistency Will Not Impede the Road of Building Responsible AI Systems
- AI-Augmented LLMs Achieve Therapist-Level Responses in Motivational Interviewing
- MPO: Multilingual Safety Alignment via Reward Gap Optimization
- Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmas
- Is Active Persona Inference Necessary for Aligning Small Models to Personal Preferences?
- Scoring Rules! Statistical and Strategic Alignment for Text Evaluation Metrics
- Pairwise Calibrated Rewards for Pluralistic Alignment
- The Value of Disagreement in AI Design, Evaluation, and Alignment
- Political Neutrality as Balanced Approval: A Large-Scale Human Evaluation of AI Responses
- Architecting Trust in Artificial Epistemic Agents
- Positive Alignment: Artificial Intelligence for Human Flourishing
- Understanding Annotator Safety Policy with Interpretability
- Relative Principals, Pluralistic Alignment, and the Structural Value Alignment Problem
- What Does the AI Doctor Value? Auditing Pluralism in the Clinical Ethics of Language Models
- Inducing Sustained Creativity and Diversity in Large Language Models
- Reward Models Inherit Value Biases from Pretraining
- Socially Grounded Agentic AI: Coordinating Plural Perspectives through Social Theory
- LoRe: Personalizing LLMs via Low-Rank Reward Modeling
- Persona-judge: Personalized Alignment of Large Language Models via Token-level Self-judgment
- Cautious Context Steering for Language Model Personalization
- DICE: A Framework for Dimensional and Contextual Evaluation of Language Models
- Societal Impacts Research Requires Benchmarks for Creative Composition Tasks
- Mixture-of-Personas Language Models for Population Simulation
Discussions
Related