Positive Alignment: Artificial Intelligence for Human Flourishing
2026/05/11 by Ruben Laukkonen, Seb Krier, Chloé Bakalar +13 · 3 voices
Biochemistry, Genetics and Molecular Biology · Computer Science · #cs.AI #cs.CY #cs.HC #q-bio.NC
paper · pdf · doi:10.48550/arxiv.2605.10310
arxiv published 2026/05/11 · arxiv updated 2026/06/19
Abstract
Existing alignment research is dominated by concerns about safety and preventing harm: safeguards, controllability, and compliance. This paradigm of alignment parallels early psychology's focus on mental illness: necessary but incomplete. What we call Positive Alignment is the development of AI systems that (i) actively support human and ecological flourishing in a pluralistic, polycentric, context-sensitive, and user-authored way while (ii) remaining safe and cooperative. It is a distinct and necessary agenda within AI alignment research. We argue that several existing failures of alignment (e.g., engagement hacking, loss of human autonomy, failures in truth-seeking, low epistemic humility, error correction, lack of diverse viewpoints, and being primarily reactive rather than proactive) may be better addressed through positive alignment, including cultivating virtues and maximizing human flourishing. We highlight a range of challenges, open questions, and technical directions (e.g., data filtering and upsampling, pre- and post-training, evaluations, collaborative value collection) for different phases of the LLM and agents lifecycle. We end with design principles for promoting disagreement and decentralization through contextual grounding, community customization, continual adaptation, and polycentric governance; that is, many legitimate centers of oversight rather than one institutional or moral chokepoint.
Citations
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment
- Legal Alignment for Safe and Ethical AI
- Emergent Introspective Awareness in Large Language Models
- CoPE: A Small Language Model for Steerable and Scalable Content Labeling
- Full-Stack Alignment: Co-Aligning AI and Institutions with Thick Models of Value
- Neural steering vectors reveal dose and exposure-dependent impacts of human-AI relationships
- Operationalizing Pluralistic Values in Large Language Model Alignment Reveals Trade-offs in Safety, Inclusivity, and Model Behavior
- (Working Paper) Good Faith Design: Religion as a Resource for Technologists
- Consistency Training Helps Stop Sycophancy and Jailbreaks
- MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes
- Measuring Epistemic Humility in Multimodal Large Language Models
- Decentralising LLM Alignment: A Case for Context, Pluralism, and Participation
- An Economy of AI Agents
- Benchmarking Deception Probes via Black-to-White Performance Boosts
- Cultivating Pluralism In Algorithmic Monoculture: The Community Alignment Dataset
- Measuring AI Alignment with Human Flourishing
- A beautiful loop: An active inference theory of consciousness
- C-VARC: A Large-Scale Chinese Value Rule Corpus for Value Alignment of Large Language Models
- Societal and technological progress as sewing an ever-growing, ever-changing, patchy, and polychrome quilt
- Artificial Intelligences: A Bridge Toward Diverse Intelligence and Humanity's Future
- Contemplative Agent
- Values in the Wild: Discovering and Analyzing Values in Real-World Language Model Interactions
- Epistemic Alignment: A Mediating Framework for User-LLM Knowledge Delivery
- A matter of principle? AI alignment as the fair treatment of claims
- Auditing language models for hidden objectives
- PluralLLM: Pluralistic Alignment in LLMs via Federated Learning
- Societal Alignment Frameworks Can Improve LLM Alignment
- Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
- Why human-AI relationships need socioaffective alignment
- Investigating machine moral judgement through the Delphi experiment
- A theory of appropriateness with applications to generative artificial intelligence
- The Road to Artificial SuperIntelligence: A Comprehensive Survey of Superalignment
- AI can help humans find common ground in democratic deliberation
- DailyDilemmas: Revealing Value Preferences of LLMs with Quandaries of Daily Life
- CulturalBench: A Robust, Diverse, and Challenging Cultural Benchmark by Human-AI CulturalTeaming
- Beyond Preferences in AI Alignment
- Personality Alignment of Large Language Models
- Dimensions of wisdom perception across twelve countries on five continents
- PERSONA: A Reproducible Testbed for Pluralistic Alignment
- ProgressGym: Alignment with a Millennium of Moral Progress
- The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
- Modular Pluralism: Pluralistic Alignment via Multi-LLM Collaboration
- External Invariants: A Cryptographic Trust Architecture for Institutional AI Inference
- HelpSteer2: Open-source dataset for training top-performing reward models
- Language Models Resist Alignment: Evidence From Data Compression
- Inverse Constitutional AI: Compressing Preferences into Principles
- The Platonic Representation Hypothesis
- Towards Guaranteed Safe AI: A Framework for Ensuring Robust and Reliable AI Systems
- Value Augmented Sampling for Language Model Alignment and Personalization
- From Persona to Personalization: A Survey on Role-Playing Language Agents
- OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
- Measuring Political Bias in Large Language Models: What Is Said and How It Is Said
- A Roadmap to Pluralistic Alignment
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- Flames: Benchmarking Value Alignment of LLMs in Chinese
- A General Theoretical Paradigm to Understand Learning from Human Preferences
- MemGPT: Towards LLMs as Operating Systems
- Consciousness in Artificial Intelligence: Insights from the Science of Consciousness
- Code of "Sirens' Whisper: Inaudible Near-Ultrasonic Jailbreaks of Speech-Driven LLMs"
- Evaluating the Moral Beliefs Encoded in LLMs
- Towards Measuring the Representation of Subjective Global Opinions in Language Models
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- NormBank: A Knowledge Bank of Situational Social Norms
- MemoryBank: Enhancing Large Language Models with Long-Term Memory
- Multi-Value Alignment in Normative Multi-Agent System: An Evolutionary Optimisation Approach
- Multi-Value Alignment in Normative Multi-Agent System: Evolutionary Optimisation Approach
- LaMP: When Large Language Models Meet Personalization
- Generative Agents: Interactive Simulacra of Human Behavior
- Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark
- Whose Opinions Do Language Models Reflect?
- Could a Large Language Model be Conscious?
- Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
- Machine Love
- The Capacity for Moral Self-Correction in Large Language Models
- Discovering Language Model Behaviors with Model-Written Evaluations
- Affective Coherence Monitoring for Transformer-Based Language Models
- Discovering Latent Knowledge in Language Models Without Supervision
- Fine-tuning language models to find agreement among humans with diverse preferences
- Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task
- Out of One, Many: Using Language Models to Simulate Human Samples
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and\n Implicit Hate Speech Detection
- Training language models to follow instructions with human feedback
- Survey of Hallucination in Natural Language Generation
- Red Teaming Language Models with Language Models
- BBQ: A Hand-Built Bias Benchmark for Question Answering
- BBQ: A hand-built bias benchmark for question answering
- TruthfulQA: Measuring How Models Mimic Human Falsehoods
- Informational Design of Dynamic Multi-Agent System
- Discipline and Punish: The Birth of the Prison
- The 2020 Five Domains Model: Including Human–Animal Interactions in Assessments of Animal Welfare
- RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language\n Models
- Learning to summarize from human feedback
- Aligning AI With Shared Human Values
- Liquid Time-constant Networks
- Artificial Intelligence, Values, and Alignment
- Risks from Learned Optimization in Advanced Machine Learning Systems
- Personal Universes: A Solution to the Multi-Agent Value Alignment Problem
- AI safety via debate
- Deep reinforcement learning from human preferences
- Concrete Problems in AI Safety
- Beyond Markets and States: Polycentric Governance of Complex Economic Systems
- The cognitive revolution: a historical perspective
- The Open Society and its Enemies
- Contemplative Superalignment
- How AI and Human Behaviors Shape Psychosocial Effects of Extended Chatbot Use: A Longitudinal Randomized Controlled Study
- Regulatory Markets: The Future of AI Governance
- Self-determination theory and the facilitation of intrinsic motivation, social development, and well-being.
- Positive psychology: An introduction.
Discussions
- A new paper presents "Positive Alignment," an AI approach that prioritizes human and ecological flourishing while ensuring safety. The authors advocate for a shift from traditional alignment concerns [bsky, 1 points, 1 comments]
- Positive Alignment: Artificial Intelligence for Human Flourishing [hn, 1 points, 0 comments]
- Amazing work by @RubenLaukkonen, @drmichaellevin, @weballergy, @verena_rieser, @AdamCElwood, @996roma, @FranklinMatija, @shamilch, @_fernando_rosas, @scychan_brains, @matybohacek, @sudoraohacker, and [bsky, 0 points, 0 comments]