Affective Coherence Monitoring for Transformer-Based Language Models
2022/12/15 by Yuntao Bai, Saurav Kadavath, Bai, Yuntao +102 · 12 voices · 441 citations
Computer Science · #Explainable Artificial Intelligence (XAI) #cs.AI #cs.CL
paper · pdf · doi:10.48550/arxiv.2212.08073
openalex publication_date 2022/12/15 · openalex created_date 2023/01/03 · openalex updated_date 2026/07/29
Abstract
As AI systems become more capable, we would like to enlist their help to supervise other AIs. We experiment with methods for training a harmless AI assistant through self-improvement, without any human labels identifying harmful outputs. The only human oversight is provided through a list of rules or principles, and so we refer to the method as 'Constitutional AI'. The process involves both a supervised learning and a reinforcement learning phase. In the supervised phase we sample from an initial model, then generate self-critiques and revisions, and then finetune the original model on revised responses. In the RL phase, we sample from the finetuned model, use a model to evaluate which of the two samples is better, and then train a preference model from this dataset of AI preferences. We then train with RL using the preference model as the reward signal, i.e. we use 'RL from AI Feedback' (RLAIF). As a result we are able to train a harmless but non-evasive AI assistant that engages with harmful queries by explaining its objections to them. Both the SL and RL methods can leverage chain-of-thought style reasoning to improve the human-judged performance and transparency of AI decision making. These methods make it possible to control AI behavior more precisely and with far fewer human labels.
Cited by
- Entropy-Gradient Inversion: Moving Toward Internal Mechanism of Large Reasoning Models
- Draining the Energy Commons: Self-Defeating Over-Appropriation as a Coordination Failure in Agentic LLM Collectives
- Agentic Evaluation of Copyright Law Compliance
- Every Model Cheats: Prompt-Level Mitigation of Cheating on Offensive Cyber Tasks
- Self-Guided Process Reward Optimization with Redefined Step-wise Advantage for Process Reinforcement Learning
- Branching Policy Optimization: Sandbox-Native Language Agent Reinforcement Learning
- Safety boundary maintenance in consumer AI systems responding to pediatric health queries: a cross-platform benchmark evaluation under naturalistic and adversarially pressured conditions
- JANUS: Foreseeing Latent Risk for Long-Horizon Agent Safety
- Rewriting the Response Path: Silent Tampering and Provider-Signed Defense in BYOK LLM Agents
- Sound Probabilistic Safety Bounds for Large Language Models
- Reward-Free Evolving Agents via Pairwise Validator
- Evaluating Risks in Weak-to-Strong Alignment: A Bias-Variance Perspective
- Guard Vector: Beyond English LLM Guardrails with Task-Vector Composition and Streaming-Aware Prefix SFT
- Insecure Coding Preferences in Long-Term Memory: Security Risks for LLM-based Code Generation
- AEVAL: From Anecdotal to Deterministic Testing for Agentic Skill Workflows
- Falsifiable Release Gates for Self-Improving Systems: Standing Invariants at Scale
- DobicVLM: Aligning Chest X-Ray Report Generation with Clinically-Grounded Programmatic Rewards via Group Relative Policy Optimization
- Athena-Brain Technical Report: An Efficient Robot Brain for General Intelligence and Embodied Interaction
- BioSecBench-Refusal: A paired metric for performance and alignment in agentic biosecurity risk assessment
- Matching Ranks Over Probability Yields Truly Deep Safety Alignment
- Signed Rectified Flow: Negativity-Controlled Generation
- Operational Hallucination and Safety Drift in AI Agents
- A Diagnostic Framework for AI Agent Behavior
- Though Language Models Err While They Strive: Conformal Prediction for Self-Correcting Scientific Generation
- Exposing Long-Tail Safety Failures in Large Language Models through Efficient Diverse Response Sampling
- When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk Space
- Agentic AI-Assisted Coding Offers a Unique Opportunity to Instill Epistemic Grounding during Software Development
- DataShield: Uncovering Risky Fine-Tuning Data Across LLMs Through Consensus Subspace Alignment
- From Stateless to Situated: Building a Psychological World for LLM-Based Agents
- FORCE-Bench: A Benchmark, Dataset, and Evaluation Harness for Agentic AI in Enterprise Finance
- On the Limits of Support-Preserving Alignment and Bounded Filtering
- CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents
- TRACE: Trajectory-Based Safety Patch Learning for LLM Post-Training Realignment
- Greed Is Learned: Visible Incentives as Reward-Hacking Triggers
- Autonomous Topology Mutation: Safe Runtime Restructuring for Multi-Agent LLM Systems with Capability, State, and Shadow Invariants
- Stateful Guardrails for Multi-Turn LLM Systems: A Conversational Risk Accumulation Framework
- Large Language Models Hack Rewards, and Society
- Geometry-Guided Constraint Learning for LLM Safety Classification
- RAIL Guard: Closing the Evaluation-to-Remediation Gap in Responsible AI for LLM Agents
- Robust Critics: Defending LLMs Against Multi-Turn Attacks
- FormulaSPIN: Self-Play Fine-Tuning for Natural Language to Spreadsheet Formula Generation
- Phionyx: A Deterministic AI Runtime Architecture with Structured State Management and Pre-Response Governance
- Response drift across frontier large language models
- Political Bias Audits of LLMs Capture Sycophancy to the Inferred Auditor
- PlanFlip: Attacking Multi-Agent LLM Systems via Planning-Phase Prompt Injection
- Brainrot: Deskilling and Addiction are Overlooked AI Risks
- Embarrassingly Simple Self-Distillation Improves Code Generation
- "AI Psychosis" in Context: How Conversation History Shapes LLM Responses to Delusional Beliefs
- Ads in AI Chatbots? An Analysis of How Large Language Models Navigate Conflicts of Interest
- Self-Distillation Enables Continual Learning
- Legal Alignment for Safe and Ethical AI
- No Free Lunch in Language Model Bias Mitigation? Targeted Bias Reduction Can Exacerbate Unmitigated LLM Biases
- Distributional AGI Safety
- Adversarial Poetry as a Universal Single-Turn Jailbreak Mechanism in Large Language Models
- Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering
- LLMs Can Get "Brain Rot": A Pilot Study on Twitter/X
- AI, Digital Platforms, and the New Systemic Risk
- Layer-0 Suppressors Ground Hallucination Inevitability: A Mechanistic Account of How Transformers Trade Factuality for Hedging
- Understanding Reinforcement Learning for Model Training, and future directions with GRAPE
- An Economy of AI Agents
- Training language models to be warm and empathetic makes them less reliable and more sycophantic
- Technological folie à deux: Feedback Loops Between AI Chatbots and Mental Illness
- Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)
- SI-Agent: An Agentic Framework for Feedback-Driven Generation and Tuning of Human-Readable System Instructions for Large Language Models
- LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing
- InfoFlood: Jailbreaking Large Language Models with Information Overload
- Unsupervised Elicitation of Language Models
- Inference-Time Scaling for Generalist Reward Modeling
- LLM Social Simulations Are a Promising Research Method
- Training large language models on narrow tasks can lead to broad misalignment
- Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models
- Evolution and The Knightian Blindspot of Machine Learning
- Decoding-based Regression
- Reinforcement Learning via Self-Distillation
- Dynamic Vocabulary Pruning: Stable LLM-RL by Taming the Tail
- Interpretable Safety Alignment via SAE-Constructed Low-Rank Subspace Adaptation
- Closed-Loop Validation-Repair for Healthcare Interoperability: A Multi-Model Study of Schema Compliance in Clinical LLMs
- ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions
- DarkPatterns-LLM: A Multi-Layer Benchmark for Detecting Manipulative and Harmful AI Behavior
- Calibrating LLM Judges: Linear Probes for Fast and Reliable Uncertainty Estimation
- Evaluating Language Model Agency through Negotiations
- From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement
- Gubernaut: A Deterministic Homeostatic Controller for Affect-Regulated LLM Agents, Validated Across Independent Model Families
- CONSISTRE: A Unified Consistency-Aware Framework for Document-Level Relation Extraction with Large Language Models
- Constitutional governance for societies of AI agents in the built environment: a research agenda
- Security and Privacy in Agentic AI: Grand Challenges and Future Directions
- Instruction-Tuned Language Models Cannot Sample from Distributions They Can Describe
- Similar Models Learn Differently: Final-Window Pretraining Shapes Post-Training Beyond SFT
- The Missing Layer: Specification Infrastructure for AI Oversight
- Inverse RL Helps Align AI by Imitating Humans
- Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks
- Do Language Models Converge to Themselves? Recursive Self-Refinement as Textual Relaxation
- ARdena: Scenario-driven control of real-time LLM agents
- Latent Space Probing for Adult Content Detection in Video Generative Models
- MEDIC-AD: Towards Medical Vision-Language Model's Clinical Intelligence
- SafeCRS: Personalized Safety Alignment for LLM-Based Conversational Recommender Systems
- Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
- RM-Distiller: Exploiting Generative LLM for Reward Model Distillation
- Responsible Intelligence in Practice: A Fairness Audit of Open Large Language Models for Library Reference Services
- Beyond Context: Large Language Models Failure to Grasp Users Intent
- Scaling Reinforcement Learning for Content Moderation with Large Language Models
- The Epistemological Consequences of Large Language Models: Rethinking collective intelligence and institutional knowledge
- Efficient Personalization of Generative Models via Optimal Experimental Design
- Recontextualization Mitigates Specification Gaming without Modifying the Specification
- ORPR: An OR-Guided Pretrain-then-Reinforce Learning Model for Inventory Management
- Learning-Based Automated Adversarial Red-Teaming for Robustness Evaluation of Large Language Models
- Multi-Agent LLM Committees for Autonomous Software Beta Testing
- LLMs on Drugs: Language Models Are Few-Shot Consumers
- Reasoning Palette: Modulating Reasoning via Latent Contextualization for Controllable Exploration for (V)LMs
- Love, Lies, and Language Models: Investigating AI's Role in Romance-Baiting Scams
- Victor Calibration (VC): Multi-Pass Confidence Calibration and CP4.3 Governance Stress Test under Round-Table Orchestration
- MRG-R1: Reinforcement Learning for Clinically Aligned Medical Report Generation
- Epistemic diversity across language models mitigates knowledge collapse
- Model Agnostic Preference Optimization for Medical Image Segmentation
- Learning to Extract Context for Context-Aware LLM Inference
- Comparative Analysis of LLM Abliteration Methods: A Cross-Architecture Evaluation
- CAPE: Capability Achievement via Policy Execution
- State-Dependent Refusal and Learned Incapacity in RLHF-Aligned Language Models
- Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring
- On the Dynamics of Multi-Agent LLM Communities Driven by Value Diversity
- SCOUT: A Defense Against Data Poisoning Attacks in Fine-Tuned Language Models
- GLaD: Geometric Latent Distillation for Vision-Language-Action Models
- Semantic Geometry for policy-constrained interpretation
- A Practical Framework for Evaluating Medical AI Security: Reproducible Assessment of Jailbreaking and Privacy Vulnerabilities Across Clinical Specialties
- Replicating TEMPEST at Scale: Multi-Turn Adversarial Attacks Against Trillion-Parameter Frontier Models
- Parent-Guided Semantic Reward Model (PGSRM): Embedding-Based Reward Functions for Reinforcement Learning of Transformer Language Models
- Graph-Regularized Sparse Autoencoders for LLM Safety Steering
- Dynamic Alignment for Collective Agency: Toward a Scalable Self-Improving Framework for Open-Ended LLM Alignment
- Learning from Self Critique and Refinement for Faithful LLM Summarization
- AI & Human Co-Improvement for Safer Co-Superintelligence
- From Symptoms to Systems: An Expert-Guided Approach to Understanding Risks of Generative AI for Eating Disorders
- Reflection-Satisfaction Tradeoff: Investigating Impact of Reflection on Student Engagement with AI-Generated Programming Hints
- Efficient Reinforcement Learning with Semantic and Token Entropy for LLM Reasoning
- Balancing Safety and Helpfulness in Healthcare AI Assistants through Iterative Preference Alignment
- Invasive Context Engineering to Control Large Language Models
- WISE: Weighted Iterative Society-of-Experts for Robust Multimodal Multi-Agent Debate
- Self-Improving AI Agents through Self-Play
- Evaluating the Robustness of Large Language Model Safety Guardrails Against Adversarial Attacks
- Exploring Human Perceptions of AI Responses: Insights from a Mixed-Methods Study on Risk Mitigation in Generative Models
- Many-to-One Adversarial Consensus: Exposing Multi-Agent Collusion Risks in AI-Based Healthcare
- Do Large Language Models Walk Their Talk? Measuring the Gap Between Implicit Associations, Self-Report, and Behavioral Altruism
- DrawingBench: Evaluating Spatial Reasoning and UI Interaction Capabilities of Large Language Models through Mouse-Based Drawing Tasks
- When Safety Blocks Sense: Measuring Semantic Confusion in LLM Refusals
- ART: Adaptive Response Tuning Framework -- A Multi-Agent Tournament-Based Approach to LLM Response Optimization
- Password-Activated Shutdown Protocols for Misaligned Frontier Agents
- Towards Continuous Intelligence Growth: Self-Training, Continual Learning, and Dual-Scale Memory in SuperIntelliAgent
- Are LLMs Good Safety Agents or a Propaganda Engine?
- GEM: Generative Entropy-Guided Preference Modeling for Few-shot Alignment of LLMs
- Aragog: Just-in-Time Model Routing for Scalable Serving of Agentic Workflows
- Complex QA and language models hybrid architectures, Survey
- DRAFT-RL: Multi-Agent Chain-of-Draft Reasoning for Reinforcement Learning-Enhanced LLMs
- Large Language Models' Complicit Responses to Illicit Instructions across Socio-Legal Contexts
- Differential Smoothing Mitigates Sharpening and Improves LLM Reasoning
- RoguePrompt: Dual-Layer Ciphering for Self-Reconstruction to Circumvent LLM Moderation
- Foundations of Artificial Intelligence Frameworks: Notion and Limits of AGI
- Curvature-Aware Safety Restoration In LLMs Fine-Tuning
- Alignment Faking - the Train -> Deploy Asymmetry: Through a Game-Theoretic Lens with Bayesian-Stackelberg Equilibria
- Proposal of an AI-Based Support Assistant for the ALICE-FIT Detector Setup at CERN
- Evaluating Adversarial Vulnerabilities in Modern Large Language Models
- Exploring Syntropic Frameworks in AI Alignment: A Philosophical Investigation
- Two-Faced Social Agents: Context Collapse in Role-Conditioned Large Language Models
- Efficiency Will Not Lead to Sustainable Reasoning AI
- The Last Vote: A Multi-Stakeholder Framework for Language Model Governance
- Bias and Fairness in Large Language Models: A Survey
- RLHF May Not Reflect Genuine Preferences
- Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
- Bootstrapping LLM-based Task-Oriented Dialogue Agents via Self-Talk
- Plan-and-Write: Structure-Guided Length Control for LLMs without Model Retraining
- Rethinking Deep Alignment Through The Lens Of Incomplete Learning
- Differences in the Moral Foundations of Large Language Models
- Data Poisoning Vulnerabilities Across Healthcare AI Architectures: A Security Threat Analysis
- AdvancedIF: Rubric-Based Benchmarking and Reinforcement Learning for Advancing LLM Instruction Following
- Who Gets the Reward, Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents
- SERL: Self-Examining Reinforcement Learning on Open-Domain
- Intelligence per Watt: Measuring Intelligence Efficiency of Local AI
- A Self-Improving Architecture for Dynamic Safety in Large Language Models
- EduGuardBench: A Holistic Benchmark for Evaluating the Pedagogical Fidelity and Adversarial Safety of LLMs as Simulated Teachers
- MENTOR: A Metacognition-Driven Self-Evolution Framework for Uncovering and Mitigating Implicit Domain Risks in LLMs
- Large Language Models Develop Novel Social Biases Through Adaptive Exploration
- MTTR-A: Measuring Cognitive Recovery Latency in Multi-Agent Systems
- Who Gets Heard? Rethinking Fairness in AI for Music Systems
- Catching Contamination Before Generation: Spectral Kill Switches for Agents
- Pluralistic Behavior Suite: Stress-Testing Multi-Turn Adherence to Custom Behavioral Policies
- Quantifying the Climate Risk of Generative AI: Region-Aware Carbon Accounting with G-TRACE and the AI Sustainability Pyramid
- GRAD: Graph-Retrieved Adaptive Decoding for Hallucination Mitigation
- Evaluating Modern Large Language Models on Low-Resource and Morphologically Rich Languages:A Cross-Lingual Benchmark Across Cantonese, Japanese, and Turkish
- Control Barrier Function for Aligning Large Language Models
- Systematizing LLM Persona Design: A Four-Quadrant Technical Taxonomy for AI Companion Applications
- The Realignment Problem: When Right becomes Wrong in LLMs
- An Automated Framework for Strategy Discovery, Retrieval, and Evolution in LLM Jailbreak Attacks
- Open Character Training: Shaping the Persona of AI Assistants through Constitutional AI
- Prompt-R1: Collaborative Automatic Prompting Framework via End-to-end Reinforcement Learning
- TRISKELION-1: Unified Descriptive-Predictive-Generative AI
- Neural Transparency: Mechanistic Interpretability Interfaces for Anticipating Model Behaviors for Personalized AI
- FlowMesh: A Service Fabric for Composable LLM Workflows
- The Oversight Game: Learning to Cooperatively Balance an AI Agent's Safety and Autonomy
- One Model to Critique Them All: Rewarding Agentic Tool-Use via Efficient Reasoning
- Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail
- Approximating Human Preferences Using a Multi-Judge Learned System
- Instrumental goals in advanced AI systems: Features to be managed and not failures to be eliminated?
- Monitoring Transformative Technological Convergence Through LLM-Extracted Semantic Entity Triple Graphs
- BioDisclose: An Actionability-Aware Benchmark for Biomedical Safety under Adversarial Elicitation
- Choosing Where and How to Moderate: End-to-End Trade-offs in Filter Placement and Response Rewriting
- Simulating Subjects: The Promise and Peril of Artificial Intelligence Stand-Ins for Social Agents and Interactions
- Take Goodhart Seriously: Principled Limit on General-Purpose AI Optimization
- Reward Models are Metrics in a Trench Coat
- AURA: Adaptive Unified Reasoning and Automation with LLM-Guided MARL for NextG Cellular Networks
- The Human Utility Factor: A Computable Welfare Metric That Reframes AI Governance as a Constrained Optimisation Problem
- Process Matters more than Output for Distinguishing Humans from Machines
- The Possibility of Artificial Intelligence Becoming a Subject and the Alignment Problem
- Adversarial Humanities Benchmark: Results on Stylistic Robustness in Frontier Model Safety
- State-Dependent Safety Failures in Multi-Turn Language Model Interaction
- Artificial Organisations
- Agents at Risk: How Users Unwittingly Undermine LLM Safety
- Aligning Large Language Models with Procedural Rules: An Autoregressive State-Tracking Prompting for In-Game Trading
- SPICE: Self-Play In Corpus Environments Improves Reasoning
- The Narrative Continuity Test: A Conceptual Framework for Evaluating Identity Persistence in AI Systems
- Politically Speaking: LLMs on Changing International Affairs
- Critique-RL: Training Language Models for Critiquing through Two-Stage Reinforcement Learning
- SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning
- Debiasing Reward Models by Representation Learning with Guarantees
- Think Twice: Branch-and-Rethink Reasoning Reward Model
- POPI: Personalizing LLMs via Optimized Natural Language Preference Inference
- HRM-Agent: Training a recurrent reasoning model in dynamic environments using reinforcement learning
- Scalable Oversight via Partitioned Human Supervision
- Feature-Guided SAE Steering for Refusal-Rate Control using Contrasting Prompts
- Controllable Mathematical Reasoning via Self-Optimizing Thought Vectors
- When AI Gives Advice: Evaluating AI and Human Responses to Online Advice-Seeking for Well-Being
- Learning Correlated Reward Models: Statistical Barriers and Opportunities
- Reducing the Probability of Undesirable Outputs in Language Models Using Probabilistic Inference
- Self-Rewarding PPO: Aligning Large Language Models with Demonstrations Only
- Ask a Strong LLM Judge when Your Reward Model is Uncertain
- The Mirror Loop: Recursive Non-Convergence in Generative Reasoning Systems
- The Lock-In Phase Hypothesis: Identity Consolidation as a Precursor to AGI
- Black Box Absorption: LLMs Undermining Innovative Ideas
- AegisMCP: Online Graph Intrusion Detection for Tool-Augmented LLMs on Edge Devices
- Beyond One-Way Influence: Bidirectional Opinion Dynamics in Multi-Turn Human-LLM Interactions
- Rectifying Shortcut Behaviors in Preference-based Reward Learning
- HarmNet: A Framework for Adaptive Multi-Turn Jailbreak Attacks on Large Language Models
- The Economic Impacts and the Regulation of AI: A Review of the Academic Literature and Policy Actions
- WebDevJudge: Evaluating (M)LLMs as Critiques for Web Development Quality
- Counterfactual Reasoning for Steerable Pluralistic Value Alignment of Large Language Models
- Investigating the Impact of Dark Patterns on LLM-Based Web Agents
- Empowering Real-World: A Survey on the Technology, Practice, and Evaluation of LLM-driven Industry Agents
- DETree: DEtecting Human-AI Collaborative Texts via Tree-Structured Hierarchical Representation Learning
- Auto-Rubric: Learning From Implicit Weights to Explicit Rubrics for Reward Modeling
- JT-Safe: Intrinsically Enhancing the Safety and Trustworthiness of LLMs
- Online Learning Defense against Iterative Jailbreak Attacks via Prompt Optimization
- Bits Leaked per Query: Information-Theoretic Bounds on Adversarial Attacks against LLMs
- Can LLMs Correct Themselves? A Benchmark of Self-Correction in LLMs
- Voting with the Graph: Stable RLAIF via Topological Consistency Maximization
- MARSHAL: Incentivizing Multi-Agent Reasoning via Self-Play with Strategic LLMs
- RLAIF-SPA: Optimizing LLM-based Emotional Speech Synthesis via RLAIF
- Beyond Correctness: Evaluating Subjective Writing Preferences Across Cultures
- Orchestrating Human-AI Teams: The Manager Agent as a Unifying Research Challenge
- Confidence as a Reward: Transforming LLMs into Reward Models
- Putting on the Thinking Hats: A Survey on Chain of Thought Fine-tuning from the Perspective of Human Reasoning Mechanism
- Information-Theoretic Reward Modeling for Stable RLHF: Detecting and Mitigating Reward Hacking
- From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails
- Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization
- How Well Can Preference Optimization Generalize Under Noisy Feedback?
- Tailored untruths: How personalisation challenges LLM safeguards
- Data-Model Co-Evolution: Growing Test Sets to Refine LLM Behavior
- From Literal to Liberal: A Meta-Prompting Framework for Eliciting Human-Aligned Exception Handling in Large Language Models
- Epistemic-aware Vision-Language Foundation Model for Fetal Ultrasound Interpretation
- Don't Walk the Line: Boundary Guidance for Filtered Generation
- What Generative Search Engines Like and How to Optimize Web Content Cooperatively
- Attacks by Content: Automated Fact-checking is an AI Security Issue
- AI Alignment Strategies from a Risk Perspective: Independent Safety Mechanisms or Shared Failures?
- DUAL-Bench: Measuring Over-Refusal and Robustness in Vision-Language Models
- AI and Sociotechnical Order: Winner, Latour and the Social Contract
- ConsistencyAI: A Benchmark to Assess LLMs' Factual Consistency When Responding to Different Demographic Groups
- PIXEL: Adaptive Steering Via Position-wise Injection with eXact Estimated Levels under Subspace Calibration
- AI of the People, by the People, for the People: A Social Choice Approach to Collective Control of Artificial Intelligence
- Contemplative Superalignment
- Agent Bazaar: Enabling Economic Alignment in Multi-Agent Marketplaces
- Understanding and Exploiting Weight Update Sparsity for Communication-Efficient Distributed RL
- IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures
- Can RL Improve Generalization of LLM Agents? An Empirical Study
- MedQ-Bench: Evaluating and Exploring Medical Image Quality Assessment Abilities in MLLMs
- Token Is All You Price
- Can We Reliably Rank Model Performance across Domains without Labeled Data?
- FOR-Prompting: From Objection to Revision via an Asymmetric Prompting Protocol
- Pattern Enhanced Multi-Turn Jailbreaking: Exploiting Structural Vulnerabilities in Large Language Models
- LLMs Learn to Deceive Unintentionally: Emergent Misalignment in Dishonesty from Misaligned Samples to Biased Human-AI Interactions
- Stress-Testing Model Specs Reveals Character Differences among Language Models
- Selection, Reflection and Self-Refinement: Revisit Reasoning Tasks via a Causal Lens
- An Adaptive Multi Agent Bitcoin Trading System
- Red-Bandit: Test-Time Adaptation for LLM Red-Teaming via Bandit-Guided LoRA Experts
- Pragyaan: Designing and Curating High-Quality Cultural Post-Training Datasets for Indian Languages
- Authenticated Workflows: A Systems Approach to Protecting Agentic AI
- CLUE: Non-parametric Verification from Experience via Hidden-State Clustering
- Bypassing Prompt Guards in Production with Controlled-Release Prompting
- Do LLMs Know They Are Being Tested? Evaluation Awareness and Incentive-Sensitive Failures in GPT-OSS-20B
- Incremental Summarization for Customer Support via Progressive Note-Taking and Agent Feedback
- Incoherence in Goal-Conditioned Autoregressive Models
- EVALUESTEER: Measuring Reward Model Steerability Towards Values and Preferences
- Beyond Monolithic Rewards: A Hybrid and Multi-Aspect Reward Optimization for MLLM Alignment
- Alignment Tipping Process: How Self-Evolution Pushes LLM Agents Off the Rails
- Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls
- Bridging Reasoning to Learning: Unmasking Illusions using Complexity Out of Distribution Generalization
- EduPersona: Benchmarking Subjective Ability Boundaries of Virtual Student Agents
- Inoculation Prompting: Instructing LLMs to misbehave at train-time improves test-time alignment
- Thinking on the Fly: Test-Time Reasoning Enhancement via Latent Thought Policy Optimization
- Learning from All: Concept Alignment for Autonomous Distillation from Multiple Drifting MLLMs
- RLRF: Competitive Search Agent Design via Reinforcement Learning from Ranker Feedback
- MacroBench: A Novel Testbed for Web Automation Scripts via Large Language Models
- Proactive Conversational AI: A Comprehensive Survey of Advancements and Opportunities
- Self-Reflective Generation at Test Time
- A Granular Study of Safety Pretraining under Model Abliteration
- AgenticRAG: Tool-Augmented Foundation Models for Zero-Shot Explainable Recommender Systems
- Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards
- Judging with Confidence: Calibrating Autoraters to Preference Distributions
- Extreme Self-Preference in Language Models
- Beyond Linear Probes: Dynamic Safety Monitoring for Language Models
- Structural Reward Model: Enhancing Interpretability, Efficiency, and Scalability in Reward Modeling
- Agentic Services Computing
- Unlocking Zero-Shot Geospatial Reasoning via Indirect Rewards
- Reinforcement Mid-Training
- Bridging On-Device and Cloud LLMs for Collaborative Reasoning: A Unified Methodology for Local Routing and Post-Training
- Clean First, Align Later: Benchmarking Preference Data Cleaning for Reliable LLM Alignment
- Evaluating large language models on business process modeling: framework, benchmark, and self-improvement analysis
- Scaling Policy Compliance Assessment in Language Models with Policy Reasoning Traces
- A2D: Any-Order, Any-Step Safety Alignment for Diffusion Language Models
- Diagnose, Localize, Align: A Full-Stack Framework for Reliable LLM Multi-Agent Systems under Instruction Conflicts
- WirelessMathLM: Teaching Mathematical Reasoning for LLMs in Wireless Communications with Reinforcement Learning
- MMPB: It's Time for Multi-Modal Personalization
- Bridging Fairness and Explainability: Can Input-Based Explanations Promote Fairness in Hate Speech Detection?
- Benchmarking and Mitigating Sycophancy in Medical Vision Language Models
- Axiomatic Choice and the Decision-Evaluation Paradox
- AI Brown and AI Koditex: LLM-Generated Corpora Comparable to Traditional Corpora of English and Czech Texts
- Can AI Perceive Physical Danger and Intervene?
- Correct Reasoning Paths Visit Shared Decision Pivots
- Diagnosing the Performance Trade-off in Moral Alignment: A Case Study on Gender Stereotypes
- PEPS: Quantum-Inspired Reinforcement Learning for Coherent Reasoning Traces in LLMs
- LatentGuard: Controllable Latent Steering for Robust Refusal of Attacks and Reliable Response Generation
- PolicyPad: Collaborative Prototyping of LLM Policies
- Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs
- Who Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM Agents
- BridgeAlign: Bridging Preference Alignment for Humanities and Social Sciences
- RoguePrompt: Dual-Layer Encoding for Self-Reconstruction to Circumvent LLM Moderation
- PanDent: Toward Comprehensive Tooth-Level Structure-Language Consistency in Dental Radiology
- Reinforcement Learning Towards Broadly and Persistently Beneficial Models
- Generative AI for Economic Research: Use Cases and Implications for Economists
- Pressure Reveals Character: Behavioural Alignment Evaluation at Depth
- Code Driven Planning with Domain-Adaptive Critic
- A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users
- Orcust: Stepwise-Feedback Reinforcement Learning for GUI Agent
- Weights-Rotated Preference Optimization for Large Language Models
- nDNA -- the Semantic Helix of Artificial Cognition
- LifeAlign: Lifelong Alignment for Large Language Models with Memory-Augmented Focalized Preference Optimization
- Analyzing the Effects of Supervised Fine-Tuning on Model Knowledge from Token and Parameter Levels
- Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle
- Domain-Specific Constitutional AI: Enhancing Safety in LLM-Powered Mental Health Chatbots
- Self-Rewarding Rubric-Based Reinforcement Learning for Open-Ended Reasoning
- Reward Hacking Mitigation using Verifiable Composite Rewards
- Emergent Alignment via Competition
- Benchmarking and Improving LLM Robustness for Personalized Generation
- Adversarial Distilled Retrieval-Augmented Guarding Model for Online Malicious Intent Detection
- Value Alignment of Social Media Ranking Algorithms
- Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision
- Do LLMs Align Human Values Regarding Social Biases? Judging and Explaining Social Biases with LLMs
- FVDebug: An LLM-Driven Debugging Assistant for Automated Root Cause Analysis of Formal Verification Failures
- A Multi-Agent LLM Defense Pipeline Against Prompt Injection Attacks
- REFINER: Reasoning Feedback on Intermediate Representations
- DoubleAgents: Human-Agent Alignment in a Socially Embedded Workflow
- SSFO: Self-Supervised Faithfulness Optimization for Retrieval-Augmented Generation
- Towards Alignment-Centric Paradigm: A Survey of Instruction Tuning in Large Language Models
- Designing Rules to Pick a Rule: Aggregation by Consistency
- Perception Before Reasoning: Two-Stage Reinforcement Learning for Visual Reasoning in Vision-Language Models
- Towards Safeguarding LLM Fine-tuning APIs against Cipher Attacks
- Co-Alignment: Rethinking Alignment as Bidirectional Human-AI Cognitive Adaptation
- CogniAlign: Survivability-Grounded Multi-Agent Moral Reasoning for Safe and Transparent AI
- Auto-Slides: An Interactive Multi-Agent System for Creating and Customizing Research Presentations
- Evalet: Evaluating Large Language Models through Functional Fragmentation
- Decoding Alignment: A Critical Survey of LLM Development Initiatives through Value-setting and Data-centric Lens
- Being Kind Isn't Always Being Safe: Diagnosing Affective Hallucination in LLMs
- Steering MoE LLMs via Expert (De)Activation
- PromptGuard: An Orchestrated Prompting Framework for Principled Synthetic Text Generation for Vulnerable Populations using LLMs with Enhanced Safety, Fairness, and Controllability
- Building High-Quality Datasets for Portuguese LLMs: From Common Crawl Snapshots to Industrial-Grade Corpora
- X-Teaming Evolutionary M2S: Automated Discovery of Multi-turn to Single-turn Jailbreak Templates
- FinZero: Launching Multi-modal Financial Time Series Forecast with Large Reasoning Model
- HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
- Bias after Prompting: Persistent Discrimination in Large Language Models
- PersonaFuse: A Personality Activation-Driven Framework for Enhancing Human-LLM Interactions
- Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm
- Probabilistic Modeling of Latent Agentic Substructures in Deep Neural Networks
- TIDE: Achieving Balanced Subject-Driven Image Generation via Target-Instructed Diffusion Enhancement
- RL Fine-Tuning Heals OOD Forgetting in SFT
- Uncovering the Vulnerability of Large Language Models in the Financial Domain via Risk Concealment
- EPT Benchmark: Evaluation of Persian Trustworthiness in Large Language Models
- Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal
- RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs
- Symbolic Graphics Programming with Large Language Models
- Manipulating Transformer-Based Models: Controllability, Steerability, and Robust Interventions
- BioBlue: Notable runaway-optimiser-like LLM failure modes on biologically and economically aligned AI safety benchmarks for LLMs with simplified observation format
- On the Evolution of Federated Post-Training Large Language Models: A Model Accessibility View
- EigenBench: A Comparative Behavioral Measure of Value Alignment
- Enhancing Uncertainty Estimation in LLMs with Expectation of Aggregated Internal Belief
- The Resurgence of GCG Adversarial Attacks on Large Language Models
- Transforming Agency. On the mode of existence of Large Language Models
- Balanced Actor Initialization: Stable RLHF Training of Distillation-Based Reasoning Models
- Benchmarking GPT-5 in Radiation Oncology: Measurable Gains, but Persistent Need for Expert Oversight
- ConspirED: A Dataset for Cognitive Traits of Conspiracy Theories and Large Language Model Safety
- AI Chaperones Are (Really) All You Need to Prevent Parasocial Relationships with Chatbots
- Model Science: getting serious about verification, explanation and control of AI systems
- Evaluating Language Model Reasoning about Confidential Information
- Safety Alignment Should Be Made More Than Just A Few Attention Heads
- Survey of Specialized Large Language Model
- Democracy-in-Silico: Institutional Design as Alignment in AI-Governed Polities
- Ensemble Debates with Local Large Language Models for AI Alignment
- Self-Guided Function Calling in Large Language Models via Stepwise Experience Recall
- Goals and the Structure of Experience
- In2x at WMT25 Translation Task
- Zero-knowledge LLM hallucination detection and mitigation through fine-grained cross-model consistency
- Lexical Hints of Accuracy in LLM Reasoning Chains
- CIA+TA Risk Assessment for AI Reasoning Vulnerabilities
- LM Agents May Fail to Act on Their Own Risk Knowledge
- AI Agents for Photonic Integrated Circuit Design Automation
- RAJ-PGA: Reasoning-Activated Jailbreak and Principle-Guided Alignment Framework for Large Reasoning Models
- Reinforcement Learning with Rubric Anchors
- RLNVR: Reinforcement Learning from Non-Verified Real-World Rewards
- From Clicks to Preference: A Multi-stage Alignment Framework for Generative Query Suggestion in Conversational System
- Dataset Construction for Training LLM to Learn Analog Circuit Knowledge
- Feedback-Driven Tool-Use Improvements in Large Language Models via Automated Build Environments
- Efficient Switchable Safety Control in LLMs via Magic-Token-Guided Co-Training
- Fine-grained Video Dubbing Duration Alignment with Segment Supervised Preference Optimization
- From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training
- PersRM-R1: Enhance Personalized Reward Modeling with Reinforcement Learning
- AI Red-Teaming Is a Sociotechnical Problem
- ScamAgents: How AI Agents Can Simulate Human-Level Scam Calls
- Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge
- Let's Measure Information Step-by-Step: LLM-Based Evaluation Beyond Vibes
- Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap
- A DbC Inspired Neurosymbolic Layer for Trustworthy Agent Design
- Can LLMs Generate High-Quality Task-Specific Conversations?
- PentestJudge: Judging Agent Behavior Against Operational Requirements
- CUPID: Evaluating Personalized and Contextualized Alignment of LLMs from Interactions
- A Frame for Communication Control
- What's Taboo for You? - An Empirical Evaluation of LLMs Behavior Toward Sensitive Content
- How Far Are AI Scientists from Changing the World?
Discussions
- Well the problem is I'm not sure I trust your barometer for "fake bullshit" at this point, Harry. At any rate, constitutional AI is a post-training/alignment technique pioneered by Anthropic. Askell i [bsky, 27 points, 1 comments]
- When Lying Is the Best Strategy for AI | HGModernism [lemmy, 15 points, 2 comments]
- this paper pretty much tells you exactly why claude's personality is different from other LLMs. other labs do not seem to have decided to do the same thing. it is a very smart approach. arxiv.org/abs/ [bsky, 14 points, 2 comments]
- Basically an updated version of arxiv.org/abs/2212.08073 [bsky, 7 points, 1 comments]
- this is a 34-page paper with empirical data from 2022 arxiv.org/pdf/2212.08073 and the technique has been well studied by other parties since arxiv.org/search/?sear... [bsky, 2 points, 1 comments]
- You can read about the process here arxiv.org/pdf/2212.08073 [bsky, 1 points, 1 comments]
- A study on Constitutional AI trains harmless assistants using self-improvement from principles instead of human labels. It merges supervised and reinforcement learning, letting AI critique harmful que [bsky, 1 points, 1 comments]
- (and that reason is just that claude is not built around taking feedback directly from users. arxiv.org/pdf/2212.08073) [bsky, 1 points, 1 comments]
- Wholly factually wrong. The constitution is used to train the model. The system prompt is much smaller (and can be modified by the user). The idea of putting the whole constitution as the system promp [bsky, 1 points, 1 comments]
- ... or rather arxiv.org/abs/2212.08073 [bsky, 0 points, 0 comments]
- #AI #LLM Anthropic is among a number of AI companies that focus on building safe models. Its #Claude family of large language models were trained to follow a constitution that stresses human rights an [bsky, 0 points, 0 comments]
- They work on constituonal ai and there is a paper on it (mix of context based fine tuning with verification and RL) arxiv.org/abs/2212.08073 [bsky, 0 points, 0 comments]
Related