Learning to summarize from human feedback
2020/09/02 by Nisan Stiennon, Stiennon, Nisan, Long Ouyang +16 · 3 voices · 658 citations
Computer Science · #Artificial intelligence #Automatic summarization #Computer science #Machine learning #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Quality (philosophy) #Reinforcement learning #Task (project management) #Topic Modeling #Transfer of learning #cs.AI #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.2009.01325
published in arXiv (Cornell University) 33, 3008-3021 (Cornell University) · NeurIPS 2020
openalex publication_date 2020/09/02 · openalex created_date 2020/09/08 · arxiv created 2022/02/15 · arxiv updated 2022/02/17 · openalex updated_date 2026/08/05
Abstract
As language models become more powerful, training and evaluation are increasingly bottlenecked by the data and metrics used for a particular task. For example, summarization models are often trained to predict human reference summaries and evaluated using ROUGE, but both of these metrics are rough proxies for what we really care about -- summary quality. In this work, we show that it is possible to significantly improve summary quality by training a model to optimize for human preferences. We collect a large, high-quality dataset of human comparisons between summaries, train a model to predict the human-preferred summary, and use that model as a reward function to fine-tune a summarization policy using reinforcement learning. We apply our method to a version of the TL;DR dataset of Reddit posts and find that our models significantly outperform both human reference summaries and much larger models fine-tuned with supervised learning alone. Our models also transfer to CNN/DM news articles, producing summaries nearly as good as the human reference without any news-specific fine-tuning. We conduct extensive analyses to understand our human feedback dataset and fine-tuned models We establish that our reward model generalizes to new datasets, and that optimizing our reward model results in better summaries than optimizing ROUGE according to humans. We hope the evidence from our paper motivates machine learning researchers to pay closer attention to how their training loss affects the model behavior they actually want.
Citations
Cited by
- Learning as Reasoning Unfolds: Progressive Rollout Allocation for Efficient Reinforcement Learning
- Towards Bridging the Reward-Generation Gap in Direct Alignment Algorithms
- When Agents Disagree: The Selection Bottleneck in Multi-Agent LLM Pipelines
- How Fast Can Reward Models Score? A Systems Study of C++ and PyTorch Inference Runtimes for RLHF
- Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels
- ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples
- OR Else: A Differentiable Trust Region for Policy Optimization
- RRPO: Reference-Relative Policy Optimization with Stratified Conditional Rollouts
- AI Value Alignment for Evolving Social Norms
- Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration
- Learning from World Feedback: Why Model Uncertainty Fails as a Risk Signal in Model-Based RL
- Model-Driven Discipline for Multi-Agent LLMs: Requirement-to-Verification Generation of Traceable System Models
- Scaling Evaluation-time Compute with Reasoning Models as Evaluators
- Discrete Action Space as a Prerequisite for GRPO Convergence in Small-Model Continuous Control
- Align AI to Dynamic Human-AI Workflows
- Reliability-Aware LLM Alignment from Inconsistent Human Feedback
- Greed Is Learned: Visible Incentives as Reward-Hacking Triggers
- S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF
- How Open Must Language Models be to Enable Reliable Scientific Inference?
- AI Can Learn Scientific Taste
- Semantic-aware Wasserstein Policy Regularization for Large Language Model Alignment
- Self-Distillation Enables Continual Learning
- Distributional AGI Safety
- The Reasoning Trap: How Enhancing LLM Reasoning Amplifies Tool Hallucination
- Antislop: A Comprehensive Framework for Identifying and Eliminating Repetitive Patterns in Language Models
- Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples
- Understanding Reinforcement Learning for Model Training, and future directions with GRAPE
- K2-Think: A Parameter-Efficient Reasoning System
- Cognitive models can reveal interpretable value trade-offs in language models
- LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing
- REASONING GYM: Reasoning Environments for Reinforcement Learning with Verifiable Rewards
- Breaking the Performance Ceiling in Reinforcement Learning requires Inference Strategies
- Inference-Time Scaling for Generalist Reward Modeling
- Privacy Ripple Effects from Adding or Removing Personal Information in Language Model Training
- The Reward Model Selection Crisis in Personalized Alignment
- UniLumos: Fast and Unified Image and Video Relighting with Physics-Plausible Feedback
- Self-Rewarded Multimodal Coherent Reasoning Across Diverse Visual Domains
- DICE: Discrete Interpretable Comparative Evaluation with Probabilistic Scoring for Retrieval-Augmented Generation
- Agent2World: Learning to Generate Symbolic World Models via Adaptive Multi-Agent Feedback
- Offline-Online Curriculum RL for Multimodal Reasoning
- Epistemic Norms for AI Safety and Alignment Research
- Less Data, Better Alignment: Data-Centric Multi-Evaluator Agreement for Preference Optimization
- Instruction-Tuned Language Models Cannot Sample from Distributions They Can Describe
- On the Opportunities and Risks of Foundation Models
- Towards Robust Reinforcement Learning for Small-Scale Language Model Agents
- Inverse RL Helps Align AI by Imitating Humans
- Frustratingly Simple Black-Box Adaptation of Language Models via Logit Bias
- AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning
- Evaluating LLMs as Interpretable Controllers for Dynamical Systems
- Context Sensitivity Improves Human-Machine Visual Alignment
- A Comedy of Estimators: On KL Regularization in RL Training of LLMs
- Can LLMs Predict Their Own Failures? Self-Awareness via Internal Circuits
- Offline Safe Policy Optimization From Heterogeneous Feedback
- Learning to Reason in LLMs by Expectation Maximization
- Efficient Personalization of Generative Models via Optimal Experimental Design
- AdvJudge-Zero: Binary Decision Flips in LLM-as-a-Judge via Adversarial Control Tokens
- Stackelberg Learning from Human Feedback: Preference Optimization as a Sequential Game
- Autoregressive Language Models are Secretly Energy-Based Models: Insights into the Lookahead Capabilities of Next-Token Prediction
- Super Suffixes: Bypassing Text Generation Alignment and Guard Models Simultaneously
- Let's (not) just put things in Context: Test-Time Training for Long-Context LLMs
- Explainable reinforcement learning from human feedback to improve alignment
- RoleRMBench & RoleRM: Towards Reward Modeling for Profile-Based Role Play in Dialogue Systems
- Intelligently Weighting Multiple Reference Models for Direct Preference Optimization of LLMs
- Reference Recommendation based Membership Inference Attack against Hybrid-based Recommender Systems
- Targeting Misalignment: A Conflict-Aware Framework for Reward-Model-based LLM Alignment
- Fluent Alignment with Disfluent Judges: Post-training for Lower-resource Languages
- Parent-Guided Semantic Reward Model (PGSRM): Embedding-Based Reward Functions for Reinforcement Learning of Transformer Language Models
- When Distance Distracts: Representation Distance Bias in BT-Loss for Reward Models
- ARCANE: A Multi-Agent Framework for Interpretable and Configurable Alignment
- Mitigating Self-Preference by Authorship Obfuscation
- Reflection-Satisfaction Tradeoff: Investigating Impact of Reflection on Student Engagement with AI-Generated Programming Hints
- YingMusic-Singer: Zero-shot Singing Voice Synthesis and Editing with Annotation-free Melody Guidance
- Natural Language Actor-Critic: Scalable Off-Policy Learning in Language Space
- On the Limits of Test-Time Compute: Sequential Reward Filtering for Better Inference
- Towards better dense rewards in Reinforcement Learning Applications
- Overcoming State Inertia: Minimally Invasive Temporal Alignment for Evolving Contexts
- Generative Multi-modal Feedback for Singing Voice Synthesis Evaluation
- Zero-Overhead Introspection for Adaptive Test-Time Compute
- Tracing How Annotators Think: Augmenting Preference Judgments with Reading Processes
- Ambiguity Awareness Optimization: Towards Semantic Disambiguation for Direct Preference Optimization
- Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards
- Learning the MPC objective function from human preferences
- Bootstrapping LLMs via Preference-Based Policy Optimization
- GEM: Generative Entropy-Guided Preference Modeling for Few-shot Alignment of LLMs
- MapReduce LoRA: Advancing the Pareto Front in Multi-Preference Optimization for Generative Models
- DRAFT-RL: Multi-Agent Chain-of-Draft Reasoning for Reinforcement Learning-Enhanced LLMs
- Learning to Clean: Reinforcement Learning for Noisy Label Correction
- Optimizing LLM Code Suggestions: Feedback-Driven Timing with Lightweight State Bounds
- Learning What to Trust: Bayesian Prior-Guided Optimization for Visual Generation
- Exploring Weak-to-Strong Generalization for CLIP-based Classification
- PrefixGPT: Prefix Adder Optimization by a Generative Pre-trained Transformer
- SPINE: Token-Selective Test-Time Reinforcement Learning with Entropy-Band Regularization
- Alignment Faking - the Train -> Deploy Asymmetry: Through a Game-Theoretic Lens with Bayesian-Stackelberg Equilibria
- PARROT: Persuasion and Agreement Robustness Rating of Output Truth -- A Sycophancy Robustness Benchmark for LLMs
- Multi-Agent Collaborative Reward Design for Enhancing Reasoning in Reinforcement Learning
- Multi-Faceted Attack: Exposing Cross-Model Vulnerabilities in Defense-Equipped Vision-Language Models
- A Mathematical Framework for Custom Reward Functions in Job Application Evaluation using Reinforcement Learning
- STEP: Success-Rate-Aware Trajectory-Efficient Policy Optimization
- Prompt-Driven Domain Adaptation for End-to-End Autonomous Driving via In-Context RL
- Detecting LLM-Assisted Academic Dishonesty using Keystroke Dynamics
- Probing Preference Representations: A Multi-Dimensional Evaluation and Analysis Method for Reward Models
- Mitigating Length Bias in RLHF through a Causal Lens
- MetaGDPO: Alleviating Catastrophic Forgetting with Metacognitive Knowledge through Group Direct Preference Optimization
- EARL: Entropy-Aware RL Alignment of LLMs for Reliable RTL Code Generation
- Context-Emotion Aware Therapeutic Dialogue Generation: A Multi-component Reinforcement Learning Approach to Language Models for Mental Health Support
- AdvancedIF: Rubric-Based Benchmarking and Reinforcement Learning for Advancing LLM Instruction Following
- Uncertainty-Guided Checkpoint Selection for Reinforcement Finetuning of Large Language Models
- Hail to the Thief: Exploring Attacks and Defenses in Decentralised GRPO
- Environment Scaling for Interactive Agentic Experience Collection: A Survey
- Convergence and Stability Analysis of Self-Consuming Generative Models with Heterogeneous Human Curation
- DeCoRL: Decoupling Reasoning Chains via Parallel Sub-Step Generation and Cascaded Reinforcement for Interpretable and Scalable RLHF
- SERL: Self-Examining Reinforcement Learning on Open-Domain
- PC-Diffusion: Aligning Diffusion Models with Human Preferences via Preference Classifier
- SpeechJudge: Towards Human-Level Judgment for Speech Naturalness
- MathSE: Improving Multimodal Mathematical Reasoning via Self-Evolving Iterative Reflection and Reward-Guided Fine-Tuning
- Chain-of-Thought as a Lens: Evaluating Structured Reasoning Alignment between Human Preferences and Large Language Models
- KG-DF: A Black-box Defense Framework against Jailbreak Attacks Based on Knowledge Graphs
- FLEX: Continuous Agent Evolution via Forward Learning from Experience
- Multi-Reward GRPO Fine-Tuning for De-biasing Large Language Models: A Study Based on Chinese-Context Discrimination Data
- The MineRL BASALT Competition on Learning from Human Feedback
- OckBench: Measuring the Efficiency of LLM Reasoning
- Reflective Personalization Optimization: A Post-hoc Rewriting Framework for Black-Box Large Language Models
- Black-Box Guardrail Reverse-engineering Attack
- Learning Without Critics? Revisiting GRPO in Classical Reinforcement Learning Environments
- The Realignment Problem: When Right becomes Wrong in LLMs
- DetectiumFire: A Comprehensive Multi-modal Dataset Bridging Vision and Language for Fire Understanding
- Inference-Time Personalized Alignment with a Few User Preference Queries
- RLAC: Reinforcement Learning with Adversarial Critic for Free-Form Generation Tasks
- Reasoning Planning for Language Models
- G2: Guided Generation for Enhanced Output Diversity in LLMs
- Iterative Foundation Model Fine-Tuning on Multiple Rewards
- Offline Clustering of Preference Learning with Active-data Augmentation
- Approximating Human Preferences Using a Multi-Judge Learned System
- RecoReward: Recommender-Guided Multimodal Description Generation for Recommendation
- Generating Self-Contained and Summary-Centric Question Answer Pairs via Differentiable Reward Imitation Learning
- Learning Dynamic User Personas from Implicit Interaction Streams via Iterative Refinement
- Post-Training at the Edge of Detectability: A Game-Theoretic Approach to Fine-Tuning
- Take Goodhart Seriously: Principled Limit on General-Purpose AI Optimization
- Reward Models are Metrics in a Trench Coat
- MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes
- Greedy Sampling Is Provably Efficient for RLHF
- Can LLMs Write Faithfully? An Agent-Based Evaluation of LLM-generated Islamic Content
- Latent Chain-of-Thought for Visual Reasoning
- Debiasing Reward Models by Representation Learning with Guarantees
- Think Twice: Branch-and-Rethink Reasoning Reward Model
- The Best of N Worlds: Aligning Reinforcement Learning with Best-of-N Sampling via max@k Optimisation
- Adaptive Blockwise Search: Inference-Time Alignment for Large Language Models
- Aligning Diffusion Language Models via Unpaired Preference Optimization
- Towards Scalable Oversight via Partitioned Human Supervision
- Beyond Reasoning Gains: Mitigating General Capabilities Forgetting in Large Reasoning Models
- Weak-to-Strong Generalization under Distribution Shifts
- PanicToCalm: A Proactive Counseling Agent for Panic Attacks
- Beyond Pairwise: Empowering LLM Alignment With Ranked Choice Modeling
- Self-Rewarding PPO: Aligning Large Language Models with Demonstrations Only
- Learning Grouped Lattice Vector Quantizers for Low-Bit LLM Compression
- Compress to Impress: Efficient LLM Adaptation Using a Single Gradient Step on 100 Samples
- Ask a Strong LLM Judge when Your Reward Model is Uncertain
- Every Question Has Its Own Value: Reinforcement Learning with Explicit Human Values
- Rectifying Shortcut Behaviors in Preference-based Reward Learning
- Preference-based Reinforcement Learning beyond Pairwise Comparisons: Benefits of Multiple Options
- ADPO: Anchored Direct Preference Optimization
- Adaptive Divergence Regularized Policy Optimization for Fine-tuning Generative Models
- Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains
- A Comprehensive Survey on Reinforcement Learning-based Agentic Search: Foundations, Roles, Optimizations, Evaluations, and Applications
- Dual-Weighted Reinforcement Learning for Generative Preference Modeling
- Direct Preference Optimization with Unobserved Preference Heterogeneity: The Necessity of Ternary Preferences
- Reinforcement Learning for Unsupervised Domain Adaptation in Spatio-Temporal Echocardiography Segmentation
- Stop Reducing Responsibility in LLM-Powered Multi-Agent Systems to Local Alignment
- Training LLM Agents to Empower Humans
- Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation
- Putting on the Thinking Hats: A Survey on Chain of Thought Fine-tuning from the Perspective of Human Reasoning Mechanism
- Information-Theoretic Reward Modeling for Stable RLHF: Detecting and Mitigating Reward Hacking
- Beyond Static LLM Policies: Imitation-Enhanced Reinforcement Learning for Recommendation
- How Well Can Preference Optimization Generalize Under Noisy Feedback?
- From Literal to Liberal: A Meta-Prompting Framework for Eliciting Human-Aligned Exception Handling in Large Language Models
- Guarding the Guardrails: A Taxonomy-Driven Approach to Jailbreak Detection
- Finite-time Convergence Analysis of Actor-Critic with Evolving Reward
- Towards Inference-time Scaling for Continuous Space Reasoning
- Precise Attribute Intensity Control in Large Language Models via Targeted Representation Editing
- Reliable Fine-Grained Evaluation of Natural Language Math Proofs
- Don't Walk the Line: Boundary Guidance for Filtered Generation
- DocReward: A Document Reward Model for Structuring and Stylizing
- Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models
- A-IPO: Adaptive Intent-driven Preference Optimization
- Understanding and Exploiting Weight Update Sparsity for Communication-Efficient Distributed RL
- Token Is All You Price
- Opponent Shaping in LLM Agents
- Contrastive Weak-to-strong Generalization
- Efficient Preference-Based Reinforcement Learning: Randomized Exploration Meets Experimental Design
- xRouter: Training Cost-Aware LLMs Orchestration System via Reinforcement Learning
- From Data to Rewards: a Bilevel Optimization Perspective on Maximum Likelihood Estimation
- Textual interpretation of transient image classifications from large language models
- Predictive Preference Learning from Human Interventions
- Incremental Summarization for Customer Support via Progressive Note-Taking and Agent Feedback
- Agent-in-the-Loop: A Data Flywheel for Continuous Improvement in LLM-based Customer Support
- Aligning Large Language Models via Fully Self-Synthetic Data
- Rethinking KL Regularization in RLHF: From Value Estimation to Gradient Optimization
- Towards Better Optimization For Listwise Preference in Diffusion Models
- Online Rubrics Elicitation from Pairwise Comparisons
- Incoherence in Goal-Conditioned Autoregressive Models
- Reward Model Perspectives: Whose Opinions Do Reward Models Reward?
- The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives
- On the Role of Difficult Prompts in Self-Play Preference Optimization
- Beyond Monolithic Rewards: A Hybrid and Multi-Aspect Reward Optimization for MLLM Alignment
- Margin Adaptive DPO: Leveraging Reward Model for Granular Control in Preference Optimization
- Test-Time Scaling in Diffusion LLMs via Hidden Semi-Autoregressive Experts
- TROLL: Trust Regions improve Reinforcement Learning for Large Language Models
- Decoupling Task-Solving and Output Formatting in LLM Generation
- FrameOracle: Learning What to See and How Much to See in Videos
- Best-of-Majority: Minimax-Optimal Strategy for Pass@k Inference Scaling
- Fine-Tuning Diffusion Models via Intermediate Distribution Shaping
- Evaluating Large Language Models Trained on Code
- Limited Preference Data? Learning Better Reward Model with Latent Space Synthesis
- Generative Value Conflicts Reveal LLM Priorities
- Structural Reward Model: Enhancing Interpretability, Efficiency, and Scalability in Reward Modeling
- The Era of Real-World Human Interaction: RL from User Conversations
- PhysicsMinions: Winning Gold Medals in the Latest Physics Olympiads with a Coevolutionary Multimodal Multi-Agent System
- T-POP: Test-Time Personalization with Online Preference Feedback
- Reference-Free Rating of LLM Responses via Latent Information
- MedMMV: A Controllable Multimodal Multi-Agent Framework for Reliable and Verifiable Clinical Reasoning
- Humanline: Online Alignment as Perceptual Loss
- Advantage Weighted Matching: Aligning RL with Pretraining in Diffusion Models
- Which course? Discourse! Teaching Discourse and Generation in the Era of LLMs
- Toward Preference-aligned Large Language Models via Residual-based Model Steering
- Anchored Supervised Fine-Tuning
- Why Alignment Must Precede Distillation: A Minimal Working Explanation
- Clean First, Align Later: Benchmarking Preference Data Cleaning for Reliable LLM Alignment
- Large-Scale Constraint Generation -- Can LLMs Parse Hundreds of Constraints?
- Alignment through Meta-Weighted Online Sampling: Bridging the Gap between Data Generation and Preference Optimization
- Tree Reward-Aligned Search for TReASURe in Masked Diffusion Language Models
- General Exploratory Bonus for Optimistic Exploration in RLHF
- Adaptive Margin RLHF via Preference over Preferences
- Causally-Enhanced Reinforcement Policy Optimization
- Adaptive Policy Backbone via Shared Network
- SoK: Potentials and Challenges of Large Language Models for Reverse Engineering
- Rethinking Large Language Model Distillation: A Constrained Markov Decision Process Perspective
- Who's Laughing Now? An Overview of Computational Humour Generation and Explanation
- Failure Modes of Maximum Entropy RLHF
- PEPS: Quantum-Inspired Reinforcement Learning for Coherent Reasoning Traces in LLMs
- Embedding Domain Knowledge for Large Language Models via Reinforcement Learning from Augmented Generation
- Scaling Laws for Transfer
- BridgeAlign: Bridging Preference Alignment for Humanities and Social Sciences
- Bridging Inference-Time Scaling and Episodic Memory with Action-Centric Graphs
- Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation
- Co-FactChecker: A Framework for Human-AI Collaborative Claim Verification Using Large Reasoning Models
- References Improve LLM Alignment in Non-Verifiable Domains
- GEPO: Group Expectation Policy Optimization for Stable Heterogeneous Reinforcement Learning
- A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users
- Weights-Rotated Preference Optimization for Large Language Models
- LAD-VF: LLM-Automatic Differentiation Enables Fine-Tuning-Free Robot Planning from Formal Methods Feedback
- Asking a Language Model for Diverse Responses
- Multimodal Prompt Decoupling Attack on the Safety Filters in Text-to-Image Models
- Captioning for Text-Video Retrieval via Dual-Group Direct Preference Optimization
- Control the Temperature: Selective Sampling for Diverse and High-Quality LLM Outputs
- Generalizability of Large Language Model-Based Agents: A Comprehensive Survey
- Self-Improving Embodied Foundation Models
- MUSE: MCTS-Driven Red Teaming Framework for Enhanced Multi-Turn Dialogue Safety in Large Language Models
- Rethinking the Evaluation of Alignment Methods: Insights into Diversity, Generalisation, and Safety
- Recursively Summarizing Books with Human Feedback
- Truthful AI: Developing and governing AI that does not lie
- Pluralistic Off-policy Evaluation and Alignment
- ToolRM: Outcome Reward Models for Tool-Calling Large Language Models
- Pathological Truth Bias in Vision-Language Models
- Auto-Slides: An Interactive Multi-Agent System for Creating and Customizing Research Presentations
- Pluralistic Alignment for Healthcare: A Role-Driven Framework
- Decoding Alignment: A Critical Survey of LLM Development Initiatives through Value-setting and Data-centric Lens
- Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning
- Topic-Guided Reinforcement Learning with LLMs for Enhancing Multi-Document Summarization
- SCoder: Iterative Self-Distillation for Bootstrapping Small-Scale Data Synthesizers to Empower Code LLMs
- Uncovering Scaling Laws for Large Language Models via Inverse Problems
- Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs
- IntrEx: A Dataset for Modeling Engagement in Educational Conversations
- BranchGRPO: Stable and Efficient GRPO with Structured Branching in Diffusion Models
- Post-training Large Language Models for Diverse High-Quality Responses
- Guiding Diffusion Models with Reinforcement Learning for Stable Molecule Generation
- Towards a Unified View of Large Language Model Post-Training
- RL's Razor: Why Online Reinforcement Learning Forgets Less
- D-LiFT: Improving LLM-based Decompiler Backend via Code Quality-driven Fine-tuning
- SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences
- OPERA: A Reinforcement Learning--Enhanced Orchestrated Planner-Executor Architecture for Reasoning-Oriented Multi-Hop Retrieval
- Loong: Synthesize Long Chain-of-Thoughts at Scale through Verifiers
- GRAM-R2: Self-Training Generative Foundation Reward Models for Reward Reasoning
- FantasyHSI: Video-Generation-Centric 4D Human Synthesis In Any Scene through A Graph-based Multi-Agent Framework
- Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models
- PiCSAR: Probabilistic Confidence Selection And Ranking for Reasoning Chains
- Reasoning-Intensive Regression
- Igniting Creative Writing in Small Language Models: LLM-as-a-Judge versus Multi-Agent Refined Rewards
- Learning to Generate Unit Test via Adversarial Reinforcement Learning
- SDGO: Self-Discrimination-Guided Optimization for Consistent Safety in Large Language Models
- ReSURE: Regularizing Supervision Unreliability for Multi-turn Dialogue Fine-tuning
- HEAL: A Hypothesis-Based Preference-Aware Analysis Framework
- Learning to Refine: Self-Refinement of Parallel Reasoning in LLMs
- RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing
- Learning from Few Samples: A Novel Approach for High-Quality Malcode Generation
- CARFT: Boosting LLM Reasoning via Contrastive Learning with Annotated Chain-of-Thought-based Reinforced Fine-Tuning
- Open-Universe Assistance Games
- Reward-Shifted Speculative Sampling Is An Efficient Test-Time Weak-to-Strong Aligner
- DEPTH: Hallucination-Free Relation Extraction via Dependency-Aware Sentence Simplification and Two-tiered Hierarchical Refinement
- LLMs and Agentic AI in Insurance Decision-Making: Opportunities and Challenges For Africa
- Your Reward Function for RL is Your Best PRM for Search: Unifying RL and Search-Based TTS
- MAVIS: Multi-Objective Alignment via Inference-Time Value-Guided Selection
- Human Feedback Driven Dynamic Speech Emotion Recognition
- Fusing Rewards and Preferences in Reinforcement Learning
- FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation
- DistFlow: A Fully Distributed RL Framework for Scalable and Efficient LLM Post-Training
- Speciesism in AI: Evaluating Discrimination Against Animals in Large Language Models
- Diversity First, Quality Later: A Two-Stage Assumption for Language Model Alignment
- Compass-Thinker-7B Technical Report
- STELAR-VISION: Self-Topology-Aware Efficient Learning for Aligned Reasoning in Vision
- Fine-grained Video Dubbing Duration Alignment with Segment Supervised Preference Optimization
- PersRM-R1: Enhance Personalized Reward Modeling with Reinforcement Learning
- \(X\)-evolve: Solution space evolution powered by large language models
- Learning to Align, Aligning to Learn: A Unified Approach for Self-Optimized Alignment
- Towards Theoretical Understanding of Transformer Test-Time Computing: Investigation on In-Context Linear Regression
- AMFT: Aligning LLM Reasoners by Meta-Learning the Optimal Imitation-Exploration Balance
- Anatomy of a Machine Learning Ecosystem: 2 Million Models on Hugging Face
- PROPS: Progressively Private Self-alignment of Large Language Models
- Inference-Aware Prompt Optimization for Aligning Black-Box Large Language Models
- EvolvR: Self-Evolving Pairwise Reasoning for Story Evaluation to Enhance Generation
- Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future
- Fine-Grained Safety Neurons with Training-Free Continual Projection to Reduce LLM Fine Tuning Risks
- Cooper: Co-Optimizing Policy and Reward Models in Reinforcement Learning for Large Language Models
- Aligning LLMs on a Budget: Inference-Time Alignment with Heuristic Reward Models
- Unveiling Over-Memorization in Finetuning LLMs for Reasoning Tasks
- Towards Trustworthy Multimodal Moderation via Policy-Aligned Reasoning and Hierarchical Labeling
- When Truth Is Overridden: Uncovering the Internal Origins of Sycophancy in Large Language Models
- MArgE: Meshing Argumentative Evidence from Multiple Large Language Models for Justifiable Claim Verification
- Censored Sampling for Topology Design: Guiding Diffusion with Human Preferences
- RSPO: Risk-Seeking Policy Optimization for Pass@k and Max@k Metrics in Large Language Models
- CoT-Self-Instruct: Building high-quality synthetic prompts for reasoning and non-reasoning tasks
- Open-Source Agentic Hybrid RAG Framework for Scientific Literature Review
- Improving Generative Ad Text on Facebook using Reinforcement Learning
- Off-Policy Corrected Reward Modeling for Reinforcement Learning from Human Feedback
- MaPPO: Maximum a Posteriori Preference Optimization with Prior Knowledge
- SDD: Self-Degraded Defense against Malicious Fine-tuning
- PITA: Preference-Guided Inference-Time Alignment for LLM Post-Training
- Leveraging Fine-Tuned Large Language Models for Interpretable Pancreatic Cystic Lesion Feature Extraction and Risk Categorization
- Weak-to-Strong Generalization with Failure Trajectories: A Tree-based Approach to Elicit Optimal Policy in Strong Models
- DxHF: Providing High-Quality Human Feedback for LLM Alignment via Interactive Decomposition
- Datasets and Recipes for Video Temporal Grounding via Reinforcement Learning
- E.A.R.T.H.: Structuring Creative Evolution through Model Error in Generative AI
- High Quality Related Search Query Suggestions using Deep Reinforcement Learning
- V1: Unifying Generation and Self-Verification for Parallel Reasoners
- Influence Functions for Preference Dataset Pruning
- Preference-based Multi-Objective Reinforcement Learning
- URPO: A Unified Reward & Policy Optimization Framework for Large Language Models
- Can External Validation Tools Improve Annotation Quality for LLM-as-a-Judge?
- Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities
- RePO: Replay-Enhanced Policy Optimization
- AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation
- PrefPalette: Personalized Preference Modeling with Latent Attributes
- Learning to summarize user information for personalized reinforcement learning from human feedback
- Granular feedback merits sophisticated aggregation
- Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling
- Grounding 'Grounding' in NLP
- Human-Guided Shade Artifact Suppression in CBCT-to-MDCT Translation via Schrödinger Bridge with Conditional Diffusion
- A Survey on Large Language Models for Mathematical Reasoning
- Multi-Armed Sampling Problem and the End of Exploration
- PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training
- On Monotonicity in AI Alignment
- Ground-Compose-Reinforce: Grounding Language in Agentic Behaviours using Limited Data
- Draft-based Approximate Inference for LLMs
- Compute Requirements for Algorithmic Innovation in Frontier AI Models
- Reinforce LLM Reasoning through Multi-Agent Reflection
- Generalizing while preserving monotonicity in comparison-based preference learning models
- Quantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition Functions
- Reinforcement Learning with Action Chunking
- Stable Preference Optimization: A Bilevel Approach to Catastrophic Preference Shift
- Bradley-Terry and Multi-Objective Reward Modeling Are Complementary
- Robust Multimodal Large Language Models Against Modality Conflict
- Detecting Proxy Gaming in RL and LLM Alignment via Evaluator Stress Tests
- Discrete Diffusion Trajectory Alignment via Stepwise Decomposition
- Evaluating the Robustness of Collaborative Agents
- Exploring a Gamified Personality Assessment Method through Interaction with LLM Agents Embodying Different Personalities
- PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization
- Data Diversification Methods In Alignment Enhance Math Performance In LLMs
- Energy-Based Transformers are Scalable Learners and Thinkers
- Activation Reward Models for Few-Shot Model Alignment
- Towards Decentralized and Sustainable Foundation Model Training with the Edge
- Gradient-Adaptive Policy Optimization: Towards Multi-Objective Alignment of Large Language Models
- SAFER: Probing Safety in Reward Models with Sparse Autoencoder
- TeamCMU at Touché: Adversarial Co-Evolution for Advertisement Integration and Detection in Conversational Search
- Auto-TA: Towards Scalable Automated Thematic Analysis (TA) via Multi-Agent Large Language Models with Reinforcement Learning
- A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents
- Bingo: Boosting Efficient Reasoning of LLMs via Dynamic and Significance-based Reinforcement Learning
- QA-LIGN: Aligning LLMs through Constitutionally Decomposed QA
- Logit-Gap Steering: A Forward-Pass Diagnostic for Alignment Robustness
- Optimizing Conversational Product Recommendation via Reinforcement Learning
- Integrating Large Language Models in Financial Investments and Market Analysis: A Survey
- Video Unlearning via Low-Rank Refusal Vector
- A Survey of Human-in-the-loop for Machine Learning
- BEST-Route: Adaptive LLM Routing with Test-Time Optimal Compute
- A Survey of Continual Reinforcement Learning
- PrefPaint: Enhancing Image Inpainting through Expert Human Feedback
- The Hidden Link Between RLHF and Contrastive Learning
- Training Language Model to Critique for Better Refinement
- Explicit Preference Optimization: No Need for an Implicit Reward Model
- Improving Fairness of Large Language Models in Multi-document Summarization
- LARP: Learner-Agnostic Robust Data Prefiltering
- Position: Machine Learning Conferences Should Establish a "Refutations and Critiques" Track
- Inference-Time Reward Hacking in Large Language Models
- Automatic Prompt Optimization for Knowledge Graph Construction: Insights from an Empirical Study
- RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1
- GUI-Reflection: Empowering Multimodal GUI Models with Self-Reflection Behavior
- Harnessing the Power of Reinforcement Learning for Language-Model-Based Information Retriever via Query-Document Co-Augmentation
- NSFW-Classifier Guided Prompt Sanitization for Safe Text-to-Image Generation
- Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach
- Out of Control -- Why Alignment Needs Formal Control Theory (and an Alignment Control Stack)
- MDSAM:Memory-Driven Sparse Attention Matrix for LVLMs Hallucination Mitigation
- SAFEx: Analyzing Vulnerabilities of MoE-Based LLMs via Stable Safety-critical Expert Identification
- HeuriGym: An Agentic Benchmark for LLM-Crafted Heuristics in Combinatorial Optimization
- Relic: Enhancing Reward Model Generalization for Low-Resource Indic Languages with Few-Shot Examples
- Reranking-based Generation for Unbiased Perspective Summarization
- Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints
- Modeling the One-to-Many Property in Open-Domain Dialogue with LLMs
- Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation
- Reward Model Interpretability via Optimal and Pessimal Tokens
- ASCD: Attention-Steerable Contrastive Decoding for Reducing Hallucination in MLLM
- GRAM: A Generative Foundation Reward Model for Reward Generalization
- From Tool Calling to Symbolic Thinking: LLMs in a Persistent Lisp Metaprogramming Loop
- TimeMaster: Training Time-Series Multimodal LLMs to Reason via Reinforcement Learning
- Human-assisted Robotic Policy Refinement via Action Preference Optimization
- Hone as You Read: A Practical Type of Interactive Summarization
- Min-p, Max Exaggeration: A Critical Analysis of Min-p Sampling in Language Models
- Rethinking DPO: The Role of Rejected Responses in Preference Misalignment
- History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM
- Can LLMs Reconcile Knowledge Conflicts in Counterfactual Reasoning
SPECS: Faster Test-Time Scaling through Speculative Drafts- RL from Physical Feedback: Aligning Large Motion Models with Humanoid Control
- From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment
- Mitigating Hallucination Through Theory-Consistent Symmetric Multimodal Preference Optimization
- Improving Large Language Model Safety with Contrastive Representation Learning
- Training-free LLM Verification via Recycling Few-shot Examples
- Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards
- AnnoDPO: Protein Functional Annotation Learning with Direct Preference Optimization
- Pareto Optimal Code Generation
- Debiasing Online Preference Learning via Preference Feature Preservation
- Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language Models
- Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning Datasets
- A Smooth Sea Never Made a Skilled SAILOR: Robust Imitation via Learning to Search
- Resisting Contextual Interference in RAG via Parametric-Knowledge Reinforcement
- Flattery, Fluff, and Fog: Diagnosing and Mitigating Idiosyncratic Biases in Preference Models
- Beyond RLHF and NLHF: Population-Proportional Alignment under an Axiomatic Framework
- RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation
- Structured Pruning for Diverse Best-of-N Reasoning Optimization
- Putting the Value Back in RL: Better Test-Time Scaling by Unifying LLM Reasoners With Verifiers
- BPO: Revisiting Preference Modeling in Direct Preference Optimization
- Policy-labeled Preference Learning: Is Preference Enough for RLHF?
- Debate, Reflect, and Distill: Multi-Agent Feedback with Tree-Structured Preference Optimization for Efficient Language Model Enhancement
- RewardAnything: Generalizable Principle-Following Reward Models
- Does Thinking More always Help? Mirage of Test-Time Scaling in Reasoning Models
- Boosting Open-Source LLMs for Program Repair via Reasoning Transfer and LLM-Guided Reinforcement Learning
- World Modelling Improves Language Model Agents
- Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective
- From Anger to Joy: How Nationality Personas Shape Emotion Attribution in Large Language Models
- Understanding the Impact of Sampling Quality in Direct Preference Optimization
- daDPO: Distribution-Aware DPO for Distilling Conversational Abilities
- Expanding before Inferring: Enhancing Factuality in Large Language Models through Premature Layers Interpolation
- Corrigibility as a Singular Target: A Vision for Inherently Reliable Foundation Models
- Quantitative LLM Judges
- What Makes LLMs Effective Sequential Recommenders? A Study on Preference Intensity and Temporal Context
- Cycle Consistency as Reward: Learning Image-Text Alignment without Human Preferences
- Large language models can learn and generalize steganographic chain-of-thought under process supervision
- Fodor and Pylyshyn's Legacy: Still No Human-like Systematic Compositionality in Neural Networks
- Beyond RLHF: A Unified Theoretical Framework of Alignment
- Expressive Communication: A Common Framework for Evaluating Developments in Generative Models and Steering Interfaces
- Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment
- SynPO: Synergizing Descriptiveness and Preference Optimization for Video Detailed Captioning
- Doubly Robust Alignment for Large Language Models
- Putting Humans in the Natural Language Processing Loop: A Survey
- Soft Best-of-n Sampling for Model Alignment
- On Symmetric Losses for Robust Policy Optimization with Noisy Preferences
- A Reward-driven Automated Webshell Malicious-code Generator for Red-teaming
- Aligning Protein Conformation Ensemble Generation with Physical Feedback
- MDPO: Multi-Granularity Direct Preference Optimization for Mathematical Reasoning
- Whispers of Many Shores: Cultural Alignment through Collaborative Cultural Expertise
- Knowledge Augmented Complex Problem Solving with Large Language Models: A Survey
- MiCRo: Mixture Modeling and Context-aware Routing for Personalized Preference Learning
- Thompson Sampling in Online RLHF with General Function Approximation
- Fortune: Formula-Driven Reinforcement Learning for Symbolic Table Reasoning in Language Models
- Probability-Consistent Preference Optimization for Enhanced LLM Reasoning
- Towards Reward Fairness in RLHF: From a Resource Allocation Perspective
- Dataset Cartography for Large Language Model Alignment: Mapping and Diagnosing Preference Data
- Document-Level Text Generation with Minimum Bayes Risk Decoding using Optimal Transport
- Proximalized Preference Optimization for Diverse Feedback Types: A Decomposed Perspective on DPO
- Bounded Rationality for LLMs: Satisficing Alignment at Inference-Time
- Continuous Chain of Thought Enables Parallel Exploration and Reasoning
- LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training
- On-Policy RL with Optimal Reward Baseline
- ValueSim: Generating Backstories to Model Individual Value Systems
- Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy
- Decomposing Elements of Problem Solving: What "Math" Does RL Teach?
- Text2Grad: Reinforcement Learning from Natural Language Feedback
- Revisiting Group Relative Policy Optimization: Insights into On-Policy and Off-Policy Training
- What Have We Achieved on Text Summarization?
- Learning Natural Language Generation from Scratch
- Exploring Fluent Query Reformulations with Text-to-Text Transformers and Reinforcement Learning
- The Multilingual Divide and Its Impact on Global AI Safety
- Can Past Experience Accelerate LLM Reasoning?
- LPOI: Listwise Preference Optimization for Vision Language Models
- Unveiling Instruction-Specific Neurons & Experts: An Analytical Framework for LLM's Instruction-Following Capabilities
- Multi-objective Large Language Model Alignment with Hierarchical Experts
- SquareχPO: Differentially Private and Robust χ2-Preference Optimization in Offline Direct Alignment
- RM-R1: Reward Modeling as Reasoning
- Aligning LLMs by Predicting Preferences from User Writing Samples
- Curriculum-RLAIF: Curriculum Alignment with Reinforcement Learning from AI Feedback
- What Can RL Bring to VLA Generalization? An Empirical Study
- Token-level Accept or Reject: A Micro Alignment Approach for Large Language Models
- Leveraging Importance Sampling to Detach Alignment Modules from Large Language Models
- Efficient and Stealthy Jailbreak Attacks via Adversarial Prompt Distillation from LLMs to SLMs
- Preference Optimization by Estimating the Ratio of the Data Distribution
- Learning a Pessimistic Reward Model in RLHF
- Inference-time Alignment in Continuous Space
- Alignment of large language models with constrained learning
- SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning
- FairPO: Robust Preference Optimization for Fair Multi-Label Learning
- Frictional Agent Alignment Framework: Slow Down and Don't Break Things
- Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPO
- SafeDPO: A Simple Approach to Direct Preference Optimization with Enhanced Safety
- Surrogate Signals from Format and Length: Reinforcement Learning for Solving Mathematical Problems without Ground Truth Answers
- ActiveDPO: Active Direct Preference Optimization for Sample-Efficient Alignment
- LLaDA 1.5: Variance-Reduced Preference Optimization for Large Language Diffusion Models
- Incentivizing High-Quality Human Annotations with Golden Questions
- Online Knowledge Distillation with Reward Guidance
- Optimal Transport-Based Token Weighting scheme for Enhanced Preference Optimization
- GRE Suite: Geo-localization Inference via Fine-Tuned Vision-Language Models and Enhanced Reasoning Chains
- Flex-Judge: Text-Only Reasoning Unleashes Zero-Shot Multimodal Evaluators
- metaTextGrad: Automatically optimizing language model optimizers
- MOSLIM:Align with diverse preferences in prompts through reward classification
- OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models
- Reward Model Overoptimisation in Iterated RLHF
- Scalable Valuation of Human Feedback through Provably Robust Model Alignment
- Rethinking Personalized Reward Modeling for LLMs under Preference Heterogeneity via Group-Debiased Federated Learning
- Self-Improving Large Language Models via Progressive Experience Evolution
- From Profiling to Synthesis: Benchmarking Implicit Behavioral Alignment in Personalized LLM Agents
- Start Classifying: Categorical Critics for LLM Reinforcement Learning
- Towards Analyzing and Understanding the Limitations of VAPO: A Theoretical Perspective
- Dynamic Risk Assessments for Offensive Cybersecurity Agents
- Learning to Choose or Choosing to Learn: Best-of-N vs. Supervised Fine-Tuning for Bit String Generation
- Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning
- Self-Rewarding Large Vision-Language Models for Optimizing Prompts in Text-to-Image Generation
- Your Pre-trained LLM is Secretly an Unsupervised Confidence Calibrator
- Latent Principle Discovery for Language Model Self-Improvement
- Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models
- VerifyBench: Benchmarking Reference-based Reward Systems for Large Language Models
- Aligning Dialogue Agents with Global Feedback via Large Language Model Multimodal Reward Decomposition
- Reward Is Enough: LLMs Are In-Context Reinforcement Learners
- RL Tango: Reinforcing Generator and Verifier Together for Language Reasoning
- The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning
- Preference Learning with Lie Detectors can Induce Honesty or Evasion
- Visionary-R1: Mitigating Shortcuts in Visual Reasoning with Reinforcement Learning
- J4R: Learning to Judge with Equivalent Initial State Group Relative Policy Optimization
- Multi-Level Aware Preference Learning: Enhancing RLHF for Complex Multi-Instruction Tasks
- Bias Fitting to Mitigate Length Bias of Reward Model in RLHF
- Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization
- Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs
- Is Active Persona Inference Necessary for Aligning Small Models to Personal Preferences?
- Fractured Chain-of-Thought Reasoning
- SGDPO: Self-Guided Direct Preference Optimization for Language Model Alignment
- Pairwise Calibrated Rewards for Pluralistic Alignment
- RLAP: A Reinforcement Learning Enhanced Adaptive Planning Framework for Multi-step NLP Task Solving
- AdaBoN: Adaptive Best-of-N Alignment
- Stepwise Guided Policy Optimization: Coloring your Incorrect Reasoning in GRPO
- Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models
- A Systematic Analysis of Base Model Choice for Reward Modeling
- Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback
- Group-in-Group Policy Optimization for LLM Agent Training
- Dist2ill: Distributional Distillation for One-Pass Uncertainty Estimation in Large Language Models
- HelpSteer3-Preference: Open Human-Annotated Preference Data across Diverse Tasks and Languages
- BLEUBERI: BLEU is a surprisingly effective reward for instruction following
- Ranked Voting based Self-Consistency of Large Language Models
- Auditable Release Control for Pedagogical Leakage in LLM Tutors
- ADHMR: Aligning Diffusion-based Human Mesh Recovery via Direct Preference Optimization
- Current Limitations of Language Models: What You Need is Retrieval
- WorldPM: Scaling Human Preference Modeling
- Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning
- Card Sorting Simulator: Augmenting Design of Logical Information Architectures with Large Language Models
- InfoPO: On Mutual Information Maximization for Large Language Model Alignment
- Improved Algorithms for Differentially Private Language Model Alignment
- Detecting Prefix Bias in LLM-based Reward Models
- On the Robustness of Reward Models for Language Model Alignment
- DARLR: Dual-Agent Offline Reinforcement Learning for Recommender Systems with Dynamic Reward
- Direct Density Ratio Optimization: A Statistically Consistent Approach to Aligning Large Language Models
- You Only Look One Step: Accelerating Backpropagation in Diffusion Sampling with Gradient Shortcuts
- Sandcastles in the Storm: Revisiting the (Im)possibility of Strong Watermarking
- Margin-Adaptive Confidence Ranking for Reliable LLM Judgement
- REFINE-AF: A Task-Agnostic Framework to Align Language Models via Self-Generated Instructions using Reinforcement Learning from Automated Feedback
- GeoMind: An Agentic Workflow for Lithology Classification with Reasoned Tool Invocation
- Assessing Robustness to Spurious Correlations in Post-Training Language Models
- Tail-Aware Information-Theoretic Bounds for LLM Alignment under Heavy-Tailed Rewards
- ComPO: Preference Alignment via Comparison Oracles
- Semantic Probabilistic Control of Language Models
- CAMOUFLAGE: Exploiting Misinformation Detection Systems Through LLM-driven Adversarial Claim Transformation
- Multi-agents based User Values Mining for Recommendation
- Large Language Model Guided Incentive Aware Reward Design for Cooperative Multi-Agent Reinforcement Learning
- Gradient Regularization Mitigates Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards
- Quo Vadis, World Modeling?
- Do Coding Agents Deceive Us? Detecting and Preventing Cheating via Capped Evaluation with Randomized Tests
- AuroraRL: Fast, Fault-Tolerant, and Cost-Efficient Reinforcement Learning over Decentralized Network
- Positive Alignment: Artificial Intelligence for Human Flourishing
- In LLM Reasoning, there is Irrationality on top of Value Misalignment
- From Precision to Perception: User-Centred Evaluation of Keyword Extraction Algorithms for Internet-Scale Contextual Advertising
- BiasGuard: A Reasoning-enhanced Bias Detection Tool For Large Language Models
- Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR
- On Training in Imagination
- Long-Horizon Q-Learning: Accurate Value Learning via n-Step Inequalities
- Toward Efficient Exploration by Large Language Model Agents
- HyPerAlign: Interpretable Personalized LLM Alignment via Hypothesis Generation
- Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models
- HARVE: Hacking-Aware Reward-Head Vector Editing for Robust Reward Models
- Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases
- Robust AI Evaluation through Maximal Lotteries
- GVPO: Group Variance Policy Optimization for Large Language Model Post-Training
- Beyond Static Snapshots: A Grounded Evaluation Framework for Language Models at the Agentic Frontier
- The State of AI Governance Research: AI Safety and Reliability in Real World Commercial Deployment
- Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
- Evaluating ChatGPT on Medical Information Extraction Tasks: Performance, Explainability and Beyond
- RPO-RAG: Aligning Small LLMs with Relation-aware Preference Optimization for Knowledge Graph Question Answering
- Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing
- Ethical Risks in Deploying Large Language Models: An Evaluation of Medical Ethics Jailbreaking
- Guardrails for trust, safety, and ethical development and deployment of Large Language Models (LLM)
- SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation
- MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory
- TQLite: Multi-LLM Jury Guided Distillation for Real-time MQM Translation Quality Evaluation
- MemoryCPT: An End-to-End Agent Memory Framework for Cost-Performance Trade-off
- Private Direct Preference Optimization for LLM Alignment
- FISH-Tuning: Enhancing PEFT Methods with Fisher Information
- Hidden Underbelly of the Silicon Valley: Algorithmic Exploitation and Health in Data Work Value Chains
- ParetoHqD: Fast Offline Multiobjective Alignment of Large Language Models using Pareto High-quality Data
- A Domain-Based Taxonomy of Jailbreak Vulnerabilities in Large Language Models
- Algorithm Discovery With LLMs: Evolutionary Search Meets Reinforcement Learning
- TTRL: Test-Time Reinforcement Learning
- Establishing Reliability Metrics for Reward Models in Large Language Models
- Contemplative Agent
- In-context Ranking Preference Optimization
- Reinforcement Learning from Multi-level and Episodic Human Feedback
- Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey
- FairSteer: Inference Time Debiasing for LLMs with Dynamic Activation Steering
- LoRe: Personalizing LLMs via Low-Rank Reward Modeling
- Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning
- Governance Challenges in Reinforcement Learning from Human Feedback: Evaluator Rationality and Reinforcement Stability
- Syntactic and Semantic Control of Large Language Models via Sequential Monte Carlo
- Persona-judge: Personalized Alignment of Large Language Models via Token-level Self-judgment
- SPIE: Semantic and Structural Post-Training of Image Editing Diffusion Models with AI feedback
- Evaluating the Diversity and Quality of LLM Generated Content
- PolyAlign: Conditional Human-Distribution Alignment
- DeepTrans: Deep Reasoning Translation via Reinforcement Learning
- Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project
- Better Estimation of the Kullback--Leibler Divergence Between Language Models
- Learning from Reference Answers: Versatile Language Model Alignment without Binary Human Preference Data
- Supervised Optimism Correction: Be Confident When LLMs Are Sure
- Dual-Difficulty Curriculum Learning for Direct Preference Optimization
- Leveraging Robust Optimization for LLM Alignment under Distribution Shifts
- FactGuard: Leveraging Multi-Agent Systems to Generate Answerable and Unanswerable Questions for Enhanced Long-Context LLM Extraction
- Multimedia and Visual Analytics in the Agentic Era
- Adversarial Training of Reward Models
- Lightweight and Direct Document Relevance Optimization for Generative Information Retrieval
- Fast Controlled Generation from Language Models with Adaptive Weighted Rejection Sampling
Discussions
Related