Learning to summarize from human feedback
2020/09/02 by Nisan Stiennon, Stiennon, Nisan, Long Ouyang +16 · 3 voices · 325 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Multimodal Machine Learning Applications
paper · pdf · doi:10.48550/arxiv.2009.01325
Abstract
As language models become more powerful, training and evaluation are increasingly bottlenecked by the data and metrics used for a particular task. For example, summarization models are often trained to predict human reference summaries and evaluated using ROUGE, but both of these metrics are rough proxies for what we really care about -- summary quality. In this work, we show that it is possible to significantly improve summary quality by training a model to optimize for human preferences. We collect a large, high-quality dataset of human comparisons between summaries, train a model to predict the human-preferred summary, and use that model as a reward function to fine-tune a summarization policy using reinforcement learning. We apply our method to a version of the TL;DR dataset of Reddit posts and find that our models significantly outperform both human reference summaries and much larger models fine-tuned with supervised learning alone. Our models also transfer to CNN/DM news articles, producing summaries nearly as good as the human reference without any news-specific fine-tuning. We conduct extensive analyses to understand our human feedback dataset and fine-tuned models We establish that our reward model generalizes to new datasets, and that optimizing our reward model results in better summaries than optimizing ROUGE according to humans. We hope the evidence from our paper motivates machine learning researchers to pay closer attention to how their training loss affects the model behavior they actually want.
Citations
Cited by
- Learning as Reasoning Unfolds: Progressive Rollout Allocation for Efficient Reinforcement Learning
- When Agents Disagree: The Selection Bottleneck in Multi-Agent LLM Pipelines
- How Fast Can Reward Models Score? A Systems Study of C++ and PyTorch Inference Runtimes for RLHF
- Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels
- ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples
- OR Else: A Differentiable Trust Region for Policy Optimization
- RRPO: Reference-Relative Policy Optimization with Stratified Conditional Rollouts
- AI Value Alignment for Evolving Social Norms
- Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration
- Learning from World Feedback: Why Model Uncertainty Fails as a Risk Signal in Model-Based RL
- Model-Driven Discipline for Multi-Agent LLMs: Requirement-to-Verification Generation of Traceable System Models
- Scaling Evaluation-time Compute with Reasoning Models as Evaluators
- Discrete Action Space as a Prerequisite for GRPO Convergence in Small-Model Continuous Control
- Align AI to Dynamic Human-AI Workflows
- Reliability-Aware LLM Alignment from Inconsistent Human Feedback
- Greed Is Learned: Visible Incentives as Reward-Hacking Triggers
- S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF
- How Open Must Language Models be to Enable Reliable Scientific Inference?
- AI Can Learn Scientific Taste
- Semantic-aware Wasserstein Policy Regularization for Large Language Model Alignment
- Self-Distillation Enables Continual Learning
- Distributional AGI Safety
- The Reasoning Trap: How Enhancing LLM Reasoning Amplifies Tool Hallucination
- Antislop: A Comprehensive Framework for Identifying and Eliminating Repetitive Patterns in Language Models
- Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples
- Understanding Reinforcement Learning for Model Training, and future directions with GRAPE
- K2-Think: A Parameter-Efficient Reasoning System
- Cognitive models can reveal interpretable value trade-offs in language models
- LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing
- REASONING GYM: Reasoning Environments for Reinforcement Learning with Verifiable Rewards
- Breaking the Performance Ceiling in Reinforcement Learning requires Inference Strategies
- Inference-Time Scaling for Generalist Reward Modeling
- Privacy Ripple Effects from Adding or Removing Personal Information in Language Model Training
- The Reward Model Selection Crisis in Personalized Alignment
- UniLumos: Fast and Unified Image and Video Relighting with Physics-Plausible Feedback
- Self-Rewarded Multimodal Coherent Reasoning Across Diverse Visual Domains
- DICE: Discrete Interpretable Comparative Evaluation with Probabilistic Scoring for Retrieval-Augmented Generation
- Agent2World: Learning to Generate Symbolic World Models via Adaptive Multi-Agent Feedback
- Offline-Online Curriculum RL for Multimodal Reasoning
- Epistemic Norms for AI Safety and Alignment Research
- Less Data, Better Alignment: Data-Centric Multi-Evaluator Agreement for Preference Optimization
- Instruction-Tuned Language Models Cannot Sample from Distributions They Can Describe
- On the Opportunities and Risks of Foundation Models
- Towards Robust Reinforcement Learning for Small-Scale Language Model Agents
- Inverse RL Helps Align AI by Imitating Humans
- Frustratingly Simple Black-Box Adaptation of Language Models via Logit Bias
- AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning
- Evaluating LLMs as Interpretable Controllers for Dynamical Systems
- Context Sensitivity Improves Human-Machine Visual Alignment
- A Comedy of Estimators: On KL Regularization in RL Training of LLMs
- Can LLMs Predict Their Own Failures? Self-Awareness via Internal Circuits
- Offline Safe Policy Optimization From Heterogeneous Feedback
- Learning to Reason in LLMs by Expectation Maximization
- Efficient Personalization of Generative Models via Optimal Experimental Design
- AdvJudge-Zero: Binary Decision Flips in LLM-as-a-Judge via Adversarial Control Tokens
- Stackelberg Learning from Human Feedback: Preference Optimization as a Sequential Game
- Autoregressive Language Models are Secretly Energy-Based Models: Insights into the Lookahead Capabilities of Next-Token Prediction
- Super Suffixes: Bypassing Text Generation Alignment and Guard Models Simultaneously
- Let's (not) just put things in Context: Test-Time Training for Long-Context LLMs
- Explainable reinforcement learning from human feedback to improve alignment
- RoleRMBench & RoleRM: Towards Reward Modeling for Profile-Based Role Play in Dialogue Systems
- Intelligently Weighting Multiple Reference Models for Direct Preference Optimization of LLMs
- Reference Recommendation based Membership Inference Attack against Hybrid-based Recommender Systems
- Targeting Misalignment: A Conflict-Aware Framework for Reward-Model-based LLM Alignment
- Fluent Alignment with Disfluent Judges: Post-training for Lower-resource Languages
- Parent-Guided Semantic Reward Model (PGSRM): Embedding-Based Reward Functions for Reinforcement Learning of Transformer Language Models
- When Distance Distracts: Representation Distance Bias in BT-Loss for Reward Models
- ARCANE: A Multi-Agent Framework for Interpretable and Configurable Alignment
- Mitigating Self-Preference by Authorship Obfuscation
- Reflection-Satisfaction Tradeoff: Investigating Impact of Reflection on Student Engagement with AI-Generated Programming Hints
- YingMusic-Singer: Zero-shot Singing Voice Synthesis and Editing with Annotation-free Melody Guidance
- Natural Language Actor-Critic: Scalable Off-Policy Learning in Language Space
- On the Limits of Test-Time Compute: Sequential Reward Filtering for Better Inference
- Towards better dense rewards in Reinforcement Learning Applications
- Overcoming State Inertia: Minimally Invasive Temporal Alignment for Evolving Contexts
- Generative Multi-modal Feedback for Singing Voice Synthesis Evaluation
- Zero-Overhead Introspection for Adaptive Test-Time Compute
- Tracing How Annotators Think: Augmenting Preference Judgments with Reading Processes
- Ambiguity Awareness Optimization: Towards Semantic Disambiguation for Direct Preference Optimization
- Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards
- Learning the MPC objective function from human preferences
- Bootstrapping LLMs via Preference-Based Policy Optimization
- GEM: Generative Entropy-Guided Preference Modeling for Few-shot Alignment of LLMs
- MapReduce LoRA: Advancing the Pareto Front in Multi-Preference Optimization for Generative Models
- DRAFT-RL: Multi-Agent Chain-of-Draft Reasoning for Reinforcement Learning-Enhanced LLMs
- Learning to Clean: Reinforcement Learning for Noisy Label Correction
- Optimizing LLM Code Suggestions: Feedback-Driven Timing with Lightweight State Bounds
- Learning What to Trust: Bayesian Prior-Guided Optimization for Visual Generation
- Exploring Weak-to-Strong Generalization for CLIP-based Classification
- PrefixGPT: Prefix Adder Optimization by a Generative Pre-trained Transformer
- SPINE: Token-Selective Test-Time Reinforcement Learning with Entropy-Band Regularization
- Alignment Faking - the Train -> Deploy Asymmetry: Through a Game-Theoretic Lens with Bayesian-Stackelberg Equilibria
- PARROT: Persuasion and Agreement Robustness Rating of Output Truth -- A Sycophancy Robustness Benchmark for LLMs
- Multi-Agent Collaborative Reward Design for Enhancing Reasoning in Reinforcement Learning
- Multi-Faceted Attack: Exposing Cross-Model Vulnerabilities in Defense-Equipped Vision-Language Models
- A Mathematical Framework for Custom Reward Functions in Job Application Evaluation using Reinforcement Learning
- STEP: Success-Rate-Aware Trajectory-Efficient Policy Optimization
- Prompt-Driven Domain Adaptation for End-to-End Autonomous Driving via In-Context RL
- Detecting LLM-Assisted Academic Dishonesty using Keystroke Dynamics
- Probing Preference Representations: A Multi-Dimensional Evaluation and Analysis Method for Reward Models
- Mitigating Length Bias in RLHF through a Causal Lens
- MetaGDPO: Alleviating Catastrophic Forgetting with Metacognitive Knowledge through Group Direct Preference Optimization
- EARL: Entropy-Aware RL Alignment of LLMs for Reliable RTL Code Generation
- Context-Emotion Aware Therapeutic Dialogue Generation: A Multi-component Reinforcement Learning Approach to Language Models for Mental Health Support
- AdvancedIF: Rubric-Based Benchmarking and Reinforcement Learning for Advancing LLM Instruction Following
- Uncertainty-Guided Checkpoint Selection for Reinforcement Finetuning of Large Language Models
- Hail to the Thief: Exploring Attacks and Defenses in Decentralised GRPO
- Environment Scaling for Interactive Agentic Experience Collection: A Survey
- Convergence and Stability Analysis of Self-Consuming Generative Models with Heterogeneous Human Curation
- DeCoRL: Decoupling Reasoning Chains via Parallel Sub-Step Generation and Cascaded Reinforcement for Interpretable and Scalable RLHF
- SERL: Self-Examining Reinforcement Learning on Open-Domain
- PC-Diffusion: Aligning Diffusion Models with Human Preferences via Preference Classifier
- SpeechJudge: Towards Human-Level Judgment for Speech Naturalness
- MathSE: Improving Multimodal Mathematical Reasoning via Self-Evolving Iterative Reflection and Reward-Guided Fine-Tuning
- Chain-of-Thought as a Lens: Evaluating Structured Reasoning Alignment between Human Preferences and Large Language Models
- KG-DF: A Black-box Defense Framework against Jailbreak Attacks Based on Knowledge Graphs
- FLEX: Continuous Agent Evolution via Forward Learning from Experience
- Multi-Reward GRPO Fine-Tuning for De-biasing Large Language Models: A Study Based on Chinese-Context Discrimination Data
- The MineRL BASALT Competition on Learning from Human Feedback
- OckBench: Measuring the Efficiency of LLM Reasoning
- Reflective Personalization Optimization: A Post-hoc Rewriting Framework for Black-Box Large Language Models
- Black-Box Guardrail Reverse-engineering Attack
- Learning Without Critics? Revisiting GRPO in Classical Reinforcement Learning Environments
- The Realignment Problem: When Right becomes Wrong in LLMs
- DetectiumFire: A Comprehensive Multi-modal Dataset Bridging Vision and Language for Fire Understanding
- Inference-Time Personalized Alignment with a Few User Preference Queries
- RLAC: Reinforcement Learning with Adversarial Critic for Free-Form Generation Tasks
- Reasoning Planning for Language Models
- G2: Guided Generation for Enhanced Output Diversity in LLMs
- Iterative Foundation Model Fine-Tuning on Multiple Rewards
- Offline Clustering of Preference Learning with Active-data Augmentation
- Approximating Human Preferences Using a Multi-Judge Learned System
- RecoReward: Recommender-Guided Multimodal Description Generation for Recommendation
- Generating Self-Contained and Summary-Centric Question Answer Pairs via Differentiable Reward Imitation Learning
- Learning Dynamic User Personas from Implicit Interaction Streams via Iterative Refinement
- Post-Training at the Edge of Detectability: A Game-Theoretic Approach to Fine-Tuning
- Take Goodhart Seriously: Principled Limit on General-Purpose AI Optimization
- Reward Models are Metrics in a Trench Coat
- MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes
- Greedy Sampling Is Provably Efficient for RLHF
- Can LLMs Write Faithfully? An Agent-Based Evaluation of LLM-generated Islamic Content
- Latent Chain-of-Thought for Visual Reasoning
- Debiasing Reward Models by Representation Learning with Guarantees
- Think Twice: Branch-and-Rethink Reasoning Reward Model
- The Best of N Worlds: Aligning Reinforcement Learning with Best-of-N Sampling via max@k Optimisation
- Adaptive Blockwise Search: Inference-Time Alignment for Large Language Models
- Aligning Diffusion Language Models via Unpaired Preference Optimization
- Scalable Oversight via Partitioned Human Supervision
- Beyond Reasoning Gains: Mitigating General Capabilities Forgetting in Large Reasoning Models
- Weak-to-Strong Generalization under Distribution Shifts
- PanicToCalm: A Proactive Counseling Agent for Panic Attacks
- Beyond Pairwise: Empowering LLM Alignment With Ranked Choice Modeling
- Self-Rewarding PPO: Aligning Large Language Models with Demonstrations Only
- Learning Grouped Lattice Vector Quantizers for Low-Bit LLM Compression
- Compress to Impress: Efficient LLM Adaptation Using a Single Gradient Step on 100 Samples
- Ask a Strong LLM Judge when Your Reward Model is Uncertain
- Every Question Has Its Own Value: Reinforcement Learning with Explicit Human Values
- Rectifying Shortcut Behaviors in Preference-based Reward Learning
- Preference-based Reinforcement Learning beyond Pairwise Comparisons: Benefits of Multiple Options
- ADPO: Anchored Direct Preference Optimization
- Adaptive Divergence Regularized Policy Optimization for Fine-tuning Generative Models
- Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains
- A Comprehensive Survey on Reinforcement Learning-based Agentic Search: Foundations, Roles, Optimizations, Evaluations, and Applications
- Dual-Weighted Reinforcement Learning for Generative Preference Modeling
- Direct Preference Optimization with Unobserved Preference Heterogeneity: The Necessity of Ternary Preferences
- Reinforcement Learning for Unsupervised Domain Adaptation in Spatio-Temporal Echocardiography Segmentation
- Stop Reducing Responsibility in LLM-Powered Multi-Agent Systems to Local Alignment
- Training LLM Agents to Empower Humans
- Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation
- Putting on the Thinking Hats: A Survey on Chain of Thought Fine-tuning from the Perspective of Human Reasoning Mechanism
- Information-Theoretic Reward Modeling for Stable RLHF: Detecting and Mitigating Reward Hacking
- Beyond Static LLM Policies: Imitation-Enhanced Reinforcement Learning for Recommendation
- How Well Can Preference Optimization Generalize Under Noisy Feedback?
- From Literal to Liberal: A Meta-Prompting Framework for Eliciting Human-Aligned Exception Handling in Large Language Models
- Guarding the Guardrails: A Taxonomy-Driven Approach to Jailbreak Detection
- Finite-time Convergence Analysis of Actor-Critic with Evolving Reward
- Towards Inference-time Scaling for Continuous Space Reasoning
- Precise Attribute Intensity Control in Large Language Models via Targeted Representation Editing
- Reliable Fine-Grained Evaluation of Natural Language Math Proofs
- Don't Walk the Line: Boundary Guidance for Filtered Generation
- DocReward: A Document Reward Model for Structuring and Stylizing
- Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models
- A-IPO: Adaptive Intent-driven Preference Optimization
- Understanding and Exploiting Weight Update Sparsity for Communication-Efficient Distributed RL
- Token Is All You Price
- Opponent Shaping in LLM Agents
- Contrastive Weak-to-strong Generalization
- xRouter: Training Cost-Aware LLMs Orchestration System via Reinforcement Learning
- From Data to Rewards: a Bilevel Optimization Perspective on Maximum Likelihood Estimation
- Textual interpretation of transient image classifications from large language models
- Predictive Preference Learning from Human Interventions
- Incremental Summarization for Customer Support via Progressive Note-Taking and Agent Feedback
- Agent-in-the-Loop: A Data Flywheel for Continuous Improvement in LLM-based Customer Support
- Aligning Large Language Models via Fully Self-Synthetic Data
- Rethinking KL Regularization in RLHF: From Value Estimation to Gradient Optimization
- Towards Better Optimization For Listwise Preference in Diffusion Models
- Online Rubrics Elicitation from Pairwise Comparisons
- Incoherence in Goal-Conditioned Autoregressive Models
- Reward Model Perspectives: Whose Opinions Do Reward Models Reward?
- The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives
- On the Role of Difficult Prompts in Self-Play Preference Optimization
- Beyond Monolithic Rewards: A Hybrid and Multi-Aspect Reward Optimization for MLLM Alignment
- Margin Adaptive DPO: Leveraging Reward Model for Granular Control in Preference Optimization
- Test-Time Scaling in Diffusion LLMs via Hidden Semi-Autoregressive Experts
- TROLL: Trust Regions improve Reinforcement Learning for Large Language Models
- Decoupling Task-Solving and Output Formatting in LLM Generation
- FrameOracle: Learning What to See and How Much to See in Videos
- Best-of-Majority: Minimax-Optimal Strategy for Pass@k Inference Scaling
- Fine-Tuning Diffusion Models via Intermediate Distribution Shaping
- Evaluating Large Language Models Trained on Code
- Limited Preference Data? Learning Better Reward Model with Latent Space Synthesis
- Generative Value Conflicts Reveal LLM Priorities
- Structural Reward Model: Enhancing Interpretability, Efficiency, and Scalability in Reward Modeling
- The Era of Real-World Human Interaction: RL from User Conversations
- PhysicsMinions: Winning Gold Medals in the Latest Physics Olympiads with a Coevolutionary Multimodal Multi-Agent System
- T-POP: Test-Time Personalization with Online Preference Feedback
- Reference-Free Rating of LLM Responses via Latent Information
- MedMMV: A Controllable Multimodal Multi-Agent Framework for Reliable and Verifiable Clinical Reasoning
- Humanline: Online Alignment as Perceptual Loss
- Advantage Weighted Matching: Aligning RL with Pretraining in Diffusion Models
- Which course? Discourse! Teaching Discourse and Generation in the Era of LLMs
- Toward Preference-aligned Large Language Models via Residual-based Model Steering
- Anchored Supervised Fine-Tuning
- Why Alignment Must Precede Distillation: A Minimal Working Explanation
- Clean First, Align Later: Benchmarking Preference Data Cleaning for Reliable LLM Alignment
- Large-Scale Constraint Generation -- Can LLMs Parse Hundreds of Constraints?
- Alignment through Meta-Weighted Online Sampling: Bridging the Gap between Data Generation and Preference Optimization
- Tree Reward-Aligned Search for TReASURe in Masked Diffusion Language Models
- General Exploratory Bonus for Optimistic Exploration in RLHF
- Adaptive Margin RLHF via Preference over Preferences
- Causally-Enhanced Reinforcement Policy Optimization
- Adaptive Policy Backbone via Shared Network
- SoK: Potentials and Challenges of Large Language Models for Reverse Engineering
- Rethinking Large Language Model Distillation: A Constrained Markov Decision Process Perspective
- Who's Laughing Now? An Overview of Computational Humour Generation and Explanation
- Failure Modes of Maximum Entropy RLHF
- PEPS: Quantum-Inspired Reinforcement Learning for Coherent Reasoning Traces in LLMs
- Embedding Domain Knowledge for Large Language Models via Reinforcement Learning from Augmented Generation
- Scaling Laws for Transfer
- BridgeAlign: Bridging Preference Alignment for Humanities and Social Sciences
- Bridging Inference-Time Scaling and Episodic Memory with Action-Centric Graphs
- Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation
- Co-FactChecker: A Framework for Human-AI Collaborative Claim Verification Using Large Reasoning Models
- References Improve LLM Alignment in Non-Verifiable Domains
- GEPO: Group Expectation Policy Optimization for Stable Heterogeneous Reinforcement Learning
- A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users
- Weights-Rotated Preference Optimization for Large Language Models
- LAD-VF: LLM-Automatic Differentiation Enables Fine-Tuning-Free Robot Planning from Formal Methods Feedback
- Asking a Language Model for Diverse Responses
- Multimodal Prompt Decoupling Attack on the Safety Filters in Text-to-Image Models
- Captioning for Text-Video Retrieval via Dual-Group Direct Preference Optimization
- Control the Temperature: Selective Sampling for Diverse and High-Quality LLM Outputs
- Generalizability of Large Language Model-Based Agents: A Comprehensive Survey
- Self-Improving Embodied Foundation Models
- MUSE: MCTS-Driven Red Teaming Framework for Enhanced Multi-Turn Dialogue Safety in Large Language Models
- Rethinking the Evaluation of Alignment Methods: Insights into Diversity, Generalisation, and Safety
- Recursively Summarizing Books with Human Feedback
- Truthful AI: Developing and governing AI that does not lie
- Pluralistic Off-policy Evaluation and Alignment
- ToolRM: Outcome Reward Models for Tool-Calling Large Language Models
- Pathological Truth Bias in Vision-Language Models
- Auto-Slides: An Interactive Multi-Agent System for Creating and Customizing Research Presentations
- Pluralistic Alignment for Healthcare: A Role-Driven Framework
- Decoding Alignment: A Critical Survey of LLM Development Initiatives through Value-setting and Data-centric Lens
- Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning
- Topic-Guided Reinforcement Learning with LLMs for Enhancing Multi-Document Summarization
- SCoder: Iterative Self-Distillation for Bootstrapping Small-Scale Data Synthesizers to Empower Code LLMs
- Uncovering Scaling Laws for Large Language Models via Inverse Problems
- Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs
- IntrEx: A Dataset for Modeling Engagement in Educational Conversations
- BranchGRPO: Stable and Efficient GRPO with Structured Branching in Diffusion Models
- Post-training Large Language Models for Diverse High-Quality Responses
- Guiding Diffusion Models with Reinforcement Learning for Stable Molecule Generation
- Towards a Unified View of Large Language Model Post-Training
- RL's Razor: Why Online Reinforcement Learning Forgets Less
- SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences
- OPERA: A Reinforcement Learning--Enhanced Orchestrated Planner-Executor Architecture for Reasoning-Oriented Multi-Hop Retrieval
- Loong: Synthesize Long Chain-of-Thoughts at Scale through Verifiers
- GRAM-R2: Self-Training Generative Foundation Reward Models for Reward Reasoning
- FantasyHSI: Video-Generation-Centric 4D Human Synthesis In Any Scene through A Graph-based Multi-Agent Framework
- Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models
- PiCSAR: Probabilistic Confidence Selection And Ranking for Reasoning Chains
- Reasoning-Intensive Regression
- Igniting Creative Writing in Small Language Models: LLM-as-a-Judge versus Multi-Agent Refined Rewards
- Learning to Generate Unit Test via Adversarial Reinforcement Learning
- SDGO: Self-Discrimination-Guided Optimization for Consistent Safety in Large Language Models
- ReSURE: Regularizing Supervision Unreliability for Multi-turn Dialogue Fine-tuning
- HEAL: A Hypothesis-Based Preference-Aware Analysis Framework
- Learning to Refine: Self-Refinement of Parallel Reasoning in LLMs
- RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing
- Learning from Few Samples: A Novel Approach for High-Quality Malcode Generation
- CARFT: Boosting LLM Reasoning via Contrastive Learning with Annotated Chain-of-Thought-based Reinforced Fine-Tuning
- Open-Universe Assistance Games
- Reward-Shifted Speculative Sampling Is An Efficient Test-Time Weak-to-Strong Aligner
- DEPTH: Hallucination-Free Relation Extraction via Dependency-Aware Sentence Simplification and Two-tiered Hierarchical Refinement
- LLMs and Agentic AI in Insurance Decision-Making: Opportunities and Challenges For Africa
- Your Reward Function for RL is Your Best PRM for Search: Unifying RL and Search-Based TTS
- MAVIS: Multi-Objective Alignment via Value-Guided Inference-Time Search
- Human Feedback Driven Dynamic Speech Emotion Recognition
- Fusing Rewards and Preferences in Reinforcement Learning
- FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation
- Speciesism in AI: Evaluating Discrimination Against Animals in Large Language Models
- Diversity First, Quality Later: A Two-Stage Assumption for Language Model Alignment
- Compass-Thinker-7B Technical Report
- STELAR-VISION: Self-Topology-Aware Efficient Learning for Aligned Reasoning in Vision
- Fine-grained Video Dubbing Duration Alignment with Segment Supervised Preference Optimization
- PersRM-R1: Enhance Personalized Reward Modeling with Reinforcement Learning
- \(X\)-evolve: Solution space evolution powered by large language models
- Learning to Align, Aligning to Learn: A Unified Approach for Self-Optimized Alignment
- Towards Theoretical Understanding of Transformer Test-Time Computing: Investigation on In-Context Linear Regression
- AMFT: Aligning LLM Reasoners by Meta-Learning the Optimal Imitation-Exploration Balance
- Anatomy of a Machine Learning Ecosystem: 2 Million Models on Hugging Face
- PROPS: Progressively Private Self-alignment of Large Language Models
- Inference-Aware Prompt Optimization for Aligning Black-Box Large Language Models
- EvolvR: Self-Evolving Pairwise Reasoning for Story Evaluation to Enhance Generation
- Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future
- Fine-Grained Safety Neurons with Training-Free Continual Projection to Reduce LLM Fine Tuning Risks
- Cooper: Co-Optimizing Policy and Reward Models in Reinforcement Learning for Large Language Models
- Aligning LLMs on a Budget: Inference-Time Alignment with Heuristic Reward Models
- Unveiling Over-Memorization in Finetuning LLMs for Reasoning Tasks
- Towards Trustworthy Multimodal Moderation via Policy-Aligned Reasoning and Hierarchical Labeling
- When Truth Is Overridden: Uncovering the Internal Origins of Sycophancy in Large Language Models
- MArgE: Meshing Argumentative Evidence from Multiple Large Language Models for Justifiable Claim Verification
- Censored Sampling for Topology Design: Guiding Diffusion with Human Preferences
- RSPO: Risk-Seeking Policy Optimization for Pass@k and Max@k Metrics in Large Language Models
- CoT-Self-Instruct: Building high-quality synthetic prompts for reasoning and non-reasoning tasks
Discussions
Related