Diversity-Enhanced Reasoning for Subjective Questions
2025/07/27 by Wang, Yumeng, Fan, Zhiyuan, Liu, Jiayu +2 · 4 citations
#Computation and Language (cs.CL) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2507.20187
Abstract
Large Reasoning Models (LRMs) with long chain-of-thought capabilities, optimized via reinforcement learning with verifiable rewards (RLVR), excel at objective reasoning tasks like mathematical problem solving and code generation. However, RLVR is known for degrading generation diversity, which causes LRMs to fall short on subjective reasoning that has multiple answers depending on different role perspectives. While recent studies recognize the importance of diversity-enhanced training in objective reasoning, limited attention has been given to subjective tasks. In this paper, we find that subjective reasoning can be improved by introducing perspective diversity and token-level diversity, with the former one providing a coherent scaffolding anchored to a real-world stakeholder group and the latter one broadening the answer search space. We propose MultiRole-R1, a diversity-enhanced training framework featuring an unsupervised data construction pipeline that synthesizes reasoning chains incorporating various role perspectives. It also employs reinforcement learning via Group Relative Policy Optimization with reward shaping, taking diversity as a reward signal in addition to verifiable reward. Training on subjective tasks solely, MultiRole-R1 increases the in-domain and out-of-domain accuracy by 14.1% and 7.64%, and even enhances the performance on advanced math reasoning such as AIME 2024. We further show that diversity is a more consistent indicator of accuracy than reasoning length.
Citations
- CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents
- Outcome-based Exploration for LLM Reasoning
- CultureCLIP: Empowering CLIP with Cultural Awareness through Synthetic Images and Contextualized Captions
- Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers
- Mathematical Proof as a Litmus Test: Revealing Failure Modes of Advanced Large Reasoning Models
- Reasoning with Exploration: An Entropy Perspective
- Revisiting Epistemic Markers in Confidence Estimation: Can Markers Accurately Reflect Large Language Models' Uncertainty?
- Let's Reason Formally: Natural-Formal Hybrid Reasoning Enhances LLM's Math Capability
- The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
- AdaCtrl: Towards Adaptive and Controllable Reasoning via Difficulty-Aware Budgeting
- Learning to Reason under Off-Policy Guidance
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
- Weight Ensembling Improves Reasoning in Language Models
- Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
- Hogwild! Inference: Parallel LLM Generation via Concurrent Attention
- Inference-Time Scaling for Generalist Reward Modeling
- Scaling Laws of Synthetic Data for Language Models
- Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models
- DAPO: An Open-Source LLM Reinforcement Learning System at Scale
- Uncovering Gaps in How Humans and LLMs Interpret Subjective Language
- Sampling-Efficient Test-Time Scaling: Self-Estimating the Best-of-N Sampling in Early Decoding
- Embracing Diversity: A Multi-Perspective Approach with Soft Labels
- The Relationship Between Reasoning and Performance in Large Language Models -- o3 (mini) Thinks Harder, Not Longer
- S*: Test Time Scaling for Code Generation
- Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?
- A Systematic Review on the Evaluation of Large Language Models in Theory of Mind Tasks
- Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
- LIMO: Less is More for Reasoning
- s1: Simple test-time scaling
- CALM: Unleashing the Cross-Lingual Self-Aligning Ability of Language Model Question Answering
- Perspective Transition of Large Language Models for Solving Subjective Tasks
- Bias in Large Language Models: Origin, Evaluation, and Mitigation
- A Comparative Study on Reasoning Patterns of OpenAI's o1 Model
- MentalArena: Self-play Training of Language Models for Diagnosis and Treatment of Mental Health Disorders
- Aligning LLMs with Individual Preferences via Interaction
- ReGenesis: LLMs can Grow into Reasoning Generalists via Self-Improvement
- OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data
- Crowd-Calibrator: Can Annotator Disagreement Inform Calibration in Subjective Tasks?
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- Chain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMs
- Mixture-of-Agents Enhances Large Language Model Capabilities
- LLM Discussion: Enhancing the Creativity of Large Language Models via Discussion Framework and Role-Play
- From Persona to Personalization: A Survey on Role-Playing Language Agents
- Annotator-Centric Active Learning for Subjective NLP Tasks
- Reinforcement Learning from Multi-role Debates as Feedback for Bias Mitigation in LLMs
- LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- Standardizing the Measurement of Text Diversity: A Tool and a Comparative Analysis of Scores
- Rethinking the Bounds of LLM Reasoning: Are Multi-Agent Discussions the Key?
- Reasoning in Conversation: Solving Subjective Tasks through Dialogue Simulation for Large Language Models
- BioXP-0.5B: Explainable Medical-AI via RL-GRPO
- Universal Self-Consistency for Large Language Model Generation
- Character-LLM: A Trainable Agent for Role-Playing
- Diversity of Thought Improves Reasoning Abilities of LLMs
- RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models
- Better Zero-Shot Reasoning with Role-Play Prompting
- Towards Measuring the Representation of Subjective Global Opinions in Language Models
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- ChatGPT is fun, but it is not funny! Humor is still challenging Large Language Models
- Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate
- ExpertPrompting: Instructing Large Language Models to be Distinguished Experts
- Improving Factuality and Reasoning in Language Models through Multiagent Debate
- Self-Refine: Iterative Refinement with Self-Feedback
- Large Language Models are Zero-Shot Reasoners
- Continuously Discovering Novel Strategies via Reward-Switching Policy Optimization
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- Training Verifiers to Solve Math Word Problems
- BBQ: A Hand-Built Bias Benchmark for Question Answering
- Aligning AI With Shared Human Values
- Language Models are Few-Shot Learners
- Diversity-Inducing Policy Gradient: Using Maximum Mean Discrepancy to Find a Set of Diverse Policies
- Diversity-Driven Exploration Strategy for Deep Reinforcement Learning
- A Diversity-Promoting Objective Function for Neural Conversation Models
- The Invisible Leash: Why RLVR May or May Not Escape Its Origin
- CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge
Cited by
Related