Guiding LLM Decision-Making with Fairness Reward Models
2025/07/15 by Zara Hall, Hall, Zara, Melanie Subbiah +7 · 1 citation
Economics, Econometrics and Finance · Social Sciences · #Artificial Intelligence in Law #Digitalization, Law, and Regulation #FOS: Computer and information sciences #Law, Economics, and Judicial Systems #Machine Learning (cs.LG)
paper · pdf · doi:10.48550/arxiv.2507.11344
openalex publication_date 2025/07/15 · openalex created_date 2025/10/08 · openalex updated_date 2026/07/28
Abstract
Large language models are increasingly used to support high-stakes decisions, potentially influencing who is granted bail or receives a loan. Naive chain-of-thought sampling can improve average decision accuracy, but has also been shown to amplify unfair bias. To address this challenge and enable the trustworthy use of reasoning models in high-stakes decision-making, we propose a framework for training a generalizable Fairness Reward Model (FRM). Our model assigns a fairness score to LLM reasoning, enabling the system to down-weight biased trajectories and favor equitable ones when aggregating decisions across reasoning chains. We show that a single Fairness Reward Model, trained on weakly supervised, LLM-annotated examples of biased versus unbiased reasoning, transfers across tasks, domains, and model families without additional fine-tuning. Applied to real-world decision-making tasks including recidivism prediction and social media moderation, we show that our approach consistently improves fairness while matching, or even surpassing, baseline accuracy.
Citations
- DecisionFlow: Advancing Large Language Model as Principled Decision Maker
- From Text to Trust: Empowering AI-assisted Decision Making with Adaptive LLM-powered Analysis
- Towards Effective Discrimination Testing for Generative AI
- Evaluating Gender Bias Transfer between Pre-trained and Prompt-Adapted Language Models
- Exploring Accuracy-Fairness Trade-off in Large Language Models
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Divine LLaMAs: Bias, Stereotypes, Stigmatization, and Emotion Representation of Religion in Large Language Models
- Ask LLMs Directly, "What shapes your bias?": Measuring Social Bias in Large Language Models
- AlphaMath Almost Zero: Process Supervision without Process
- Aggregation of Reasoning: A Hierarchical Framework for Enhancing Answer Selection in Large Language Models
- Prompting Techniques for Reducing Social Bias in LLMs through System 1 and System 2 Cognitive Processes
- Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing
- Measuring Gender and Racial Biases in Large Language Models
- V-STaR: Training Verifiers for Self-Taught Reasoners
- BioXP-0.5B: Explainable Medical-AI via RL-GRPO
- Evaluating Gender Bias in Large Language Models via Chain-of-Thought Prompting
- Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
- Evaluating and Mitigating Discrimination in Language Model Decisions
- "Kelly is a Warm Person, Joseph is a Role Model": Gender Biases in LLM-Generated Reference Letters
- Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training
- Bias and Fairness in Large Language Models: A Survey
- Bias and Fairness in Large Language Models: A Survey
- Gender bias and stereotypes in Large Language Models
- Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment
- Scaling Relationship on Learning Mathematical Reasoning with Large Language Models
- Unveiling Gender Bias in Terms of Profession Across LLMs: Analyzing and Addressing Sociological Implications
- Let's Verify Step by Step
- Evaluation of African American Language Bias in Natural Language Generation
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting
- Fairness-guided Few-shot Prompting for Large Language Models
- On Second Thought, Let's Not Think Step by Step! Bias and Toxicity in Zero-Shot Reasoning
- Solving math word problems with process- and outcome-based feedback
- STaR: Bootstrapping Reasoning With Reasoning
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- BBQ: A Hand-Built Bias Benchmark for Question Answering
- Do Prompt-Based Models Really Understand the Meaning of their Prompts?
- On the Opportunities and Risks of Foundation Models
- On the Dangers of Stochastic Parrots
- Nuanced Metrics for Measuring Unintended Bias with Real Data for Text\n Classification
- Proximal Policy Optimization Algorithms
- Fair prediction with disparate impact: A study of bias in recidivism\n prediction instruments
- Fair Prediction with Disparate Impact: A Study of Bias in Recidivism Prediction Instruments
- Equality of Opportunity in Supervised Learning
Cited by
Related