HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
2025/09/10 by Benjamin Sturgeon, Daniel Samuelson, Sturgeon, Benjamin +5 · 1 voice · 4 citations
Social Sciences · Medicine · Computer Science · #Ethics and Social Impacts of AI #Artificial Intelligence in Healthcare and Education #Explainable Artificial Intelligence (XAI)
paper · pdf · doi:10.48550/arxiv.2509.08494
Abstract
As humans delegate more tasks and decisions to artificial intelligence (AI), we risk losing control of our individual and collective futures. Relatively simple algorithmic systems already steer human decision-making, such as social media feed algorithms that lead people to unintentionally and absent-mindedly scroll through engagement-optimized content. In this paper, we develop the idea of human agency by integrating philosophical and scientific theories of agency with AI-assisted evaluation methods: using large language models (LLMs) to simulate and validate user queries and to evaluate AI responses. We develop HumanAgencyBench (HAB), a scalable and adaptive benchmark with six dimensions of human agency based on typical AI use cases. HAB measures the tendency of an AI assistant or agent to Ask Clarifying Questions, Avoid Value Manipulation, Correct Misinformation, Defer Important Decisions, Encourage Learning, and Maintain Social Boundaries. We find low-to-moderate agency support in contemporary LLM-based assistants and substantial variation across system developers and dimensions. For example, while Anthropic LLMs most support human agency overall, they are the least supportive LLMs in terms of Avoid Value Manipulation. Agency support does not appear to consistently result from increasing LLM capabilities or instruction-following behavior (e.g., RLHF), and we encourage a shift towards more robust safety and alignment targets.
Citations
- Can Reasoning Help Large Language Models Capture Human Annotator Disagreement?
- LLM Social Simulations Are a Promising Research Method
- Toward an Evaluation Science for Generative AI Systems
- What do Large Language Models Say About Animals? Investigating Risks of Animal Harm in Generated Text
- Multi-turn Evaluation of Anthropomorphic Behaviours in Large Language Models
- Gradual Disempowerment: Systemic Existential Risks from Incremental AI Development
- Evaluating Generative AI Systems is a Social Science Measurement Challenge
- Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations
- Hidden Persuaders: LLMs' Political Leaning and Their Influence on Voters
- The Dark Side of AI Companionship: A Taxonomy of Harmful Algorithmic Behaviors in Human-AI Relationships
- AI can help humans find common ground in democratic deliberation
- Tutor CoPilot: A Human-AI Approach for Scaling Real-Time Expertise
- AI Consciousness and Public Perceptions: Four Futures
- Perceptions of Sentient AI and Other Digital Minds: Evidence from the AI, Morality, and Sentience (AIMS) Survey
- On LLMs-Driven Synthetic Data Generation, Curation, and Evaluation: A Survey
- The Impossibility of Fair LLMs
- The Ethics of Advanced AI Assistants
- An Audit on the Perspectives and Challenges of Hallucinations in NLP
- AI alignment: Assessing the global impact of recommender systems
- Generative AI for Synthetic Data Generation: Methods, Challenges and the Future
- Bias in Language Models: Beyond Trick Tests and Toward RUTEd Evaluation
- The Reasons that Agents Act: Intention and Instrumental Goals
- Two Types of AI Existential Risk: Decisive and Accumulative
- Thousands of AI Authors on the Future of AI
- CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation
- AI Supported Degradation of the Self Concept: A Theoretical Framework Grounded in Established Cognitive and Computational Mechanisms
- Bridging the Gulf of Envisioning: Cognitive Design Challenges in LLM Interfaces
- Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
- FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets
- Towards Measuring the Representation of Subjective Global Opinions in Language Models
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- Intent-aligned AI systems deplete human agency: the need for agency foundations research in AI safety
- On the Risk of Misinformation Pollution with Large Language Models
- Whose Opinions Do Language Models Reflect?
- Towards Reasoning in Large Language Models: A Survey
- Discovering Language Model Behaviors with Model-Written Evaluations
- Affective Coherence Monitoring for Transformer-Based Language Models
- Discovering Agents
- Highly accurate protein structure prediction with AlphaFold
- Artificial Intelligence, Values, and Alignment
- Optimal Policies Tend to Seek Power
- Defining Agency: Individuality, Normativity, Asymmetry, and Spatio-temporality in Action
- KANT ON THE THEORY AND PRACTICE OF AUTONOMY
Cited by
Discussions
Related