When combinations of humans and AI are useful: A systematic review and meta-analysis
2024/05/09 by Michelle Vaccaro, Abdullah Almaatouq, Thomas Malone +1 · 5 voices · 416 citations
Computer Science · Medicine · Psychology · Social Sciences · #Artificial Intelligence in Healthcare and Education #Artificial intelligence #Association (psychology) #Biology #Cognitive psychology #Computer science #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #Human studies #MEDLINE #Machine learning #Medicine #Meta-analysis #Pathology #Psychology #Set (abstract data type) #Systematic review #Web of science #cs.AI #cs.CY #cs.HC
paper · pdf · doi:10.1038/s41562-024-02024-1
published in Nature Human Behaviour 8(12), 2293-2303 (Nature Portfolio)
openalex publication_date 2024/10/28 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
Abstract
Inspired by the increasing use of artificial intelligence (AI) to augment humans, researchers have studied human-AI systems involving different tasks, systems and populations. Despite such a large body of work, we lack a broad conceptual understanding of when combinations of humans and AI are better than either alone. Here we addressed this question by conducting a preregistered systematic review and meta-analysis of 106 experimental studies reporting 370 effect sizes. We searched an interdisciplinary set of databases (the Association for Computing Machinery Digital Library, the Web of Science and the Association for Information Systems eLibrary) for studies published between 1 January 2020 and 30 June 2023. Each study was required to include an original human-participants experiment that evaluated the performance of humans alone, AI alone and human-AI combinations. First, we found that, on average, human-AI combinations performed significantly worse than the best of humans or AI alone (Hedges' g = -0.23; 95% confidence interval, -0.39 to -0.07). Second, we found performance losses in tasks that involved making decisions and significantly greater gains in tasks that involved creating content. Finally, when humans outperformed AI alone, we found performance gains in the combination, but when AI outperformed humans alone, we found losses. Limitations of the evidence assessed here include possible publication bias and variations in the study designs analysed. Overall, these findings highlight the heterogeneity of the effects of human-AI collaboration and point to promising avenues for improving human-AI systems.
Cited by
- Method-enforcing AI for product-service system design: Conformance gains and acceptance tradeoffs in an expert evaluation
- A question of alignment – AI, GenAI and applied linguistics
- How Can AI Augment Access to Justice? Public Defenders' Perspectives on Responsible AI Adoption
- Generative AI at Work
- Artificially intelligent agents in the social and behavioral sciences: A history and outlook
- BrainPilot: Automating Brain Discovery with Agentic Research
- Nonuniformity Principle in Human-AI Coworking
- Align AI to Dynamic Human-AI Workflows
- Introducing AI to an Online Petition Platform Changed Outputs but not Outcomes
- Magentic-UI: Towards Human-in-the-loop Agentic Systems
- Explanations are a Means to an End: Decision Theoretic Explanation Evaluation
- Educational Strategies for Clinical Supervision of Artificial Intelligence Use
- A systematic review of generative AI in education: Empirical insights from a human– AI interaction perspective
- Performance and Metacognition Disconnect when Reasoning in Human-AI Interaction
- AI-assisted teams outperform AI-led teams but not human-only teams in assessing research reproducibility in quantitative social science
- Tree-Based Formalization of Multi-Agent Complementarity in Human-AI Interactions
- Alignment, Exploration, and Novelty in Human-AI Interaction
- Fostering human learning is crucial for boosting human-AI synergy
- Reliable agent engineering should integrate machine-compatible organizational principles
- Everything is Context: Agentic File System Abstraction for Context Engineering
- Humans incorrectly reject confident accusatory AI judgments
- A Meta-Analysis of the Persuasive Power of Large Language Models
- CentaurEval: Benchmarking Human-in-the-Loop Value in Agentic Coding
- Towards Synergistic Teacher-AI Interactions with Generative Artificial Intelligence
- Personality Pairing Improves Human-AI Collaboration
- When Thinking Pays Off: Incentive Alignment for Human-AI Collaboration
- Making LLMs Reliable When It Matters Most: A Five-Layer Architecture for High-Stakes Decisions
- When Machines Join the Moral Circle: The Persona Effect of Generative AI Agents in Collaborative Reasoning
- Human–robot collaboration in surgery at the nexus of knowledge, agency, and ownership
- Human-AI Collaboration with Misaligned Preferences
- Human-AI Complementarity: A Goal for Amplified Oversight
- Who Has The Final Say? Conformity Dynamics in ChatGPT's Selections
- AI as Equalizer or Amplifier? Task Complexity as the Moderating Factor for Human Expertise in Hybrid Intelligence Systems
- Multi-species approaches are necessary to optimize conservation: integrating a genetic algorithm with stochastic models in the Murray–Darling Basin (Australia)
- A meta-analysis on reactions to algorithmic decision-making in human resource management
- AI-based research mentors: Plausible scenarios and ethical issues
- How ensembling AI and public managers improves decision-making
- Unpacking the Paradoxes of Trust in Uncertain Times
- Human-AI Collaborative Uncertainty Quantification
- Partnering with Generative AI: Experimental Evaluation of Human-Led and Model-Led Interaction in Human-AI Co-Creation
- How Do AI Agents Do Human Work? Comparing AI and Human Workflows Across Diverse Occupations
- To Ask or Not to Ask: Learning to Require Human Feedback
- Multi-Hop Question Answering: When Can Humans Help, and Where do They Struggle?
- MetaMuse: Algorithm Generation via Creative Ideation
- Position: Human Factors Reshape Adversarial Analysis in Human-AI Decision-Making Systems
- Toward Complementary Intelligence: Integrating Cognitive and Machine AI
- Empirically derived evaluation requirements for responsible deployments of AI in safety-critical settings
- This human study did not involve human subjects: Validating LLM simulations as behavioral evidence
- AI Error Difficulty Modulates the Effectiveness of Explainability in Decision Support Systems
- Developing a collaborative tool to foster communication in sustainability research
- Through the Lens of Human-Human Collaboration: A Configurable Research Platform for Exploring Human-Agent Collaboration
- DuetUI: A Bidirectional Context Loop for Human-Agent Co-Generation of Task-Oriented Interfaces
- Co-Alignment: Rethinking Alignment as Bidirectional Human-AI Cognitive Adaptation
- Reliability, Embeddedness, and Agency: A Utility-Driven Mathematical Framework for Agent-Centric AI Adoption
- Improving police behavior through artificial intelligence: Pre‐registered experimental results in two large US agencies
- Effect of AI Performance, Risk Perception, and Trust on Human Dependence in Deepfake Detection AI system
- Where is AIED Headed? Key Topics and Emerging Frontiers (2020-2024)
- System 0: Transforming Artificial Intelligence into a Cognitive Extension
- Statistical Tests for Replacing Human Decision Makers with Algorithms
- A Mathematical Framework for AI-Human Integration in Work
- Human-Centered Human-AI Collaboration (HCHAC)
- A Task-Driven Human-AI Collaboration: When to Automate, When to Collaborate, When to Challenge
- The norm architects: How AI delegation norms form before anyone designs them
- Investigating the Effects of LLM Use on Critical Thinking Under Time Constraints: Access Timing and Time Availability
- Human Capital, Not Model Benchmarks, Predicts Hybrid Intelligence in Forecasting
- What Does AI Do for Cultural Interpretation? A Randomized Experiment on Close Reading Poems with Exposure to AI Interpretation
- Human Tool: An MCP-Style Framework for Human-Agent Collaboration
- Feedback by Design: Understanding and Overcoming User Feedback Barriers in Conversational Agents
- Scaffolding Human-AI Collaboration: A Field Experiment on Behavioral Protocols and Cognitive Reframing
- Generative-AI and the transformation of workforce. A job postings-driven analysis
- Co-Designing Collaborative Generative AI Tools for Freelancers
- The Augmentation Trap: AI Productivity and the Cost of Cognitive Offloading
- The Fallback as Signal: Preserved Human Skill, Liability, and Competence Signaling in Credence-Good Markets under Improving AI
- The Future of Work is Blended, Not Hybrid
- Sustainability via LLM Right-sizing
- Evaluating Trust in AI, Human, and Co-produced Feedback Among Undergraduate Students
- Has the Creativity of Large-Language Models peaked? An analysis of inter- and intra-LLM variability
- SciSciGPT: Advancing Human-AI Collaboration in the Science of Science
- Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality
- Examining human-AI interaction in real-world healthcare beyond the laboratory. [europepmc]
- New Perspective on Digital Well-Being by Distinguishing Digital Competency From Dependency: Network Approach. [europepmc]
- A Current Review of Generative AI in Medicine: Core Concepts, Applications, and Current Limitations. [europepmc]
- Beyond the gender data gap: co-creating equitable digital patient twins. [europepmc]
- A large language model improves clinicians' diagnostic performance in complex critical illness cases. [europepmc]
- Empirically derived evaluation requirements for responsible deployments of AI in safety-critical settings. [europepmc]
- AI-Based EMG Reporting: A Randomized Controlled Trial. [europepmc]
- Physicians' Attitudes Toward Artificial Intelligence in Medicine: Mixed Methods Survey and Interview Study. [europepmc]
- Enhanced metagenomic strategies for elucidating the complexities of gut microbiota: a review. [europepmc]
- LLMs outperform outsourced human coders on complex textual analysis. [europepmc]
- Human-large language model collaboration in clinical medicine: a systematic review and meta-analysis. [europepmc]
- The influence, promise, and potential perils of artificial intelligence in veterinary medicine: a call for improved awareness and literacy. [europepmc]
- Comparing artificial intelligence and healthcare professional performance in surgical and interventional video analysis: a systematic review and meta-analysis. [europepmc]
- Large language models enhance diagnostic reasoning of medical students in rheumatology: a randomized controlled trial. [europepmc]
- Warning people about the risk of AI error mitigates human acquisition of AI bias. [europepmc]
- An Entropy-Based Framework for Hybrid Coalitions in Game Theory-Part I: Human Arbitration. [europepmc]
- Global English-language-dominated discourse on artificial intelligence in healthcare: a three-year longitudinal analysis of the #AIinHealthcare movement on X. [europepmc]
- The pediatric AI readiness framework: bridging evidence to practice in pediatric artificial intelligence. [europepmc]
- Intelligent Reasoning Cues: A Framework and Case Study of the Roles of AI Information in Complex Decisions. [europepmc]
- The Impact of Artificial Intelligence Usage on Affective Work Well-Being: A Self-Determination Theory Perspective. [europepmc]
- Ethical and Legal Implications of Implementing AI in Gastrointestinal Endoscopy. [europepmc]
- Human-AI collaboration for dysphagia rehabilitation from effectiveness to implementation complexity: a systematic review. [europepmc]
- Psychological and social factors influencing AI painting tool adoption among young-old adults in the context of creative aging. [europepmc]
- A scoping review on pediatric sepsis prediction technologies in healthcare. [europepmc]
- Algorithm, expert, or both? Evaluating the role of feature selection methods on user preferences and reliance. [europepmc]
Discussions
- When combinations of humans and AI are useful: A systematic review and meta-analysis
(co-author Malone wrote “what makes things fun to learn?” at PARC in early 1980s and led MIT work on group collabo [bsky, 6 points, 0 comments]
- Os traduzco una parte de este artículo tan interesante que me pasó @jordi-a.bsky.social , en el que se analiza el rendimiento de sistemas IA, humanos y el combo humano-IA. Fascinante.
arxiv.org/abs/24 [bsky, 3 points, 1 comments]
- doi.org/10.1038/s415... [bsky, 3 points, 0 comments]
- When combinations of humans and # AI are useful: A systematic review and meta-analysis
(co-author Malone wrote “what makes things fun to learn?” at PARC in early 1980s and led MIT work on group colla [bsky, 2 points, 0 comments]
- Interessante:
arxiv.org/abs/2405.06087 [bsky, 0 points, 1 comments]
Related