ChatGPT outperforms crowd workers for text-annotation tasks
2023/03/27 by Fabrizio Gilardi, Meysam Alizadeh, Maël Kubli · 1 voice · 196 citations
Computer Science · Social Sciences · #Topic Modeling #Natural Language Processing Techniques #Misinformation and Its Impacts
paper · pdf · doi:10.1073/pnas.2305016120
Abstract
Many NLP applications require manual text annotations for a variety of tasks, notably to train classifiers or evaluate the performance of unsupervised models. Depending on the size and degree of complexity, the tasks may be conducted by crowd workers on platforms such as MTurk as well as trained annotators, such as research assistants. Using four samples of tweets and news articles ( n = 6,183), we show that ChatGPT outperforms crowd workers for several annotation tasks, including relevance, stance, topics, and frame detection. Across the four datasets, the zero-shot accuracy of ChatGPT exceeds that of crowd workers by about 25 percentage points on average, while ChatGPT’s intercoder agreement exceeds that of both crowd workers and trained annotators for all tasks. Moreover, the per-annotation cost of ChatGPT is less than 0.003—about thirty times cheaper than MTurk. These results demonstrate the potential of large language models to drastically increase the efficiency of text classification.
Cited by
- Examining the role of artwork orientation on Picasso painting auction prices
- Going Against the Grain: Climate Change as a Wedge Issue for the Radical Right
- Contextualizing Misinformation: A User-Centric Approach to Linguistic and Topical Patterns in News Consumption
- Who is transitioning to green? Introducing a text-based indicator to measure green skill transferability
- Computational Text Analysis for Building and Testing Social Theory
- Using LLMs for measurement in diplomatic speeches: The evolving politics of UN sanctions in Security Council debates
- Large Language Models for Text Classification: From Zero-Shot Learning to Instruction-Tuning
- From Codebooks to Promptbooks: Extracting Information from Text with Generative Large Language Models
- A tutorial on open-source large language models for behavioral science
- Validating open-source machine translation for quantitative text analysis
- Nostalgia in European Party Politics: A Text-Based Measurement Approach
- Unfit for stranding assessment: a panel-scale multimodal-LLM audit of building-decarbonisation disclosure (BeDA)
- Measuring Politicians' Public Personality Traits using Computational Text Analysis: A Multi-Method Feasibility Study for Agency and Communion
- Seeded Topic Models in Digital Archives: Analyzing Interpretations of Immigration in Swedish Newspapers, 1945–2019
- Updating “The Future of Coding”: Qualitative Coding with Generative Large Language Models
- Simulating narrative identity with large language models: a multi-study examination of linguistic and psychological patterns
- From Assistance to Autonomy -- A Researcher Study on the Potential of AI Support for Qualitative Data Analysis
- REGARD: Regional Affective Differences in Large Language Models
- Trusting sovereign language models as scientific instruments: evidence from Portugal's AMALIA
- "Not in My Backyard": LLMs Uncover Online and Offline Social Biases Against Homelessness
- Auditing Differential Visibility of Political Content on TikTok
- When LLMs Over-Answer: Measuring and Mitigating Quality Issues in LLM-Based Hardware Description Language Question Answering
- Mapping the Narrow Corridor with Large Language Models
- Manufactured Divisiveness: Decomposing the Hostile Content of Seven Social Media Influence Operations
- Design-Based Supervised Learning with Noisy Human Labels
- Validating LLMs in social science: Epistemic threats and emerging norms
- AI Research Agents Narrow Scientific Exploration
- RE-AD: Real-Time Requirement Adherence for Data Labeling
- Large Language Models Reproduce Racial Stereotypes When Used for Text Annotation
- Measuring the State of Open Science in Transportation Using Large Language Models
- Randomness in large language models: What researchers need to know (and report)
- Computational Turing Test Reveals Systematic Differences Between Human and AI Language
- Sycophantic AI Decreases Prosocial Intentions and Promotes Dependence
- What is a protest anyway? Codebook conceptualization is still a first-order concern in LLM-era classification
- Measuring Negative Campaigning across Languages with Large Language Models: A Study of 18 Million Tweets in 19 Countries
- Just Put a Human in the Loop? Investigating LLM-Assisted Annotation for Subjective Tasks
- Simulating Society Requires Simulating Thought
- ELEPHANT: Measuring and understanding social sycophancy in LLMs
- The Art of Audience Engagement: LLM-Based Thin-Slicing of Scientific Talks
- Hostility on Twitter in the aftermath of terror attacks
- Large language models for scientometric mapping of scientific controversy: A validated hybrid AI–Human framework
- Can GPT-4 learn to analyse moves in research article abstracts?
- Elite Political Discourse has Become More Toxic in Western Countries
- Gender disparities in the impact of generative artificial intelligence: Evidence from academia.
- Expertise Elevates AI Usage: Experimental Evidence Comparing Laypeople and Professional Artists
- Instruction-Following Evaluation of Large Vision-Language Models
- Measuring complex constructs in large-scale text with computational social mixed methods
- Mapping (A)Ideology: A Taxonomy of European Parties Using Generative LLMs as Zero-Shot Learners
- Efficiency vs. understanding: a critical examination of ChatGPT’s performance in context-sensitive annotation tasks
- BenCSSmark: Making the Social Sciences Count in LLM Research
- "F*** You Biden": Cross-Partisan Electoral Toxicity on X
- International organizations in national parliamentary debates
- Separating Clicks from Baits: Using Large Language Models to Detect Misleading YouTube Thumbnails
- The simulation of judgment in LLMs
- Beyond the Prompt: An Empirical Study of Cursor Rules
- OLAF: Towards Robust LLM-Based Annotation Framework in Empirical Software Engineering
- Toxicity Ahead: Forecasting Conversational Derailment on GitHub
- Explainable Ethical Assessment on Human Behaviors by Generating Conflicting Social Norms
- Echoes of influence: media systems and the representation of interest groups in artificial intelligence policy debates
- Can GPT replace human raters? Validity and reliability of machine-generated norms for metaphors
- Large Language Models have Chain-of-Affect
- Mining Legal Arguments to Study Judicial Formalism
- Are generative AI text annotations systematically biased?
- Can LLMs Evaluate What They Cannot Annotate? Revisiting LLM Reliability in Hate Speech Detection
- Automated Data Enrichment using Confidence-Aware Fine-Grained Debate among Open-Source LLMs for Mental Health and Online Safety
- Empirical Prompt Engineering for Construct Identification with Large Language Models
- Understanding Down Syndrome Stereotypes in LLM-Based Personas
- A Comparison of Human and ChatGPT Classification Performance on Complex Social Media Data
- MegaChat: A Synthetic Persian Q&A Dataset for High-Quality Sales Chatbot Evaluation
- SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition
- Generative AI in Sociological Research: State of the Discipline
- Applying Large Language Models to Characterize Public Narratives
- Towards Consistent Detection of Cognitive Distortions: LLM-Based Annotation and Dataset-Agnostic Evaluation
- AI in the Pen: How Real-time AI Writing Guidance Shapes Online Reviews
- MTQ-Eval: Multilingual Text Quality Evaluation for Language Models
- Who Owns the Argument? A Practical Test for Author Responsibility in <scp>AI</scp> ‐Assisted Scholarly Writing
- Increasing AI Explainability by LLM Driven Standard Processes
- Can LLM Annotations Replace User Clicks for Learning to Rank?
- Who Is the Story About? Protagonist Entity Recognition in News
- Using language models to label clusters of scientific documents
- Evaluating the Impact of LLM-Assisted Annotation in a Perspectivized Setting: the Case of FrameNet Annotation
- Approximating Human Preferences Using a Multi-Judge Learned System
- A Low-Cost Human-in-the-Loop Investigation of Toxicity on GitHub at Scale
- Nudging Sustainable Choices through LLM-Generated Recommendation Explanations
- Positioning Political Texts with Large Language Models by Asking and Averaging
- Durably reducing conspiracy beliefs through dialogues with AI
- Legitimation Strategies of Transnational Private Institutions: Evidence From the International Organization for Standardization
- Scaling Open-Ended Survey Responses Using LLM-Paired Comparisons
- Who Speaks to Whom? An LLM-Based Social Network Analysis of Tragic Plays
- Generative AI and the augmentation of information practices in knowledge work
- Spelling correction with large language models to reduce measurement error in open-ended survey responses
- Large Language Models Outperform Expert Coders and Supervised Classifiers at Annotating Political Social Media Messages
- Media coverage of climate activist groups in Germany
- Generating units of cultural analysis with large language models: methods and validation for scalable cross-cultural research
- Under the (neighbor)hood: Hyperlocal Surveillance on Nextdoor
- Enhancing Hate Speech Detection with Fine-Tuned Large Language Models Requires High-Quality Data
- The Extended Morality as Cooperation Dictionary (eMACD): A Crowd-Sourced Approach via the Moral Narrative Analyzer Platform
- Can Large Language Models Transform Computational Social Science?
- The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers
- CresOWLve: Benchmarking Creative Problem-Solving Over Real-World Knowledge
- SHIELD: Suppressing Hallucinations In LVLM Encoders via Bias and Vulnerability Defense
- Self-Reflection Protects Behavior from Volatile Beliefs Linked to Paranoia
- Quilt-1M: One Million Image-Text Pairs for Histopathology
- Machine-Assisted Grading of Nationwide School-Leaving Essay Exams with LLMs and Statistical NLP
- LAMUS: A Large-Scale Corpus for Legal Argument Mining from U.S. Caselaw using LLMs
- The use of LLMs to annotate data in management research: Foundational guidelines and warnings
- Generative Large Language Models (gLLMs) in Content Analysis: A Practical Guide for Communication Research
- Large Language Model Agent Personality and Response Appropriateness: Evaluation by Human Linguistic Experts, LLM-as-Judge, and Natural Language Processing Model
- Farmers’ Voices in European Protests: Diverse Complaints, Emotional Tones, and Policy Responses
- The end of experimental research as we know it? A perspective on generative artificial intelligence in communication science
- Prompt selection matters: enhancing text annotations for social sciences with large language models
- Cross-Platform Short-Video Diplomacy: Topic and Sentiment Analysis of China-US Relations on Douyin and TikTok
- Ask a Strong LLM Judge when Your Reward Model is Uncertain
- Black Box Absorption: LLMs Undermining Innovative Ideas
- Algorithmic Fairness in NLP: Persona-Infused LLMs for Human-Centric Hate Speech Detection
- Online In-Context Distillation for Low-Resource Vision Language Models
- Zero‐ and few‐shot prompting of generative large language models provides weak assessment of risk of bias in clinical trials
- Using natural language processing to analyse text data in behavioural science
- Reliability of Large Language Model Generated Clinical Reasoning in Assisted Reproductive Technology: Blinded Comparative Evaluation Study
- Latent Topic Synthesis: Leveraging LLMs for Electoral Ad Analysis
- DPRF: A Generalizable Dynamic Persona Refinement Framework for Optimizing Behavior Alignment Between Personalized LLM Role-Playing Agents and Humans
- MAFA: A Multi-Agent Framework for Enterprise-Scale Annotation with Configurable Task Adaptation
- LiteraryQA: Towards Effective Evaluation of Long-document Narrative QA
- Missing the Margins: A Systematic Literature Review on the Demographic Representativeness of LLMs
- Stable LLM Ensemble: Interaction between Example Representativeness and Diversity
- Repurposing Annotation Guidelines to Instruct LLM Annotators: A Case Study
- FlexAC: Towards Flexible Control of Associative Reasoning in Multimodal Large Language Models
- D-CoDe: Scaling Image-Pretrained VLMs to Video via Dynamic Compression and Question Decomposition
- Populism Meets AI: Advancing Populism Research with LLMs
- Embracing Dialectic Intersubjectivity: Coordination of Differential Perspectives in Content Analysis With LLM Persona Simulation
- Unpacking Discourses on Childbirth and Parenthood in Popular Social Media Platforms Across China, Japan, and South Korea
- The Illusion of Artificial Inclusion
- Limited Preference Data? Learning Better Reward Model with Latent Space Synthesis
- Unspoken Hints: Accuracy Without Acknowledgement in LLM Reasoning
- AI in business research
- Building Benchmarks from the Ground Up: Community-Centered Evaluation of LLMs in Healthcare Chatbot Settings
- Topic modeling of video and image data: a visual semantic unsupervised approach
- Which course? Discourse! Teaching Discourse and Generation in the Era of LLMs
- Googling Politics? Comparing Five Computational Methods to Identify Political and News-related Searches from Web Browser Histories
- From silicon to solutions: AI's impending impact on research and discovery
- MMPlanner: Zero-Shot Multimodal Procedural Planning with Chain-of-Thought Object State Reasoning
- The role of empirical evidence in philosophy
- Private Ownership of Public Trust Wildlife Habitat in Montana, U.S.A. (2004–2023)
- Large language models as first-pass filters for corpus annotation: semantic disambiguation of Galician <i>pobo</i>
- Lessons from complex systems science for AI governance
- Demystifying hashtag hijacking in the public opinion game: attention, narratives, and social bots
- A scoping review of ChatGPT research in accounting and finance
- AI for social science and social science of AI: A survey
- Conducting Qualitative Interviews with AI
- From joy to fear in scientific titles: automated emotion recognition and the citation payoff
- Fake news as a rhetorical weapon: Strategic delegitimization and selective amplification in Italian newspapers
- A cultural explanation for parole decisions in the United States
- Why Biden-era clean energy investment policies had limited political returns
- The Task Space: An Integrative Framework for Team Research
- Political DEBATE: Efficient Zero-shot and Few-shot Classifiers for Political Text
- RelRepair: Enhancing Automated Program Repair by Retrieving Relevant Code
- Building Data-Driven Occupation Taxonomies: A Bottom-Up Multi-Stage Approach via Semantic Clustering and Multi-Agent Collaboration
- Codebook LLMs: Evaluating LLMs as Measurement Tools for Political Science Concepts
- Computational Analysis of Conversation Dynamics through Participant Responsivity
- The Impact of Annotator Personas on LLM Behavior Across the Perspectivism Spectrum
- We Argue to Agree: Towards Personality-Driven Argumentation-Based Negotiation Dialogue Systems for Tourism
- Emulating Public Opinion: A Proof-of-Concept of AI-Generated Synthetic Survey Responses for the Chilean Case
- Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals
- Explicit Reasoning Makes Better Judges: A Systematic Study on Accuracy, Efficiency, and Robustness
- Timing the Message: Language-Based Notifications for Time-Critical Assistive Settings
- PersonaFuse: A Personality Activation-Driven Framework for Enhancing Human-LLM Interactions
- CURE: Controlled Unlearning for Robust Embeddings -- Mitigating Conceptual Shortcuts in Pre-Trained Language Models
- ThumbnailTruth: A Multi-Modal LLM Approach for Detecting Misleading YouTube Thumbnails Across Diverse Cultural Settings
- Evaluating the Robustness of Retrieval-Augmented Generation to Adversarial Evidence in the Health Domain
- Leveraging Media Frames to Improve Normative Diversity in News Recommendations
- LLM-Supported Content Analysis of Motivated Reasoning on Climate Change
- PromptSleuth: Detecting Prompt Injection via Semantic Intent Invariance
- Automated Quality Assessment for LLM-Based Complex Qualitative Coding: A Confidence-Diversity Framework
- Guidelines for Empirical Studies in Software Engineering involving Large Language Models
- Synthetic Replacements for Human Survey Data? The Perils of Large Language Models
- The Promise of Large Language Models in Digital Health: Evidence from Sentiment Analysis in Online Health Communities
- Exploring the Feasibility of LLMs for Automated Music Emotion Annotation
- SYNAPSE-G: Bridging Large Language Models and Graph Learning for Rare Event Classification
- Evaluating Large Language Models as Expert Annotators
- Navigating the Risks of Using Large Language Models for Text Annotation in Social Science Research
- Do Ethical AI Principles Matter to Users? A Large-Scale Analysis of User Sentiment and Satisfaction
- A Confidence-Diversity Framework for Calibrating AI Judgement in Accessible Qualitative Coding Tasks
- CoT-Self-Instruct: Building high-quality synthetic prompts for reasoning and non-reasoning tasks
- What's Taboo for You? - An Empirical Evaluation of LLMs Behavior Toward Sensitive Content
- How Exposed Are UK Jobs to Generative AI? Developing and Applying a Novel Task-Based Index
- Large Language Models Are Democracy Coders with Attitudes
- Introducing HALC: A general pipeline for finding optimal prompting strategies for automated coding with LLMs in the computational social sciences
- Modern Uyghur Dependency Treebank (MUDT): An Integrated Morphosyntactic Framework for a Low-Resource Language
- Doubling Your Data in Minutes: Ultra-fast Tabular Data Generation via LLM-Induced Dependency Graphs
- Hybrid Annotation for Propaganda Detection: Integrating LLM Pre-Annotations with Human Intelligence
- AQuilt: Weaving Logic and Self-Inspection into Low-Cost, High-Relevance Data Synthesis for Specialist LLMs
- VeriMinder: Mitigating Analytical Vulnerabilities in NL2SQL
- Machine-assisted quantitizing designs: augmenting humanities and social sciences with artificial intelligence
- DepreSym: A Depression Symptom Annotated Corpus and the Role of Large Language Models as Assessors of Psychological Markers
- Durably reducing conspiracy beliefs through dialogues with AI
- Re‐Imagining the Epistemic Possibilities of <scp>GPT</scp> for Public Administration Research in Competitive Settings
- Large language models identify immigration attitudes in online discourse regardless of language
Discussions
Related