Evaluating Online Labor Markets for Experimental Research: Amazon.com's Mechanical Turk
2012/01/01 by Adam J. Berinsky, Gregory A. Huber, Gabriel S. Lenz +1 · 51 citations
Computer Science · Economics, Econometrics and Finance · Social Sciences · #Experimental Behavioral Economics Studies #Mobile Crowdsensing and Crowdsourcing #Sports Analytics and Performance
paper · pdf · doi:10.1093/pan/mpr057
openalex publication_date 2012/01/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/03
Abstract
We examine the trade-offs associated with using Amazon.com 's Mechanical Turk (MTurk) interface for subject recruitment. We first describe MTurk and its promise as a vehicle for performing low-cost and easy-to-field experiments. We then assess the internal and external validity of experiments performed using MTurk, employing a framework that can be used to evaluate other subject pools. We first investigate the characteristics of samples drawn from the MTurk population. We show that respondents recruited in this manner are often more representative of the U.S. population than in-person convenience samples—the modal sample in published experimental political science—but less representative than subjects in Internet-based panels or national probability samples. Finally, we replicate important published experimental work using MTurk samples.
Citations
Cited by
- Self-reported autonomic symptoms linking childhood maltreatment to relationship outcomes
- Not All Explanations are Created Equal: Investigating the Pitfalls of Current XAI Evaluation
- Meaning Beyond Numbers: Introducing the Plot Staircase to Measure Graphical Preferences
- Do Voters Understand the Benefits of Taxes?
- Why Inequalities Persist: Parties’ (Non)Responses to Economic Inequality, 1970–2020
- Popular financial reporting increases understanding, interest, and trust: Experimental evidence
- Experimental Evidence for Differences in the Prosocial Effects of Binge-Watched Versus Appointment-Viewed Television Programs
- A Diamond is Foreverr and Other Fairy Tales: The Relationship between Wedding Expenses and Marriage Duration
- Examining attitudes toward public participation across sectors: An experimental study of food assistance
- Building, hosting, recruiting: A brief introduction to running behavioral experiments online
- Sample characteristics for quantitative analyses in Body Image: Issues of generalisability
- War on Aisle 5: Casualties, National Identity, and Consumer Behavior
- Text as Data: The Promise and Pitfalls of Automatic Content Analysis Methods for Political Texts
- Electric Vehicle Charging and Car Dependency
- The Demand for Season of Birth
- Entertaining Beliefs in Economic Mobility
- Crowdsourcing Reliable Local Data
- A Large-scale Analysis of the Marketplace Characteristics in Fiverr
- How should sugar-sweetened beverage health warnings be designed? A randomized experiment
- Bridging the Gap between National Weather Service Heat Terminology and Public Understanding
- Real Solutions for Fake News? Measuring the Effectiveness of General Warnings and Fact-Check Tags in Reducing Belief in False Stories on Social Media
- Estimates of Non-Heterosexual Prevalence: The Roles of Anonymity and Privacy in Survey Methodology
- Gender, race, and political ambition: how intersectionality and frames influence interest in political office
- Cultural dispositions and economic choice: How field-specific logics shape ‘rational’ economic behaviour
- Internalised White Ideal, Skin Tone Surveillance, and Hair Surveillance Predict Skin and Hair Dissatisfaction and Skin Bleaching among African American and Indian Women
- Computational Psychiatry in Borderline Personality Disorder
- Revisiting white backlash: Does race affect death penalty opinion?
- The influence of leg-to-body ratio, arm-to-body ratio and intra-limb ratio on male human attractiveness
- Bias Blind Spot: Structure, Measurement, and Consequences
- What Do I Need to Vote? Bureaucratic Discretion and Discrimination by Local Election Officials
- Hostile media bias on social media: Testing the effect of user comments on perceptions of news bias and credibility
- Out of control or right on the money? Funder self-efficacy and crowd bias in equity crowdfunding
- When Playing the Woman Card is Playing Trump: Assessing the Efficacy of Framing Campaigns as Historic
- Who Cares What They Wear? Media, Gender, and the Influence of Candidate Appearance
- Breaking monotony with meaning: Motivation in crowdsourcing markets
- “Who are these people?” Evaluating the demographic characteristics and political preferences of MTurk survey respondents
- Extreme party animals: Effects of political identification and ideological extremity
- You Can Leave Your Glasses on
- Counting polyamorists who count: Prevalence and definitions of an under-researched form of consensual nonmonogamy
- The Gag Reflex: Disgust Rhetoric and Gay Rights in American Politics
- How Black Are Lakisha and Jamal? Racial Perceptions from Names Used in Correspondence Audit Studies
- Misinformation exposure per se may not trigger cynicism, but perceived prevalence of misinformation and hostile information does
- Empathy and the Hostile Media Phenomenon
- Halo Effects and the Attractiveness Premium in Perceptions of Political Expertise
- Public perceptions and acceptance of induced earthquakes related to energy development
- Demonstrating Anticipatory Deflection and a Preemptive Measure to Manage It: An Extension of Affect Control Theory
- Nonlinear analysis of EEG complexity in episode and remission phase of recurrent depression
- Factors affecting social presence and word-of-mouth in corporate social responsibility communication: Tone of voice, message framing, and online medium type
- Public opinion on nuclear energy and nuclear weapons: The attitudinal nexus in the United States
- Anti-Semitism and opposition to Israeli government policies: the roles of prejudice and information
- Saving Media or Trading on Trust?
Related