2026/01/13 by Joshua Claassen, Jan Karem Höhne, Ruben L. Bach +1 · 1 voice · 1 citation
Psychology · Social Sciences · #Encoder #Exploit #Identification (biology) #Mental Health via Writing #Narrative #Open source #Proxy (statistics) #Quality (philosophy) #Social Media in Health Education #Social media #Survey Methodology and Nonresponse
paper · doi:10.1177/08944393251408022
published in Social Science Computer Review (SAGE Publishing)
openalex publication_date 2026/01/13 · openalex created_date 2026/01/14 · openalex updated_date 2026/05/21
Online survey participants are frequently recruited through social media platforms, opt-in online access panels, and river sampling approaches. Such online surveys are threatened by bots that shift survey outcomes and exploit incentives. In this proof-of-concept study, we advance the identification of bots driven by Large Language Models (LLMs) through the prediction of LLM-generated text in open narrative responses. We conducted an online survey on same-gender partnerships, including three open narrative questions, and recruited 1512 participants through Facebook. In addition, we utilized two LLM-driven bots, each of which responded to the open narrative questions 400 times. Open narrative responses synthesized by our bots were labeled as containing LLM-generated text (“yes”). Facebook responses were assigned a proxy label (“unclear”) as they may contain bots themselves. Using this binary label as ground truth, we fine-tuned prediction models relying on the “Bidirectional Encoder Representations from Transformers” (BERT) model, resulting in an impressive prediction performance: The models accurately identified between 97% and 100% of bot responses. However, prediction performance decreases if the models make predictions about questions they were not fine-tuned with. Our study contributes to the ongoing discussion on bots and extends the methodological toolkit for protecting the quality and integrity of online survey data.