vix.ing · top · new · best · stats

Identifying Bots Through LLM-Generated Text in Open Narrative Responses: A Proof-of-Concept Study

2026/01/13 by Joshua Claassen, Jan Karem Höhne, Ruben L. Bach +1 · 1 voice · 1 citation
Psychology · Social Sciences · #Encoder #Exploit #Identification (biology) #Mental Health via Writing #Narrative #Open source #Proxy (statistics) #Quality (philosophy) #Social Media in Health Education #Social media #Survey Methodology and Nonresponse

paper · doi:10.1177/08944393251408022

published in Social Science Computer Review (SAGE Publishing)

openalex publication_date 2026/01/13 · openalex created_date 2026/01/14 · openalex updated_date 2026/05/21

Abstract

Online survey participants are frequently recruited through social media platforms, opt-in online access panels, and river sampling approaches. Such online surveys are threatened by bots that shift survey outcomes and exploit incentives. In this proof-of-concept study, we advance the identification of bots driven by Large Language Models (LLMs) through the prediction of LLM-generated text in open narrative responses. We conducted an online survey on same-gender partnerships, including three open narrative questions, and recruited 1512 participants through Facebook. In addition, we utilized two LLM-driven bots, each of which responded to the open narrative questions 400 times. Open narrative responses synthesized by our bots were labeled as containing LLM-generated text (“yes”). Facebook responses were assigned a proxy label (“unclear”) as they may contain bots themselves. Using this binary label as ground truth, we fine-tuned prediction models relying on the “Bidirectional Encoder Representations from Transformers” (BERT) model, resulting in an impressive prediction performance: The models accurately identified between 97% and 100% of bot responses. However, prediction performance decreases if the models make predictions about questions they were not fine-tuned with. Our study contributes to the ongoing discussion on bots and extends the methodological toolkit for protecting the quality and integrity of online survey data.

Citations

Cited by

Discussions

Related