2013/10/24 by Thamar Solorio, Solorio, Thamar, Ragib Hasan +3 · 1 voice · 2 citations
Computer Science · Social Sciences · #Authorship Attribution and Profiling #Computation and Language (cs.CL) #Computers and Society (cs.CY) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Natural Language Processing Techniques #Wikis in Education and Collaboration #cs.CL #cs.CR #cs.CY
paper · pdf · doi:10.48550/arxiv.1310.6772
4 pages, under submission at LREC 2014
arxiv created 2013/10/24 · openalex publication_date 2013/10/24 · arxiv published 2013/10/24 · arxiv updated 2013/10/28 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
This paper describes the corpus of sockpuppet cases we gathered from Wikipedia. A sockpuppet is an online user account created with a fake identity for the purpose of covering abusive behavior and/or subverting the editing regulation process. We used a semi-automated method for crawling and curating a dataset of real sockpuppet investigation cases. To the best of our knowledge, this is the first corpus available on real-world deceptive writing. We describe the process for crawling the data and some preliminary results that can be used as baseline for benchmarking research. The dataset will be released under a Creative Commons license from our project website: http://docsig.cis.uab.edu.