2019/12/14 by Akash Kumar Gautam, Puneet Mathur, Gautam, Akash +10 · 1 citation
Computer Science · Social Sciences · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Hate Speech and Cyberbullying Detection #Social Media and Politics #Social and Information Networks (cs.SI)
paper · pdf · doi:10.48550/arxiv.1912.06927
openalex publication_date 2019/12/14 · openalex created_date 2022/07/26 · openalex updated_date 2026/07/28
In this paper, we present a dataset containing 9,973 tweets related to the\nMeToo movement that were manually annotated for five different linguistic\naspects: relevance, stance, hate speech, sarcasm, and dialogue acts. We present\na detailed account of the data collection and annotation processes. The\nannotations have a very high inter-annotator agreement (0.79 to 0.93 k-alpha)\ndue to the domain expertise of the annotators and clear annotation\ninstructions. We analyze the data in terms of geographical distribution, label\ncorrelations, and keywords. Lastly, we present some potential use cases of this\ndataset. We expect this dataset would be of great interest to psycholinguists,\nsocio-linguists, and computational linguists to study the discursive space of\ndigitally mobilized social movements on sensitive issues like sexual\nharassment.\n