2023/11/16 by Maria Antoniak, Joel Mire, Antoniak, Maria +7 · 5 citations
Computer Science · Health Professions · #Computation and Language (cs.CL) #Digital Storytelling and Education #FOS: Computer and information sciences #Video Analysis and Summarization
paper · pdf · doi:10.48550/arxiv.2311.09675
openalex publication_date 2023/11/16 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Story detection in online communities is a challenging task as stories are scattered across communities and interwoven with non-storytelling spans within a single text. We address this challenge by building and releasing the StorySeeker toolkit, including a richly annotated dataset of 502 Reddit posts and comments, a detailed codebook adapted to the social media context, and models to predict storytelling at the document and span levels. Our dataset is sampled from hundreds of popular English-language Reddit communities ranging across 33 topic categories, and it contains fine-grained expert annotations, including binary story labels, story spans, and event spans. We evaluate a range of detection methods using our data, and we identify the distinctive textual features of online storytelling, focusing on storytelling spans. We illuminate distributional characteristics of storytelling on a large community-centric social media platform, and we also conduct a case study on r/ChangeMyView, where storytelling is used as one of many persuasive strategies, illustrating that our data and models can be used for both inter- and intra-community research. Finally, we discuss implications of our tools and analyses for narratology and the study of online communities.