2019/09/09 by Charuta Pethe, Steven Skiena, Pethe, Charuta +1 · 1 citation
Computer Science · Mathematics · Psychology · #Artificial intelligence #Authorship Attribution and Profiling #Characterization (materials science) #Computation and Language (cs.CL) #Computer science #FOS: Computer and information sciences #Hate Speech and Cyberbullying Detection #History #Information retrieval #Machine learning #Mathematics #Natural language processing #Popularity #Psychology #Representativeness heuristic #Sequence (biology) #Social psychology #Statistics #Style (visual arts) #Subject (documents) #Task (project management) #Topic Modeling #Value (mathematics) #World Wide Web #cs.CL
paper · pdf · doi:10.48550/arxiv.1909.04002
published in arXiv (Cornell University) (Cornell University) · 11 pages, 4 figures. Accepted at EMNLP-IJCNLP 2019 as a long paper
arxiv created 2019/09/09 · openalex publication_date 2019/09/09 · arxiv updated 2019/09/10 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/06
The sequence of documents produced by any given author varies in style and content, but some documents are more typical or representative of the source than others. We quantify the extent to which a given short text is characteristic of a specific person, using a dataset of tweets from fifteen celebrities. Such analysis is useful for generating excerpts of high-volume Twitter profiles, and understanding how representativeness relates to tweet popularity. We first consider the related task of binary author detection (is x the author of text T?), and report a test accuracy of 90.37% for the best of five approaches to this problem. We then use these models to compute characterization scores among all of an author's texts. A user study shows human evaluators agree with our characterization model for all 15 celebrities in our dataset, each with p-value < 0.05. We use these classifiers to show surprisingly strong correlations between characterization scores and the popularity of the associated texts. Indeed, we demonstrate a statistically significant correlation between this score and tweet popularity (likes/replies/retweets) for 13 of the 15 celebrities in our study.