2018/06/30 by Ted Underwood, Underwood, Ted
Computer Science · Social Sciences · #Advanced Text Analysis Techniques #Computation and Language (cs.CL) #Computational and Text Analysis Methods #Computers and Society (cs.CY) #Data Analysis with R #Digital Libraries (cs.DL) #FOS: Computer and information sciences
paper · pdf · doi:10.48550/arxiv.1807.00181
openalex publication_date 2018/06/30 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Measuring similarity is a basic task in information retrieval, and now often a building-block for more complex arguments about cultural change. But do measures of textual similarity and distance really correspond to evidence about cultural proximity and differentiation? To explore that question empirically, this paper compares textual and social measures of the similarities between genres of English-language fiction. Existing measures of textual similarity (cosine similarity on tf-idf vectors or topic vectors) are also compared to new strategies that use supervised learning to anchor textual measurement in a social context.