vix.ing · top · new · best · stats · spec

An NLP approach to quantify dynamic salience of predefined topics in a text corpus

2021/08/16 by A. Bock, Anthony Palladino, Bock, A. +11
Computer Science · Physics and Astronomy · Social Sciences · #Advanced Text Analysis Techniques #Complex Network Analysis Techniques #Computation and Language (cs.CL) #Computational and Text Analysis Methods #FOS: Computer and information sciences

paper · pdf · doi:10.48550/arxiv.2108.07345

openalex publication_date 2021/08/16 · openalex created_date 2021/08/30 · openalex updated_date 2026/07/28

Abstract

The proliferation of news media available online simultaneously presents a valuable resource and significant challenge to analysts aiming to profile and understand social and cultural trends in a geographic location of interest. While an abundance of news reports documenting significant events, trends, and responses provides a more democratized picture of the social characteristics of a location, making sense of an entire corpus to extract significant trends is a steep challenge for any one analyst or team. Here, we present an approach using natural language processing techniques that seeks to quantify how a set of pre-defined topics of interest change over time across a large corpus of text. We found that, given a predefined topic, we can identify and rank sets of terms, or n-grams, that map to those topics and have usage patterns that deviate from a normal baseline. Emergence, disappearance, or significant variations in n-gram usage present a ground-up picture of a topic's dynamic salience within a corpus of interest.

Related