vix.ing · top · new · best · stats · spec

Graph-based Topic Extraction from Vector Embeddings of Text Documents:\n Application to a Corpus of News Articles

2020/10/28 by M. Tarık Altuncu, Altuncu, M. Tarik, Sophia N. Yaliraki +3 · 1 citation
Computer Science · Physics and Astronomy · Social Sciences · #Advanced Text Analysis Techniques #Artificial Intelligence (cs.AI) #Complex Network Analysis Techniques #Computation and Language (cs.CL) #Computational and Text Analysis Methods #FOS: Computer and information sciences #Machine Learning (cs.LG)

paper · pdf · doi:10.48550/arxiv.2010.15067

openalex publication_date 2020/10/28 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Production of news content is growing at an astonishing rate. To help manage\nand monitor the sheer amount of text, there is an increasing need to develop\nefficient methods that can provide insights into emerging content areas, and\nstratify unstructured corpora of text into `topics' that stem intrinsically\nfrom content similarity. Here we present an unsupervised framework that brings\ntogether powerful vector embeddings from natural language processing with tools\nfrom multiscale graph partitioning that can reveal natural partitions at\ndifferent resolutions without making a priori assumptions about the number of\nclusters in the corpus. We show the advantages of graph-based clustering\nthrough end-to-end comparisons with other popular clustering and topic\nmodelling methods, and also evaluate different text vector embeddings, from\nclassic Bag-of-Words to Doc2Vec to the recent transformers based model Bert.\nThis comparative work is showcased through an analysis of a corpus of US news\ncoverage during the presidential election year of 2016.\n

Cited by

Related