vix.ing · top · new · best · stats · spec

Clustering and Classification in Text Collections Using Graph Modularity

2011/05/29 by Grigory Pivovarov, Pivovarov, Grigory, Sergei Trunov +1
Computer Science · Physics and Astronomy · #68U99 #Advanced Graph Neural Networks #Complex Network Analysis Techniques #Digital Libraries (cs.DL) #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Web Data Mining and Analysis #cs.DL #cs.IR #msc:68U99

paper · pdf · doi:10.48550/arxiv.1105.5789

11 pages, submitted to JMLR

arxiv created 2011/05/29 · openalex publication_date 2011/05/29 · arxiv updated 2011/05/31 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

A new fast algorithm for clustering and classification of large collections of text documents is introduced. The new algorithm employs the bipartite graph that realizes the word-document matrix of the collection. Namely, the modularity of the bipartite graph is used as the optimization functional. Experiments performed with the new algorithm on a number of text collections had shown a competitive quality of the clustering (classification), and a record-breaking speed.

Citations

Related