2021/02/06 by Ying Chen, Lei, Hao, Chen, Ying
Computer Science · #Advanced Text Analysis Techniques #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Text and Document Classification Technologies #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2102.04449
openalex publication_date 2021/02/06 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We propose a Concentrated Document Topic Model(CDTM) for unsupervised text classification, which is able to produce a concentrated and sparse document topic distribution. In particular, an exponential entropy penalty is imposed on the document topic distribution. Documents that have diverse topic distributions are penalized more, while those having concentrated topics are penalized less. We apply the model to the benchmark NIPS dataset and observe more coherent topics and more concentrated and sparse document-topic distributions than Latent Dirichlet Allocation(LDA).