2015/12/07 by Bei Chen, Chen, Bei, Jun Zhu +9
Computer Science · #Advanced Text Analysis Techniques #Computation and Language (cs.CL) #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling #cs.CL #cs.IR #cs.LG
paper · pdf · doi:10.48550/arxiv.1512.02009
Accepted by AAAI 2016
arxiv created 2015/12/07 · openalex publication_date 2015/12/07 · arxiv updated 2015/12/08 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Modeling document structure is of great importance for discourse analysis and related applications. The goal of this research is to capture the document intent structure by modeling documents as a mixture of topic words and rhetorical words. While the topics are relatively unchanged through one document, the rhetorical functions of sentences usually change following certain orders in discourse. We propose GMM-LDA, a topic modeling based Bayesian unsupervised model, to analyze the document intent structure cooperated with order information. Our model is flexible that has the ability to combine the annotations and do supervised learning. Additionally, entropic regularization can be introduced to model the significant divergence between topics and intents. We perform experiments in both unsupervised and supervised settings, results show the superiority of our model over several state-of-the-art baselines.