2021/09/12 by Julia Fukuyama, Kris Sankaran, Fukuyama, Julia +3 · 1 citation
Biochemistry, Genetics and Molecular Biology · Computer Science · #Applications (stat.AP) #Bioinformatics and Genomic Networks #Computation (stat.CO) #Data Analysis with R #FOS: Computer and information sciences #Gene expression and cancer classification
paper · doi:10.48550/arxiv.2109.05541
openalex publication_date 2021/09/12 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Topic modeling is a popular method used to describe biological count data. With topic models, the user must specify the number of topics K. Since there is no definitive way to choose K and since a true value might not exist, we develop a method, which we call topic alignment, to study the relationships across models with different K. In addition, we present three diagnostics based on the alignment. These techniques can show how many topics are consistently present across different models, if a topic is only transiently present, or if a topic splits into more topics when K increases. This strategy gives more insight into the process of generating the data than choosing a single value of K would. We design a visual representation of these cross-model relationships, show the effectiveness of these tools for interpreting the topics on simulated and real data, and release an accompanying R package, alto.