2025/06/08 by Yijia Dai, Zhaolin Gao, Dai, Yijia +7 · 2 citations
Computer Science · #Artificial Intelligence (cs.AI) #Data set #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Hidden Markov model #Key (lock) #Language model #Machine Learning (cs.LG) #Machine Learning in Healthcare #Markov model #Markov process #Multimodal Machine Learning Applications #Set (abstract data type) #Training set
paper · pdf · doi:10.48550/arxiv.2506.07298
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2025/06/08 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
Hidden Markov Models (HMMs) are foundational tools for modeling sequential data with latent Markovian structure, yet fitting them to real-world data remains computationally challenging. In this work, we show that pre-trained large language models (LLMs) can effectively model data generated by HMMs via in-context learning (ICL)\unicodex2013their ability to infer patterns from examples within a prompt. On a diverse set of synthetic HMMs, LLMs achieve predictive accuracy approaching the theoretical optimum. We uncover novel scaling trends influenced by HMM properties, and offer theoretical conjectures for these empirical observations. We also provide practical guidelines for scientists on using ICL as a diagnostic tool for complex data. On real-world animal decision-making tasks, ICL achieves competitive performance with models designed by human experts. To our knowledge, this is the first demonstration that ICL can learn and predict HMM-generated sequences\unicodex2013an advance that deepens our understanding of in-context learning in LLMs and establishes its potential as a powerful tool for uncovering hidden structure in complex scientific data.