vix.ing · top · new · best · stats

Your thoughts tell who you are: Characterize the reasoning patterns of LRMs

2025/09/29 by Y Chen, Chen, Yida, Yuning Mao +17
Computer Science · Social Sciences · #Analytic reasoning #Artificial Intelligence (cs.AI) #Artificial Intelligence in Law #Case-based reasoning #Computation and Language (cs.CL) #FOS: Computer and information sciences #Focus (optics) #Machine Learning (cs.LG) #Multi-Agent Systems and Negotiation #Qualitative reasoning #Reasoning system #Software Engineering Techniques and Practices #TRACE (psycholinguistics) #Task (project management) #Taxonomy (biology)

paper · pdf · doi:10.48550/arxiv.2509.24147

published in arXiv (Cornell University) (Cornell University)

openalex publication_date 2025/09/29 · openalex created_date 2025/10/19 · openalex updated_date 2026/08/05

Abstract

Current comparisons of large reasoning models (LRMs) focus on macro-level statistics such as task accuracy or reasoning length. Whether different LRMs reason differently remains an open question. To address this gap, we introduce the LLM-proposed Open Taxonomy (LOT), a classification method that uses a generative language model to compare reasoning traces from two LRMs and articulate their distinctive features in words. LOT then models how these features predict the source LRM of a reasoning trace based on their empirical distributions across LRM outputs. Iterating this process over a dataset of reasoning traces yields a human-readable taxonomy that characterizes how models think. We apply LOT to compare the reasoning of 12 open-source LRMs on tasks in math, science, and coding. LOT identifies systematic differences in their thoughts, achieving 80-100% accuracy in distinguishing reasoning traces from LRMs that differ in scale, base model family, or objective domain. Beyond classification, LOT's natural-language taxonomy provides qualitative explanations of how LRMs think differently. Finally, in a case study, we link the reasoning differences to performance: aligning the reasoning style of smaller Qwen3 models with that of the largest Qwen3 during test time improves their accuracy on GPQA by 3.3-5.7%.

Citations

Related