2024/11/18 by Sreeram Vennam, David Valente, Vennam, Sreeram +5 · 1 citation
Computer Science · Social Sciences · #Computation and Language (cs.CL) #Education and Critical Thinking Development #FOS: Computer and information sciences #I.2.6 #Intelligent Tutoring Systems and Adaptive Learning #Machine Learning (cs.LG)
paper · pdf · doi:10.48550/arxiv.2411.11371
openalex publication_date 2024/11/18 · openalex created_date 2024/11/21 · openalex updated_date 2026/07/28
Thinking Tokens (TT) have been proposed as an unsupervised method to facilitate reasoning in language models. However, despite their conceptual appeal, our findings show that TTs marginally improves performance and consistently underperforms compared to Chain-of-Thought (CoT) reasoning across multiple benchmarks. We hypothesize that this underperformance stems from the reliance on a single embedding for TTs, which results in inconsistent learning signals and introduces noisy gradients. This paper provides a comprehensive empirical analysis to validate this hypothesis and discusses the implications for future research on unsupervised reasoning in LLMs.