2025/10/28 by Lin, Baizhou, Fang, Yuetong, Xu, Renjing +2
#Audio and Speech Processing (eess.AS) #B.7 #C.3 #FOS: Computer and information sciences #FOS: Electrical engineering #Hardware Architecture (cs.AR) #I.2 #Sound (cs.SD) #electronic engineering #information engineering
paper · doi:10.48550/arxiv.2510.24282
The Tsetlin Machine (TM) has recently attracted attention as a low-power alternative to neural networks due to its simple and interpretable inference mechanisms. However, its performance on speech-related tasks remains limited. This paper proposes TsetlinKWS, the first algorithm-hardware co-design framework for the Convolutional Tsetlin Machine (CTM) on the 12-keyword spotting task. Firstly, we introduce a novel Mel-Frequency Spectral Coefficient and Spectral Flux (MFSC-SF) feature extraction scheme together with spectral convolution, enabling the CTM to reach its first-ever competitive accuracy of 87.35% on the 12-keyword spotting task. Secondly, we develop an Optimized Grouped Block-Compressed Sparse Row (OG-BCSR) algorithm that achieves a remarkable 9.84× reduction in model size, significantly improving the storage efficiency on CTMs. Finally, we propose a state-driven architecture tailored for the CTM, which simultaneously exploits data reuse and sparsity to achieve high energy efficiency. The full system is evaluated in 65 nm process technology, consuming 16.58 μW at 0.7 V with a compact 0.63 mm2 core area. TsetlinKWS requires only 907k logic operations per inference, representing a 10× reduction compared to the state-of-the-art KWS accelerators, positioning the CTM as a highly-efficient candidate for ultra-low-power speech applications.