vix.ing · top · new · best · stats · spec

Keyword Spotting for Hearing Assistive Devices Robust to External\n Speakers

2019/06/22 by Iván López‐Espejo, López-Espejo, Iván, Zheng‐Hua Tan +3
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.1906.09417

openalex publication_date 2019/06/22 · openalex created_date 2022/07/22 · openalex updated_date 2026/07/28

Abstract

Keyword spotting (KWS) is experiencing an upswing due to the pervasiveness of\nsmall electronic devices that allow interaction with them via speech. Often,\nKWS systems are speaker-independent, which means that any person --user or\nnot-- might trigger them. For applications like KWS for hearing assistive\ndevices this is unacceptable, as only the user must be allowed to handle them.\nIn this paper we propose KWS for hearing assistive devices that is robust to\nexternal speakers. A state-of-the-art deep residual network for small-footprint\nKWS is regarded as a basis to build upon. By following a multi-task learning\nscheme, this system is extended to jointly perform KWS and users'\nown-voice/external speaker detection with a negligible increase in the number\nof parameters. For experiments, we generate from the Google Speech Commands\nDataset a speech corpus emulating hearing aids as a capturing device. Our\nresults show that this multi-task deep residual network is able to achieve a\nKWS accuracy relative improvement of around 32% with respect to a system that\ndoes not deal with external speakers.\n

Related