2019/12/14 by Beaufays Francoise, Sim, Khe Chai, Beaufays, Françoise +20
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.1912.09251
openalex publication_date 2019/12/14 · openalex created_date 2020/07/16 · openalex updated_date 2026/07/28
We study the effectiveness of several techniques to personalize end-to-end\nspeech models and improve the recognition of proper names relevant to the user.\nThese techniques differ in the amounts of user effort required to provide\nsupervision, and are evaluated on how they impact speech recognition\nperformance. We propose using keyword-dependent precision and recall metrics to\nmeasure vocabulary acquisition performance. We evaluate the algorithms on a\ndataset that we designed to contain names of persons that are difficult to\nrecognize. Therefore, the baseline recall rate for proper names in this dataset\nis very low: 2.4%. A data synthesis approach we developed brings it to 48.6%,\nwith no need for speech input from the user. With speech input, if the user\ncorrects only the names, the name recall rate improves to 64.4%. If the user\ncorrects all the recognition errors, we achieve the best recall of 73.5%. To\neliminate the need to upload user data and store personalized models on a\nserver, we focus on performing the entire personalization workflow on a mobile\ndevice.\n