2021/10/04 by Julio Cesar Duarte, Júlio César Duarte, Duarte, Julio Cesar +2 · 3 citations
Computer Science · Engineering · #Artificial intelligence #Artificial neural network #Audio and Speech Processing (eess.AS) #Classifier (UML) #Computation and Language (cs.CL) #Computer science #FOS: Computer and information sciences #FOS: Electrical engineering #Human–computer interaction #Machine learning #Music and Audio Processing #Process (computing) #Set (abstract data type) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech recognition #Task (project management) #Word error rate #cs.CL #cs.SD #eess.AS #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2110.01425
published in arXiv (Cornell University) (Cornell University) · Tech report series Monografias em Ciência da Computação, september, 2021, Dep. Informática PUC-Rio, RJ, BRAZIL, ISSN 0103-9741
arxiv created 2021/10/04 · openalex publication_date 2021/10/04 · arxiv updated 2021/10/05 · openalex created_date 2022/10/31 · openalex updated_date 2026/08/08
Automatic speech recognition systems are part of people's daily lives, embedded in personal assistants and mobile phones, helping as a facilitator for human-machine interaction while allowing access to information in a practically intuitive way. Such systems are usually implemented using machine learning techniques, especially with deep neural networks. Even with its high performance in the task of transcribing text from speech, few works address the issue of its recognition in noisy environments and, usually, the datasets used do not contain noisy audio examples, while only mitigating this issue using data augmentation techniques. This work aims to present the process of building a dataset of noisy audios, in a specific case of degenerated audios due to interference, commonly present in radio transmissions. Additionally, we present initial results of a classifier that uses such data for evaluation, indicating the benefits of using this dataset in the recognizer's training process. Such recognizer achieves an average result of 0.4116 in terms of character error rate in the noisy set (SNR = 30).