2022/03/13 by Vishwanath Pratap Singh, Hardik B. Sailor, Singh, Vishwanath Pratap +5 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Signal Processing (eess.SP) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2203.06600
openalex publication_date 2022/03/13 · openalex created_date 2022/05/05 · openalex updated_date 2026/07/28
Training a robust Automatic Speech Recognition (ASR) system for children's\nspeech recognition is a challenging task due to inherent differences in\nacoustic attributes of adult and child speech and scarcity of publicly\navailable children's speech dataset. In this paper, a novel segmental spectrum\nwarping and perturbations in formant energy are introduced, to generate a\nchildren-like speech spectrum from that of an adult's speech spectrum. Then,\nthis modified adult spectrum is used as augmented data to improve end-to-end\nASR systems for children's speech recognition. The proposed data augmentation\nmethods give 6.5% and 6.1% relative reduction in WER on children dev and test\nsets respectively, compared to the vocal tract length perturbation (VTLP)\nbaseline system trained on Librispeech 100 hours adult speech dataset. When\nchildren's speech data is added in training with Librispeech set, it gives a\n3.7 % and 5.1% relative reduction in WER, compared to the VTLP baseline system.\n