vix.ing · top · new · best · stats · spec

Mispronunciation Detection in Non-native (L2) English with Uncertainty\n Modeling

2021/01/16 by Daniel Korzekwa, Jaime Lorenzo-Trueba, Korzekwa, Daniel +9 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2101.06396

openalex publication_date 2021/01/16 · openalex created_date 2022/07/25 · openalex updated_date 2026/07/28

Abstract

A common approach to the automatic detection of mispronunciation in language\nlearning is to recognize the phonemes produced by a student and compare it to\nthe expected pronunciation of a native speaker. This approach makes two\nsimplifying assumptions: a) phonemes can be recognized from speech with high\naccuracy, b) there is a single correct way for a sentence to be pronounced.\nThese assumptions do not always hold, which can result in a significant amount\nof false mispronunciation alarms. We propose a novel approach to overcome this\nproblem based on two principles: a) taking into account uncertainty in the\nautomatic phoneme recognition step, b) accounting for the fact that there may\nbe multiple valid pronunciations. We evaluate the model on non-native (L2)\nEnglish speech of German, Italian and Polish speakers, where it is shown to\nincrease the precision of detecting mispronunciations by up to 18% (relative)\ncompared to the common approach.\n

Cited by

Related