vix.ing · top · new · best · stats · spec

Bayesian Models for Unit Discovery on a Very Low Resource Language

2018/02/16 by Lucas Ondel, Pierre Godard, P. Godard +20
Computer Science · Mathematics · #Artificial intelligence #Baseline (sea) #Bayesian probability #Computation and Language (cs.CL) #Computer science #FOS: Computer and information sciences #Field (mathematics) #Hidden Markov model #Language model #Linguistics #Machine Learning and Algorithms #Machine learning #Mathematics #Mathematics education #Natural Language Processing Techniques #Natural language processing #Resource (disambiguation) #Segmentation #Speech Recognition and Synthesis #Unit (ring theory) #Word (group theory) #cs.CL

paper · pdf · doi:10.48550/arxiv.1802.06053

Accepted to ICASSP 2018

openalex publication_date 2018/02/16 · arxiv created 2018/02/20 · arxiv updated 2018/02/21 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Developing speech technologies for low-resource languages has become a very active research field over the last decade. Among others, Bayesian models have shown some promising results on artificial examples but still lack of in situ experiments. Our work applies state-of-the-art Bayesian models to unsupervised Acoustic Unit Discovery (AUD) in a real low-resource language scenario. We also show that Bayesian models can naturally integrate information from other resourceful languages by means of informative prior leading to more consistent discovered units. Finally, discovered acoustic units are used, either as the 1-best sequence or as a lattice, to perform word segmentation. Word segmentation results show that this Bayesian approach clearly outperforms a Segmental-DTW baseline on the same corpus.

Citations

Related