2019/04/12 by Dagoberto Porras, Porras, Dagoberto, Alexander Sepúlveda +3
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Biological sciences #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Tissues and Organs (q-bio.TO) #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.1904.06083
openalex publication_date 2019/04/12 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Speech sounds are produced as the coordinated movement of the speaking\norgans. There are several available methods to model the relation of\narticulatory movements and the resulting speech signal. The reverse problem is\noften called as acoustic-to-articulatory inversion (AAI). In this paper we have\nimplemented several different Deep Neural Networks (DNNs) to estimate the\narticulatory information from the acoustic signal. There are several previous\nworks related to performing this task, but most of them are using\nElectroMagnetic Articulography (EMA) for tracking the articulatory movement.\nCompared to EMA, Ultrasound Tongue Imaging (UTI) is a technique of higher\ncost-benefit if we take into account equipment cost, portability, safety and\nvisualized structures. Seeing that, our goal is to train a DNN to obtain UT\nimages, when using speech as input. We also test two approaches to represent\nthe articulatory information: 1) the EigenTongue space and 2) the raw\nultrasound image. As an objective quality measure for the reconstructed UT\nimages, we use MSE, Structural Similarity Index (SSIM) and Complex-Wavelet SSIM\n(CW-SSIM). Our experimental results show that CW-SSIM is the most useful error\nmeasure in the UTI context. We tested three different system configurations: a)\nsimple DNN composed of 2 hidden layers with 64x64 pixels of an UTI file as\ntarget; b) the same simple DNN but with ultrasound images projected to the\nEigenTongue space as the target; c) and a more complex DNN composed of 5 hidden\nlayers with UTI files projected to the EigenTongue space. In a subjective\nexperiment the subjects found that the neural networks with two hidden layers\nwere more suitable for this inversion task.\n