vix.ing · top · new · best · stats · spec

DNN-based Acoustic-to-Articulatory Inversion using Ultrasound Tongue\n Imaging

2019/04/12 by Dagoberto Porras, Porras, Dagoberto, Alexander Sepúlveda +3
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Biological sciences #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Tissues and Organs (q-bio.TO) #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.1904.06083

openalex publication_date 2019/04/12 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Speech sounds are produced as the coordinated movement of the speaking\norgans. There are several available methods to model the relation of\narticulatory movements and the resulting speech signal. The reverse problem is\noften called as acoustic-to-articulatory inversion (AAI). In this paper we have\nimplemented several different Deep Neural Networks (DNNs) to estimate the\narticulatory information from the acoustic signal. There are several previous\nworks related to performing this task, but most of them are using\nElectroMagnetic Articulography (EMA) for tracking the articulatory movement.\nCompared to EMA, Ultrasound Tongue Imaging (UTI) is a technique of higher\ncost-benefit if we take into account equipment cost, portability, safety and\nvisualized structures. Seeing that, our goal is to train a DNN to obtain UT\nimages, when using speech as input. We also test two approaches to represent\nthe articulatory information: 1) the EigenTongue space and 2) the raw\nultrasound image. As an objective quality measure for the reconstructed UT\nimages, we use MSE, Structural Similarity Index (SSIM) and Complex-Wavelet SSIM\n(CW-SSIM). Our experimental results show that CW-SSIM is the most useful error\nmeasure in the UTI context. We tested three different system configurations: a)\nsimple DNN composed of 2 hidden layers with 64x64 pixels of an UTI file as\ntarget; b) the same simple DNN but with ultrasound images projected to the\nEigenTongue space as the target; c) and a more complex DNN composed of 5 hidden\nlayers with UTI files projected to the EigenTongue space. In a subjective\nexperiment the subjects found that the neural networks with two hidden layers\nwere more suitable for this inversion task.\n

Related