2020/12/14 by Lisa van Staden, van Staden, Lisa, Herman Kamper +1 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Speech Recognition and Synthesis #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2012.07387
openalex publication_date 2020/12/14 · openalex created_date 2022/03/05 · openalex updated_date 2026/07/28
Many speech processing tasks involve measuring the acoustic similarity\nbetween speech segments. Acoustic word embeddings (AWE) allow for efficient\ncomparisons by mapping speech segments of arbitrary duration to\nfixed-dimensional vectors. For zero-resource speech processing, where\nunlabelled speech is the only available resource, some of the best AWE\napproaches rely on weak top-down constraints in the form of automatically\ndiscovered word-like segments. Rather than learning embeddings at the segment\nlevel, another line of zero-resource research has looked at representation\nlearning at the short-time frame level. Recent approaches include\nself-supervised predictive coding and correspondence autoencoder (CAE) models.\nIn this paper we consider whether these frame-level features are beneficial\nwhen used as inputs for training to an unsupervised AWE model. We compare\nframe-level features from contrastive predictive coding (CPC), autoregressive\npredictive coding and a CAE to conventional MFCCs. These are used as inputs to\na recurrent CAE-based AWE model. In a word discrimination task on English and\nXitsonga data, all three representation learning approaches outperform MFCCs,\nwith CPC consistently showing the biggest improvement. In cross-lingual\nexperiments we find that CPC features trained on English can also be\ntransferred to Xitsonga.\n