vix.ing · top · new · best · stats · spec

Languages in Whisper-Style Speech Encoders Align Both Phonetically and Semantically

2025/05/26 by Ryan Soh-Eun Shim, Shim, Ryan Soh-Eun, Domenico De Cristofaro +8 · 1 voice · 2 citations
Computer Science · #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and dialogue systems #cs.CL

paper · pdf · doi:10.48550/arxiv.2505.19606

openalex publication_date 2025/05/26 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Cross-lingual alignment in pretrained language models enables knowledge transfer across languages. Similar alignment has been reported in Whisper-style speech encoders, based on spoken translation retrieval using representational similarity. However, prior work does not control for phonetic overlap between equivalent utterances, which may artificially support retrieval. We conduct pronunciation-controlled experiments to test whether cross-lingual alignment arises from semantic rather than phonetic similarity. Results show that spoken translation retrieval remains strongly above chance without phonetic cues in the final layers of encoders trained with a speech translation objective, most clearly for models additionally trained on translation. We further test early-exiting the encoder to induce representations we hypothesize to be less tied to language-specific semantics. These experiments indeed reveal performance gains in automatic speech recognition on low-resource languages unseen during training.

Cited by

Discussions

Related