2025/06/06 by Guillaume Wisniewski, Séverine Guillaume, Wisniewski, Guillaume +2
Computer Science · #Advanced Text Analysis Techniques #Artificial Intelligence (cs.AI) #Artificial neural network #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Invariant (physics) #Keyword spotting #Natural Language Processing Techniques #Property (philosophy) #Robustness (evolution) #Similarity (geometry) #Sound (cs.SD) #Spotting #Topic Modeling #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2506.11096
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2025/06/06 · openalex created_date 2025/10/11 · openalex updated_date 2026/08/05
Pretrained speech representations like wav2vec2 and HuBERT exhibit strong anisotropy, leading to high similarity between random embeddings. While widely observed, the impact of this property on downstream tasks remains unclear. This work evaluates anisotropy in keyword spotting for computational documentary linguistics. Using Dynamic Time Warping, we show that despite anisotropy, wav2vec2 similarity measures effectively identify words without transcription. Our results highlight the robustness of these representations, which capture phonetic structures and generalize across speakers. Our results underscore the importance of pretraining in learning rich and invariant speech representations.