2025/09/29 by Santiago Barreda, T. Florian Jaeger · 1 voice
Computer Science · Psychology · #Music and Audio Processing #Phonetics and Phonology Research #Speech and Audio Processing
paper · doi:10.1515/lingvan-2024-0239
openalex publication_date 2025/09/29 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/25
Abstract Normalization of the speech signal onto comparatively invariant phonetic representations is critical to speech perception. Assumptions about this process also play a central role in phonetics and phonology. Yet, despite the importance of these assumptions, popular normalization accounts continue to incorrectly assume that the relevant normalization parameters are “known” – for example, because the researcher can estimate them from a fully balanced set of recordings. Listeners, however, have to incrementally infer these parameters from the speech input. We reintroduce a seminal, but still underappreciated, model of this inference process – the Probabilistic Sliding Template Model of vowel normalization and perception (PSTM) – and show why researchers cannot afford to ignore it. We provide the R package STM, which makes it trivial to apply the PSTM to new data. We present the first quantitative assessment of the PSTM against human perception. We show that the model excels at predicting listeners’ vowel categorization behavior, clearly outperforming the popular Lobanov normalization.