2026/07/01 by Yumeng Shen, Stanislav Mulík, María Teresa Bajo Molina +2
Neuroscience · Psychology · #Language Development and Disorders #Neurobiology of Language and Bilingualism #Phonetics and Phonology Research
paper · doi:10.1016/j.csl.2026.102038
openalex publication_date 2026/07/01 · openalex created_date 2026/07/30 · openalex updated_date 2026/07/30
Automatic speech recognition (ASR) systems are increasingly used as transcription tools in psycholinguistic research. The present study examines whether Whisper largev3 parallels human sensitivity to morphosyntactic variability in speech production. We focus on Spanish grammatical gender agreement in regular and dual-gendered nouns (DGNs). DGNs are feminine nouns that begin with stressed /a/ and often take masculine determiners (el agua), creating ambiguity between prescriptive rules and real-world usage. Using a controlled sentence-repetition paradigm, we compared native Spanish speakers and Whisper large-v3 in repetition/transcription accuracy and correction behavior when processing determiner-noun gender mismatches. Both humans and Whisper corrected prescriptively ungrammatical input for regular nouns. For DGNs, humans exhibited determiner-specific variability in correction shaped by collocational frequency, whereas Whisper defaulted to faithful transcription regardless of grammaticality or collocational frequency. These results indicate that Whisper behaves as though it favors prescriptively dominant forms in grammatically stable contexts, but shows little adjustment in the absence of reliable distributional cues. Confidence-score and n-best analyses suggested that, in rare DGN correction cases, the prescriptively incorrect determiner produced by humans often remained available among alternative hypotheses, but received a lower decoding score than the prescriptively correct form. Although Whisper can serve as an efficient transcription tool, high transcription accuracy does not entail human-like language processing. Structured human verification and linguistically informed evaluation remain essential when Whisper is used to transcribe or analyze morphosyntactic variation. More broadly, structured grammatical variability provides a testbed for improving Whisper evaluation and developing computational models that better align with human spokenlanguage behavior.