vix.ing · top · new · best · stats · spec

The Sound of Silencing: Identities and Ideologies in Commercial Text-To-Speech

2026/04/13 by Alice Ross, Nina Markl, Catherine Lai +1 · 1 voice
Arts and Humanities · Social Sciences · #Media, Communication, and Education #Music History and Culture #Rhetoric and Communication Studies

paper · doi:10.1145/3772363.3798657

openalex created_date 2026/04/13 · openalex publication_date 2026/04/13 · openalex updated_date 2026/07/29

Abstract

Text-to-speech (TTS) technology allows the synthesis of speech that is frequently described as highly ‘natural’ and, in some contexts, indistinguishable from human speech. Voice interfaces using such synthesised speech are increasingly encountered in a wide range of contexts. Recognising that listeners are likely to hear human-like voices as belonging to different demographic/social groups, and that these social judgments exist within ideological frameworks, we note a lack of diversity in popularly used English-speaking TTS voices, and caution that decisions taken in the design and deployment of voice interfaces risk perpetuating, or even exacerbating, existing social biases. Drawing upon sociolinguistic theory, we carry out a novel experiment to investigate these issues in a leading commercial TTS system, concluding that the system’s output disproportionately reproduces white, male, US-accented speech when prompted to convey competence. This work aims to encourage further research applying sociolinguistic knowledge to the study of human-computer interaction with speech technology.

Citations

Discussions

Related