vix.ing · top · new · best · stats · spec

Investigating Polyglot Speech Foundation Models for Learning Collective Emotion from Crowds

2025/09/19 by Orchid Chetia Phukan, Phukan, Orchid Chetia, Mohd Mujtaba Akhtar +14
Physics and Astronomy · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Opinion Dynamics and Social Influence #Sound (cs.SD) #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2509.16329

openalex publication_date 2025/09/19 · openalex created_date 2025/10/16 · openalex updated_date 2026/07/28

Abstract

This paper investigates the polyglot (multilingual) speech foundation models (SFMs) for Crowd Emotion Recognition (CER). We hypothesize that polyglot SFMs, pre-trained on diverse languages, accents, and speech patterns, are particularly adept at navigating the noisy and complex acoustic environments characteristic of crowd settings, thereby offering a significant advantage for CER. To substantiate this, we perform a comprehensive analysis, comparing polyglot, monolingual, and speaker recognition SFMs through extensive experiments on a benchmark CER dataset across varying audio durations (1 sec, 500 ms, and 250 ms). The results consistently demonstrate the superiority of polyglot SFMs, outperforming their counterparts across all audio lengths and excelling even with extremely short-duration inputs. These findings pave the way for adaptation of SFMs in setting up new benchmarks for CER.

Citations

Related