vix.ing · top · new · best · stats

BUT Systems for WildSpoof Challenge: SASV in the Wild

2025/12/14 by Peng, Junyi, Jin Li, Li, Jin +8
Computer Science · #Bridge (graph theory) #Domain (mathematical analysis) #Encoder #Feature (linguistics) #Feature extraction #Minification #Music and Audio Processing #Ranging #Speech Recognition and Synthesis #Speech and Audio Processing

paper · pdf · doi:10.48550/arxiv.2512.12851

published in arXiv (Cornell University) (Cornell University)

openalex publication_date 2025/12/14 · openalex created_date 2025/12/17 · openalex updated_date 2026/08/05

Abstract

This paper presents the BUT submission to the WildSpoof Challenge, focusing on the Spoofing-robust Automatic Speaker Verification (SASV) track. We propose a SASV framework designed to bridge the gap between general audio understanding and specialized speech analysis. Our subsystem integrates diverse Self-Supervised Learning front-ends ranging from general audio models (e.g., Dasheng) to speech-specific encoders (e.g., WavLM). These representations are aggregated via a lightweight Multi-Head Factorized Attention back-end for corresponding subtasks. Furthermore, we introduce a feature domain augmentation strategy based on Distribution Uncertainty to explicitly model and mitigate the domain shift caused by unseen neural vocoders and recording environments. By fusing these robust CM scores with state-of-the-art ASV systems, our approach achieves superior minimization of the a-DCFs and EERs.

Citations

Related