2020/10/23 by Seyyed Saeed Sarfjoo, Sarfjoo, Seyyed Saeed, Srikanth Madikeri +3 · 1 citation
Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2010.12277
openalex publication_date 2020/10/23 · openalex created_date 2022/07/25 · openalex updated_date 2026/07/28
To better model the contextual information and increase the generalization\nability of Speech Activity Detection (SAD) system, this paper leverages a\nmulti-lingual Automatic Speech Recognition (ASR) system to perform SAD.\nSequence discriminative training of Acoustic Model (AM) using Lattice-Free\nMaximum Mutual Information (LF-MMI) loss function, effectively extracts the\ncontextual information of the input acoustic frame. Multi-lingual AM training,\ncauses the robustness to noise and language variabilities. The index of maximum\noutput posterior is considered as a frame-level speech/non-speech decision\nfunction. Majority voting and logistic regression are applied to fuse the\nlanguage-dependent decisions. The multi-lingual ASR is trained on 18 languages\nof BABEL datasets and the built SAD is evaluated on 3 different languages. On\nout-of-domain datasets, the proposed SAD model shows significantly better\nperformance with respect to baseline models. On the Ester2 dataset, without\nusing any in-domain data, this model outperforms the WebRTC, phoneme recognizer\nbased VAD (Phn Rec), and Pyannote baselines (respectively by 7.1, 1.7, and 2.7%\nabsolute) in Detection Error Rate (DetER) metrics. Similarly, on the LiveATC\ndataset, this model outperforms the WebRTC, Phn Rec, and Pyannote baselines\n(respectively by 6.4, 10.0, and 3.7% absolutely) in DetER metrics.\n