vix.ing · top · new · best · stats · spec

MM-KWS: Multi-modal Prompts for Multilingual User-defined Keyword Spotting

2024/06/11 by Zhiqi Ai, Zhiyong Chen, Ai, Zhiqi +3 · 2 citations
Computer Science · #Advanced Text Analysis Techniques #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Digital Communication and Language #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech and dialogue systems #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2406.07310

openalex publication_date 2024/06/11 · openalex created_date 2024/06/13 · openalex updated_date 2026/07/28

Abstract

In this paper, we propose MM-KWS, a novel approach to user-defined keyword spotting leveraging multi-modal enrollments of text and speech templates. Unlike previous methods that focus solely on either text or speech features, MM-KWS extracts phoneme, text, and speech embeddings from both modalities. These embeddings are then compared with the query speech embedding to detect the target keywords. To ensure the applicability of MM-KWS across diverse languages, we utilize a feature extractor incorporating several multilingual pre-trained models. Subsequently, we validate its effectiveness on Mandarin and English tasks. In addition, we have integrated advanced data augmentation tools for hard case mining to enhance MM-KWS in distinguishing confusable words. Experimental results on the LibriPhrase and WenetPhrase datasets demonstrate that MM-KWS outperforms prior methods significantly.

Cited by

Related