vix.ing · top · new · best · stats · spec

Visually Grounded Keyword Detection and Localisation for Low-Resource Languages

2023/02/01 by Kayode Olaleye, Olaleye, Kayode Kolawole
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech and Audio Processing #Text and Document Classification Technologies #Video Analysis and Summarization #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2302.00765

openalex publication_date 2023/02/01 · openalex created_date 2023/02/04 · openalex updated_date 2026/07/28

Abstract

This study investigates the use of Visually Grounded Speech (VGS) models for keyword localisation in speech. The study focusses on two main research questions: (1) Is keyword localisation possible with VGS models and (2) Can keyword localisation be done cross-lingually in a real low-resource setting? Four methods for localisation are proposed and evaluated on an English dataset, with the best-performing method achieving an accuracy of 57%. A new dataset containing spoken captions in Yoruba language is also collected and released for cross-lingual keyword localisation. The cross-lingual model obtains a precision of 16% in actual keyword localisation and this performance can be improved by initialising from a model pretrained on English data. The study presents a detailed analysis of the model's success and failure modes and highlights the challenges of using VGS models for keyword localisation in low-resource settings.

Related