vix.ing · top · new · best · stats · spec

Harnessing GANs for Zero-shot Learning of New Classes in Visual Speech\n Recognition

2019/01/29 by Yaman Kumar, Kumar, Yaman, Dhruva Sahrawat +13 · 1 citation
Computer Science · Engineering · #Advanced Adaptive Filtering Techniques #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Music and Audio Processing #Speech and Audio Processing

paper · pdf · doi:10.48550/arxiv.1901.10139

openalex publication_date 2019/01/29 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Visual Speech Recognition (VSR) is the process of recognizing or interpreting\nspeech by watching the lip movements of the speaker. Recent machine learning\nbased approaches model VSR as a classification problem; however, the scarcity\nof training data leads to error-prone systems with very low accuracies in\npredicting unseen classes. To solve this problem, we present a novel approach\nto zero-shot learning by generating new classes using Generative Adversarial\nNetworks (GANs), and show how the addition of unseen class samples increases\nthe accuracy of a VSR system by a significant margin of 27% and allows it to\nhandle speaker-independent out-of-vocabulary phrases. We also show that our\nmodels are language agnostic and therefore capable of seamlessly generating,\nusing English training data, videos for a new language (Hindi). To the best of\nour knowledge, this is the first work to show empirical evidence of the use of\nGANs for generating training samples of unseen classes in the domain of VSR,\nhence facilitating zero-shot learning. We make the added videos for new classes\npublicly available along with our code.\n

Cited by

Related