vix.ing · top · new · best · stats · spec

Kaizen: Continuously improving teacher using Exponential Moving Average\n for semi-supervised speech recognition

2021/06/14 by Vimal Manohar, Tatiana Likhomanenko, Manohar, Vimal +13 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2106.07759

openalex publication_date 2021/06/14 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

In this paper, we introduce the Kaizen framework that uses a continuously\nimproving teacher to generate pseudo-labels for semi-supervised speech\nrecognition (ASR). The proposed approach uses a teacher model which is updated\nas the exponential moving average (EMA) of the student model parameters. We\ndemonstrate that it is critical for EMA to be accumulated with full-precision\nfloating point. The Kaizen framework can be seen as a continuous version of the\niterative pseudo-labeling approach for semi-supervised training. It is\napplicable for different training criteria, and in this paper we demonstrate\nits effectiveness for frame-level hybrid hidden Markov model-deep neural\nnetwork (HMM-DNN) systems as well as sequence-level Connectionist Temporal\nClassification (CTC) based models.\n For large scale real-world unsupervised public videos in UK English and\nItalian languages the proposed approach i) shows more than 10% relative word\nerror rate (WER) reduction over standard teacher-student training; ii) using\njust 10 hours of supervised data and a large amount of unsupervised data closes\nthe gap to the upper-bound supervised ASR system that uses 650h or 2700h\nrespectively.\n

Cited by

Related