vix.ing · top · new · best · stats · spec

Mask-based Neural Beamforming for Moving Speakers with Self-Attention-based Tracking

2022/05/07 by Tsubasa Ochiai, Marc Delcroix, Ochiai, Tsubasa +5 · 1 citation
Computer Science · Earth and Planetary Sciences · Engineering · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Indoor and Outdoor Localization Technologies #Sound (cs.SD) #Speech and Audio Processing #Underwater Acoustics Research #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2205.03568

openalex publication_date 2022/05/07 · openalex created_date 2022/05/22 · openalex updated_date 2026/07/28

Abstract

Beamforming is a powerful tool designed to enhance speech signals from the direction of a target source. Computing the beamforming filter requires estimating spatial covariance matrices (SCMs) of the source and noise signals. Time-frequency masks are often used to compute these SCMs. Most studies of mask-based beamforming have assumed that the sources do not move. However, sources often move in practice, which causes performance degradation. In this paper, we address the problem of mask-based beamforming for moving sources. We first review classical approaches to tracking a moving source, which perform online or blockwise computation of the SCMs. We show that these approaches can be interpreted as computing a sum of instantaneous SCMs weighted by attention weights. These weights indicate which time frames of the signal to consider in the SCM computation. Online or blockwise computation assumes a heuristic and deterministic way of computing these attention weights that, although simple, may not result in optimal performance. We thus introduce a learning-based framework that computes optimal attention weights for beamforming. We achieve this using a neural network implemented with self-attention layers. We show experimentally that our proposed framework can greatly improve beamforming performance in moving source situations while maintaining high performance in non-moving situations, thus enabling the development of mask-based beamformers robust to source movements.

Cited by

Related