vix.ing · top · new · best · stats · spec

An Exploration of Length Generalization in Transformer-Based Speech Enhancement

2024/06/17 by Qiquan Zhang, Hongxu Zhu, Zhang, Qiquan +7 · 3 citations
Computer Science · Engineering · #Advanced Adaptive Filtering Techniques #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Speech and Audio Processing #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2406.11401

openalex publication_date 2024/06/17 · openalex created_date 2024/06/19 · openalex updated_date 2026/07/28

Abstract

The use of Transformer architectures has facilitated remarkable progress in speech enhancement. Training Transformers using substantially long speech utterances is often infeasible as self-attention suffers from quadratic complexity. It is a critical and unexplored challenge for a Transformer-based speech enhancement model to learn from short speech utterances and generalize to longer ones. In this paper, we conduct comprehensive experiments to explore the length generalization problem in speech enhancement with Transformer. Our findings first establish that position embedding provides an effective instrument to alleviate the impact of utterance length on Transformer-based speech enhancement. Specifically, we explore four different position embedding schemes to enable length generalization. The results confirm the superiority of relative position embeddings (RPEs) over absolute PE (APEs) in length generalization.

Cited by

Related