vix.ing · top · new · best · stats · spec

Convolutive Transfer Function Invariant SDR training criteria for\n Multi-Channel Reverberant Speech Separation

2020/11/30 by Christoph Boeddeker, Boeddeker, Christoph, Wangyou Zhang +15 · 1 citation
Computer Science · Engineering · #Advanced Adaptive Filtering Techniques #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Indoor and Outdoor Localization Technologies #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2011.15003

openalex publication_date 2020/11/30 · openalex created_date 2022/07/25 · openalex updated_date 2026/07/28

Abstract

Time-domain training criteria have proven to be very effective for the\nseparation of single-channel non-reverberant speech mixtures. Likewise,\nmask-based beamforming has shown impressive performance in multi-channel\nreverberant speech enhancement and source separation. Here, we propose to\ncombine neural network supported multi-channel source separation with a\ntime-domain training objective function. For the objective we propose to use a\nconvolutive transfer function invariant Signal-to-Distortion Ratio (CI-SDR)\nbased loss. While this is a well-known evaluation metric (BSS Eval), it has not\nbeen used as a training objective before. To show the effectiveness, we\ndemonstrate the performance on LibriSpeech based reverberant mixtures. On this\ntask, the proposed system approaches the error rate obtained on single-source\nnon-reverberant input, i.e., LibriSpeech testclean, with a difference of only\n1.2 percentage points, thus outperforming a conventional permutation invariant\ntraining based system and alternative objectives like Scale Invariant\nSignal-to-Distortion Ratio by a large margin.\n

Cited by

Related