vix.ing · top · new · best · stats · spec

An automatic mixing speech enhancement system for multi-track audio

2024/04/27 by Xiaojing Liu, Liu, Xiaojing, Hongwei Ai +3
Computer Science · Engineering · #Advanced Adaptive Filtering Techniques #Advanced Data Compression Techniques #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2404.17821

openalex publication_date 2024/04/27 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

We propose a speech enhancement system for multitrack audio. The system will minimize auditory masking while allowing one to hear multiple simultaneous speakers. The system can be used in multiple communication scenarios e.g., teleconferencing, invoice gaming, and live streaming. The ITU-R BS.1387 Perceptual Evaluation of Audio Quality (PEAQ) model is used to evaluate the amount of masking in the audio signals. Different audio effects e.g., level balance, equalization, dynamic range compression, and spatialization are applied via an iterative Harmony searching algorithm that aims to minimize the masking. In the subjective listening test, the designed system can compete with mixes by professional sound engineers and outperforms mixes by existing auto-mixing systems.

Related