Marc Delcroix
- Improving speaker discrimination of target speech extraction with time-domain SpeakerBeam
2020/01/23 by Marc Delcroix, Delcroix, Marc, Tsubasa Ochiai +11 · 13 citations
Computer Science · #Speech and Audio Processing #Music and Audio Processing #Speech Recognition and Synthesis
- Neural Target Speech Extraction: An overview
2023/05/01 by Kateřina Žmolíková, Katerina Zmolikova, Marc Delcroix +5 · 14 citations
Computer Science · #Speech and Audio Processing #Music and Audio Processing #Speech Recognition and Synthesis
- Integrating end-to-end neural and clustering-based diarization: Getting the best of both worlds
2020/10/26 by Keisuke Kinoshita, Marc Delcroix, Kinoshita, Keisuke +3 · 7 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
- Far-Field Automatic Speech Recognition
2020/09/09 by Reinhold Haeb‐Umbach, Reinhold Haeb-Umbach, Jahn Heymann +4 · 8 citations
Computer Science · #Speech and Audio Processing #Speech Recognition and Synthesis #Music and Audio Processing
- SA-SDR: A novel loss function for separation of meeting style data
2021/10/29 by Thilo von Neumann, von Neumann, Thilo, Keisuke Kinoshita +7 · 6 citations
Computer Science · #Speech and Audio Processing #Speech Recognition and Synthesis #Music and Audio Processing
- Listen to What You Want: Neural Network-based Universal Sound Selector
2020/06/10 by Tsubasa Ochiai, Ochiai, Tsubasa, Marc Delcroix +9 · 5 citations
Computer Science · Engineering · #Advanced Adaptive Filtering Techniques #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- SoundBeam: Target sound extraction conditioned on sound-class labels and enrollment clues for increased performance and continuous learning
2022/04/08 by Marc Delcroix, Jorge Bennasar Vázquez, Delcroix, Marc +9 · 6 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Improving noise robust automatic speech recognition with single-channel time-domain enhancement network
2020/03/09 by Keisuke Kinoshita, Tsubasa Ochiai, Kinoshita, Keisuke +5 · 4 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Target Speech Extraction with Conditional Diffusion Model
2023/08/08 by Naoyuki Kamo, Kamo, Naoyuki, Marc Delcroix +3 · 4 citations
Computer Science · Health Professions · #Speech and Audio Processing #Speech Recognition and Synthesis #Infant Health and Development
- Far-Field Automatic Speech Recognition
2020/09/20 by Reinhold Haeb‐Umbach, Jahn Heymann, Haeb-Umbach, Reinhold +9 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- SpeechGLUE: How Well Can Self-Supervised Speech Models Capture Linguistic Knowledge?
2023/06/14 by Takanori Ashihara, Ashihara, Takanori, Takafumi Moriya +13 · 3 citations
Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Speech and dialogue systems
- Few-shot learning of new sound classes for target sound extraction
2021/06/14 by Marc Delcroix, Delcroix, Marc, Jorge Bennasar Vázquez +7 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- SpeakerBeam-SS: Real-time Target Speaker Extraction with Lightweight Conv-TasNet and State Space Modeling
2024/07/01 by Hiroshi Sato, Sato, Hiroshi, Takafumi Moriya +15 · 4 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Signal Processing (eess.SP) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Interaural time difference loss for binaural target sound extraction
2024/08/01 by Carlos Hernandez-Olivan, Hernandez-Olivan, Carlos, Marc Delcroix +9 · 3 citations
Computer Science · Earth and Planetary Sciences · Engineering · #Advanced SAR Imaging Techniques #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech and Audio Processing #Underwater Acoustics Research #electronic engineering #information engineering
- SoundBeam meets M2D: Target Sound Extraction with Audio Foundation Model
2024/09/19 by Carlos Hernandez-Olivan, Hernandez-Olivan, Carlos, Marc Delcroix +11 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- What Do Self-Supervised Speech and Speaker Models Learn? New Findings From a Cross Model Layer-Wise Analysis
2024/01/31 by Takanori Ashihara, Marc Delcroix, Ashihara, Takanori +9 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Mask-based Neural Beamforming for Moving Speakers with Self-Attention-based Tracking
2022/05/07 by Tsubasa Ochiai, Marc Delcroix, Ochiai, Tsubasa +5 · 1 citation
Computer Science · Earth and Planetary Sciences · Engineering · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Indoor and Outdoor Localization Technologies #Sound (cs.SD) #Speech and Audio Processing #Underwater Acoustics Research #electronic engineering #information engineering
- Probing Self-supervised Learning Models with Target Speech Extraction
2024/02/17 by Junyi Peng, Peng, Junyi, Marc Delcroix +11 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Listen only to me! How well can target speech extraction handle false alarms?
2022/04/11 by Marc Delcroix, Keisuke Kinoshita, Delcroix, Marc +9 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Array Geometry-Robust Attention-Based Neural Beamformer for Moving Speakers
2024/02/05 by Marvin Tammen, Tammen, Marvin, Tsubasa Ochiai +9 · 1 citation
Computer Science · #Advanced Data Compression Techniques #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Transfer Learning from Pre-trained Language Models Improves End-to-End Speech Summarization
2023/06/07 by Kohei Matsuura, Matsuura, Kohei, Takanori Ashihara +11 · 1 citation
Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Topic Modeling
- Iterative Shallow Fusion of Backward Language Model for End-to-End Speech Recognition
2023/10/17 by Atsunori Ogawa, Takafumi Moriya, Ogawa, Atsunori +7 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Streaming Target-Speaker ASR with Neural Transducer
2022/09/09 by Takafumi Moriya, Moriya, Takafumi, Hiroshi Sato +7 · 1 citation
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
- Alignment-Free Training for Transducer-based Multi-Talker ASR
2024/09/30 by Takafumi Moriya, Moriya, Takafumi, Shota Horiguchi +13 · 2 citations
Engineering · Computer Science · #Ultrasonics and Acoustic Wave Propagation #Fault Detection and Control Systems #Speech Recognition and Synthesis
- TS-SUPERB: A Target Speech Processing Benchmark for Speech Self-Supervised Learning Models
2025/05/10 by Junyi Peng, Takanori Ashihara, Peng, Junyi +11 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
- Multi-Stream Extension of Variational Bayesian HMM Clustering (MS-VBx) for Combined End-to-End and Vector Clustering-based Diarization
2023/05/23 by Marc Delcroix, Naohiro Tawara, Delcroix, Marc +15 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- NTT Multi-Speaker ASR System for the DASR Task of CHiME-8 Challenge
2024/09/09 by Naoyuki Kamo, Kamo, Naoyuki, Naohiro Tawara +33 · 2 citations
Computer Science · Engineering · #Advanced Data Processing Techniques #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Robotics and Automated Systems #Speech Recognition and Synthesis #electronic engineering #information engineering
- Continuous Speech Separation Using Speaker Inventory for Long Multi-talker Recording
2020/12/17 by Cong Han, Han, Cong, Yi Luo +19 · 1 citation
Computer Science · #Speech and Audio Processing #Speech Recognition and Synthesis #Music and Audio Processing
- SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings
2026/07/16 by Shuai Wang, Zihan Qian, Ke Zhang +9
#eess.AS #cs.SD