Tomohiro Nakatani
- Improving speaker discrimination of target speech extraction with time-domain SpeakerBeam
2020/01/23 by Marc Delcroix, Tsubasa Ochiai, Delcroix, Marc +11 · 19 citations
Computer Science · #Speech and Audio Processing #Music and Audio Processing #Speech Recognition and Synthesis
- Far-Field Automatic Speech Recognition
2020/09/09 by Reinhold Haeb‐Umbach, Reinhold Haeb-Umbach, Jahn Heymann +4 · 9 citations
Computer Science · #Speech and Audio Processing #Speech Recognition and Synthesis #Music and Audio Processing
- Improving noise robust automatic speech recognition with single-channel time-domain enhancement network
2020/03/09 by Keisuke Kinoshita, Kinoshita, Keisuke, Tsubasa Ochiai +5 · 4 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Far-Field Automatic Speech Recognition
2020/09/20 by Reinhold Haeb‐Umbach, Jahn Heymann, Haeb-Umbach, Reinhold +9 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Interaural time difference loss for binaural target sound extraction
2024/08/01 by Carlos Hernandez-Olivan, Hernandez-Olivan, Carlos, Marc Delcroix +9 · 4 citations
Computer Science · Earth and Planetary Sciences · Engineering · #Advanced SAR Imaging Techniques #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech and Audio Processing #Underwater Acoustics Research #electronic engineering #information engineering
- Speaker activity driven neural speech extraction
2021/01/14 by Marc Delcroix, Delcroix, Marc, Kateřina Žmolíková +7 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Listen only to me! How well can target speech extraction handle false alarms?
2022/04/11 by Marc Delcroix, Keisuke Kinoshita, Delcroix, Marc +9 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- SoundBeam meets M2D: Target Sound Extraction with Audio Foundation Model
2024/09/19 by Carlos Hernandez-Olivan, Hernandez-Olivan, Carlos, Marc Delcroix +11 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- NTT Multi-Speaker ASR System for the DASR Task of CHiME-8 Challenge
2024/09/09 by Naoyuki Kamo, Naohiro Tawara, Kamo, Naoyuki +33 · 4 citations
Computer Science · Engineering · #Advanced Data Processing Techniques #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Robotics and Automated Systems #Speech Recognition and Synthesis #electronic engineering #information engineering
- Array Geometry-Robust Attention-Based Neural Beamformer for Moving Speakers
2024/02/05 by Marvin Tammen, Tsubasa Ochiai, Tammen, Marvin +9 · 2 citations
Computer Science · #Advanced Data Compression Techniques #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Convolutive Transfer Function Invariant SDR training criteria for\n Multi-Channel Reverberant Speech Separation
2020/11/30 by Christoph Boeddeker, Boeddeker, Christoph, Wangyou Zhang +15 · 1 citation
Computer Science · Engineering · #Advanced Adaptive Filtering Techniques #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Indoor and Outdoor Localization Technologies #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- Exploiting Attention-based Sequence-to-Sequence Architectures for Sound Event Localization
2021/02/28 by Christopher Schymura, Tsubasa Ochiai, Schymura, Christopher +11 · 1 citation
Computer Science · Earth and Planetary Sciences · Neuroscience · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Hearing Loss and Rehabilitation #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #Underwater Acoustics Research #electronic engineering #information engineering
- Mask-based Neural Beamforming for Moving Speakers with Self-Attention-based Tracking
2022/05/07 by Tsubasa Ochiai, Ochiai, Tsubasa, Marc Delcroix +5 · 1 citation
Computer Science · Earth and Planetary Sciences · Engineering · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Indoor and Outdoor Localization Technologies #Sound (cs.SD) #Speech and Audio Processing #Underwater Acoustics Research #electronic engineering #information engineering
- Independent Deeply Learned Tensor Analysis for Determined Audio Source Separation
2021/06/10 by Naoki Narisawa, Narisawa, Naoki, Rintaro Ikeshita +11 · 1 citation
Computer Science · Engineering · #Advanced Adaptive Filtering Techniques #Audio and Speech Processing (eess.AS) #Blind Source Separation Techniques #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- PILOT: Introducing Transformers for Probabilistic Sound Event\n Localization
2021/06/07 by Christopher Schymura, Schymura, Christopher, Benedikt Bönninghoff +13 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- ISS2: An Extension of Iterative Source Steering Algorithm for Majorization-Minimization-Based Independent Vector Analysis
2022/02/02 by Rintaro Ikeshita, Ikeshita, Rintaro, Tomohiro Nakatani +1 · 1 citation
Computer Science · Engineering · #Advanced Adaptive Filtering Techniques #Audio and Speech Processing (eess.AS) #Blind Source Separation Techniques #FOS: Electrical engineering #Signal Processing (eess.SP) #Speech and Audio Processing #electronic engineering #information engineering
- Multi-Stream Extension of Variational Bayesian HMM Clustering (MS-VBx) for Combined End-to-End and Vector Clustering-based Diarization
2023/05/23 by Marc Delcroix, Delcroix, Marc, Naohiro Tawara +15 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Description and Discussion on DCASE 2025 Challenge Task 4: Spatial Semantic Segmentation of Sound Scenes
2025/06/12 by Masahiro Yasuda, Yasuda, Masahiro, Binh Thien Nguyen +22 · 5 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering