vix.ing · top · new · best · stats · spec

Nakatani, Tomohiro

  1. Improving speaker discrimination of target speech extraction with time-domain SpeakerBeam
    2020/01/23 by Marc Delcroix, Tsubasa Ochiai, Delcroix, Marc +11 · 13 citations
    Computer Science · #Speech and Audio Processing #Music and Audio Processing #Speech Recognition and Synthesis
  2. Improving noise robust automatic speech recognition with single-channel time-domain enhancement network
    2020/03/09 by Keisuke Kinoshita, Tsubasa Ochiai, Kinoshita, Keisuke +5 · 4 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  3. Target Speech Extraction with Conditional Diffusion Model
    2023/08/08 by Naoyuki Kamo, Kamo, Naoyuki, Marc Delcroix +3 · 4 citations
    Computer Science · Health Professions · #Speech and Audio Processing #Speech Recognition and Synthesis #Infant Health and Development
  4. Far-Field Automatic Speech Recognition
    2020/09/20 by Reinhold Haeb‐Umbach, Jahn Heymann, Haeb-Umbach, Reinhold +9 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  5. Interaural time difference loss for binaural target sound extraction
    2024/08/01 by Carlos Hernandez-Olivan, Hernandez-Olivan, Carlos, Marc Delcroix +9 · 4 citations
    Computer Science · Earth and Planetary Sciences · Engineering · #Advanced SAR Imaging Techniques #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech and Audio Processing #Underwater Acoustics Research #electronic engineering #information engineering
  6. SoundBeam meets M2D: Target Sound Extraction with Audio Foundation Model
    2024/09/19 by Carlos Hernandez-Olivan, Hernandez-Olivan, Carlos, Marc Delcroix +11 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  7. Maximum likelihood convolutional beamformer for simultaneous denoising and dereverberation
    2019/08/06 by Nakatani, Tomohiro, Kinoshita, Keisuke · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  8. Integration of variational autoencoder and spatial clustering for adaptive multi-channel neural speech separation
    2020/11/24 by Zmolikova, Katerina, Delcroix, Marc, Burget, Lukáš +2 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #electronic engineering #information engineering
  9. Exploiting Attention-based Sequence-to-Sequence Architectures for Sound Event Localization
    2021/02/28 by Schymura, Christopher, Ochiai, Tsubasa, Delcroix, Marc +4 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  10. Mask-based Neural Beamforming for Moving Speakers with Self-Attention-based Tracking
    2022/05/07 by Tsubasa Ochiai, Marc Delcroix, Ochiai, Tsubasa +5 · 1 citation
    Computer Science · Earth and Planetary Sciences · Engineering · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Indoor and Outdoor Localization Technologies #Sound (cs.SD) #Speech and Audio Processing #Underwater Acoustics Research #electronic engineering #information engineering
  11. Independent Deeply Learned Tensor Analysis for Determined Audio Source Separation
    2021/06/10 by Naoki Narisawa, Narisawa, Naoki, Rintaro Ikeshita +11 · 1 citation
    Computer Science · Engineering · #Advanced Adaptive Filtering Techniques #Audio and Speech Processing (eess.AS) #Blind Source Separation Techniques #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  12. PILOT: Introducing Transformers for Probabilistic Sound Event\n Localization
    2021/06/07 by Christopher Schymura, Schymura, Christopher, Benedikt Bönninghoff +13 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  13. ISS2: An Extension of Iterative Source Steering Algorithm for Majorization-Minimization-Based Independent Vector Analysis
    2022/02/02 by Ikeshita, Rintaro, Nakatani, Tomohiro · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Signal Processing (eess.SP) #electronic engineering #information engineering
  14. Listen only to me! How well can target speech extraction handle false alarms?
    2022/04/11 by Marc Delcroix, Delcroix, Marc, Keisuke Kinoshita +9 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  15. Description and Discussion on DCASE 2025 Challenge Task 4: Spatial Semantic Segmentation of Sound Scenes
    2025/06/12 by Yasuda, Masahiro, Nguyen, Binh Thien, Harada, Noboru +10 · 5 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  16. Array Geometry-Robust Attention-Based Neural Beamformer for Moving Speakers
    2024/02/05 by Marvin Tammen, Tammen, Marvin, Tsubasa Ochiai +9 · 1 citation
    Computer Science · #Advanced Data Compression Techniques #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  17. Multi-Stream Extension of Variational Bayesian HMM Clustering (MS-VBx) for Combined End-to-End and Vector Clustering-based Diarization
    2023/05/23 by Marc Delcroix, Delcroix, Marc, Naohiro Tawara +15 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  18. NTT Multi-Speaker ASR System for the DASR Task of CHiME-8 Challenge
    2024/09/09 by Naoyuki Kamo, Naohiro Tawara, Kamo, Naoyuki +33 · 2 citations
    Computer Science · Engineering · #Advanced Data Processing Techniques #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Robotics and Automated Systems #Speech Recognition and Synthesis #electronic engineering #information engineering