Tawara, Naohiro
- Improving speaker discrimination of target speech extraction with time-domain SpeakerBeam
2020/01/23 by Marc Delcroix, Delcroix, Marc, Tsubasa Ochiai +11 · 10 citations
Computer Science · #Speech and Audio Processing #Music and Audio Processing #Speech Recognition and Synthesis
- Advances in integration of end-to-end neural and clustering-based diarization for real conversational speech
2021/05/19 by Kinoshita, Keisuke, Delcroix, Marc, Tawara, Naohiro · 8 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Integrating end-to-end neural and clustering-based diarization: Getting the best of both worlds
2020/10/26 by Keisuke Kinoshita, Marc Delcroix, Kinoshita, Keisuke +3 · 6 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
- Interaural time difference loss for binaural target sound extraction
2024/08/01 by Carlos Hernandez-Olivan, Hernandez-Olivan, Carlos, Marc Delcroix +9 · 3 citations
Computer Science · Earth and Planetary Sciences · Engineering · #Advanced SAR Imaging Techniques #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech and Audio Processing #Underwater Acoustics Research #electronic engineering #information engineering
- SoundBeam meets M2D: Target Sound Extraction with Audio Foundation Model
2024/09/19 by Carlos Hernandez-Olivan, Marc Delcroix, Hernandez-Olivan, Carlos +11 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Mamba-based Segmentation Model for Speaker Diarization
2024/10/09 by Plaquet, Alexis, Tawara, Naohiro, Delcroix, Marc +3 · 3 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Discriminative Training of VBx Diarization
2023/10/04 by Klement, Dominik, Diez, Mireia, Landini, Federico +4 · 2 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Iterative Shallow Fusion of Backward Language Model for End-to-End Speech Recognition
2023/10/17 by Atsunori Ogawa, Ogawa, Atsunori, Takafumi Moriya +7 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- BLSTM-Based Confidence Estimation for End-to-End Speech Recognition
2023/12/22 by Ogawa, Atsunori, Tawara, Naohiro, Kano, Takatomo +1 · 1 citation
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
- NTT speaker diarization system for CHiME-7: multi-domain, multi-microphone End-to-end and vector clustering diarization
2023/09/22 by Tawara, Naohiro, Delcroix, Marc, Ando, Atsushi +1 · 1 citation
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Multi-Stream Extension of Variational Bayesian HMM Clustering (MS-VBx) for Combined End-to-End and Vector Clustering-based Diarization
2023/05/23 by Delcroix, Marc, Tawara, Naohiro, Diez, Mireia +6 · 1 citation
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Applying LLMs for Rescoring N-best ASR Hypotheses of Casual Conversations: Effects of Domain Adaptation and Context Carry-over
2024/06/27 by Ogawa, Atsunori, Kamo, Naoyuki, Matsuura, Kohei +5 · 1 citation
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
- NTT Multi-Speaker ASR System for the DASR Task of CHiME-8 Challenge
2024/09/09 by Naoyuki Kamo, Kamo, Naoyuki, Naohiro Tawara +33 · 2 citations
Computer Science · Engineering · #Advanced Data Processing Techniques #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Robotics and Automated Systems #Speech Recognition and Synthesis #electronic engineering #information engineering