vix.ing · top · new · best · stats · spec

Tawara, Naohiro

  1. Improving speaker discrimination of target speech extraction with time-domain SpeakerBeam
    2020/01/23 by Marc Delcroix, Delcroix, Marc, Tsubasa Ochiai +11 · 10 citations
    Computer Science · #Speech and Audio Processing #Music and Audio Processing #Speech Recognition and Synthesis
  2. Advances in integration of end-to-end neural and clustering-based diarization for real conversational speech
    2021/05/19 by Kinoshita, Keisuke, Delcroix, Marc, Tawara, Naohiro · 8 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  3. Integrating end-to-end neural and clustering-based diarization: Getting the best of both worlds
    2020/10/26 by Keisuke Kinoshita, Marc Delcroix, Kinoshita, Keisuke +3 · 6 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
  4. Interaural time difference loss for binaural target sound extraction
    2024/08/01 by Carlos Hernandez-Olivan, Hernandez-Olivan, Carlos, Marc Delcroix +9 · 3 citations
    Computer Science · Earth and Planetary Sciences · Engineering · #Advanced SAR Imaging Techniques #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech and Audio Processing #Underwater Acoustics Research #electronic engineering #information engineering
  5. SoundBeam meets M2D: Target Sound Extraction with Audio Foundation Model
    2024/09/19 by Carlos Hernandez-Olivan, Marc Delcroix, Hernandez-Olivan, Carlos +11 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  6. Mamba-based Segmentation Model for Speaker Diarization
    2024/10/09 by Plaquet, Alexis, Tawara, Naohiro, Delcroix, Marc +3 · 3 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  7. Discriminative Training of VBx Diarization
    2023/10/04 by Klement, Dominik, Diez, Mireia, Landini, Federico +4 · 2 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  8. Iterative Shallow Fusion of Backward Language Model for End-to-End Speech Recognition
    2023/10/17 by Atsunori Ogawa, Ogawa, Atsunori, Takafumi Moriya +7 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  9. BLSTM-Based Confidence Estimation for End-to-End Speech Recognition
    2023/12/22 by Ogawa, Atsunori, Tawara, Naohiro, Kano, Takatomo +1 · 1 citation
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
  10. NTT speaker diarization system for CHiME-7: multi-domain, multi-microphone End-to-end and vector clustering diarization
    2023/09/22 by Tawara, Naohiro, Delcroix, Marc, Ando, Atsushi +1 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  11. Multi-Stream Extension of Variational Bayesian HMM Clustering (MS-VBx) for Combined End-to-End and Vector Clustering-based Diarization
    2023/05/23 by Delcroix, Marc, Tawara, Naohiro, Diez, Mireia +6 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  12. Applying LLMs for Rescoring N-best ASR Hypotheses of Casual Conversations: Effects of Domain Adaptation and Context Carry-over
    2024/06/27 by Ogawa, Atsunori, Kamo, Naoyuki, Matsuura, Kohei +5 · 1 citation
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
  13. NTT Multi-Speaker ASR System for the DASR Task of CHiME-8 Challenge
    2024/09/09 by Naoyuki Kamo, Kamo, Naoyuki, Naohiro Tawara +33 · 2 citations
    Computer Science · Engineering · #Advanced Data Processing Techniques #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Robotics and Automated Systems #Speech Recognition and Synthesis #electronic engineering #information engineering