vix.ing · top · new · best · stats · spec

Tomoki Hayashi

  1. ESPnet: End-to-End Speech Processing Toolkit
    2018/03/30 by Shinji Watanabe, Watanabe, Shinji, Takaaki Hori +21 · 84 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Music and Audio Processing #Natural Language Processing Techniques #Speech Recognition and Synthesis
  2. ESPnet-TTS: Unified, Reproducible, and Integratable Open Source End-to-End Text-to-Speech Toolkit
    2019/10/24 by Tomoki Hayashi, Ryuichi Yamamoto, Hayashi, Tomoki +15 · 13 citations
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Music and Audio Processing
  3. Any-to-One Sequence-to-Sequence Voice Conversion using Self-Supervised Discrete Speech Representations
    2020/10/23 by Wen-Chin Huang, Huang, Wen-Chin, Yi-Chiao Wu +5 · 5 citations
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Music and Audio Processing
  4. Muskits: an End-to-End Music Processing Toolkit for Singing Voice Synthesis
    2022/05/09 by Jiatong Shi, Shuai Guo, Shi, Jiatong +21 · 6 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
  5. Non-Parallel Voice Conversion with Cyclic Variational Autoencoder
    2019/07/24 by Patrick Lumban Tobing, Tobing, Patrick Lumban, Yi-Chiao Wu +7 · 4 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Advanced Data Compression Techniques
  6. ESPnet2-TTS: Extending the Edge of TTS Research
    2021/10/15 by Tomoki Hayashi, Hayashi, Tomoki, Ryuichi Yamamoto +17 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
  7. DiscreTalk: Text-to-Speech as a Machine Translation Problem
    2020/05/12 by Tomoki Hayashi, Shinji Watanabe, Hayashi, Tomoki +1 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  8. Improvement of Serial Approach to Anomalous Sound Detection by Incorporating Two Binary Cross-Entropies for Outlier Exposure
    2022/06/13 by Ibuki Kuroyanagi, Tomoki Hayashi, Kuroyanagi, Ibuki +5 · 2 citations
    Computer Science · #Anomaly Detection Techniques and Applications #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  9. The 2020 ESPnet update: new features, broadened applications, performance improvements, and future plans
    2020/12/23 by Shinji Watanabe, Florian Boyer, Watanabe, Shinji +27 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  10. Cycle-consistency training for end-to-end speech recognition
    2018/11/02 by Takaaki Hori, Ramón Fernández Astudillo, Hori, Takaaki +9 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  11. End-to-End Automatic Speech Recognition Integrated With CTC-Based Voice Activity Detection
    2020/02/03 by Takenori Yoshimura, Tomoki Hayashi, Yoshimura, Takenori +5 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  12. Non-autoregressive sequence-to-sequence voice conversion
    2021/04/14 by Tomoki Hayashi, Wen-Chin Huang, Hayashi, Tomoki +5 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  13. S3PRL-VC: Open-source Voice Conversion Framework with Self-supervised Speech Representations
    2021/10/12 by Wen-Chin Huang, Shuwen Yang, Huang, Wen-Chin +9 · 1 citation
    Computer Science · #Speech Recognition and Synthesis #Topic Modeling #Natural Language Processing Techniques
  14. Pretraining Techniques for Sequence-to-Sequence Voice Conversion
    2020/08/07 by Wen-Chin Huang, Huang, Wen-Chin, Tomoki Hayashi +7 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  15. Voice Transformer Network: Sequence-to-Sequence Voice Conversion Using Transformer with Text-to-Speech Pretraining
    2019/12/14 by Wen-Chin Huang, Tomoki Hayashi, Huang, Wen-Chin +7 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  16. Quasi-Periodic WaveNet Vocoder: A Pitch Dependent Dilated Convolution Model for Parametric Speech Generation
    2019/07/01 by Yi-Chiao Wu, Wu, Yi-Chiao, Tomoki Hayashi +7 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering