vix.ing · top · new · best · stats · spec

Yuki Saito

  1. The T05 System for The VoiceMOS Challenge 2024: Transfer Learning from Deep Image Classifier to Naturalness MOS Prediction of High-Quality Synthetic Speech
    2024/09/14 by Wataru Nakata, Baba, Kaito, Nakata, Wataru +4 · 27 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  2. On permutation-invariant neural networks
    2024/03/26 by Masanari Kimura, Ryotaro Shimizu, Kimura, Masanari +7 · 4 citations
    Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Neural Networks and Applications
  3. Voice Conversion Using Sequence-to-Sequence Learning of Context Posterior Probabilities
    2017/04/10 by Hiroyuki Miyoshi, Yuki Saito, Miyoshi, Hiroyuki +5 · 2 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Topic Modeling
  4. Statistical Parametric Speech Synthesis Incorporating Generative Adversarial Networks
    2017/09/23 by Yuki Saito, Saito, Yuki, Shinnosuke Takamichi +3 · 1 citation
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  5. STUDIES: Corpus of Japanese Empathetic Dialogue Speech Towards Friendly Voice Agent
    2022/03/28 by Yuki Saito, Yuto Nishimura, Saito, Yuki +7 · 1 citation
    Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG) #Natural Language Processing Techniques #Social Robot Interaction and HRI #Sound (cs.SD) #Speech and dialogue systems
  6. A Fashion Item Recommendation Model in Hyperbolic Space
    2024/09/04 by Ryotaro Shimizu, Yu Wang, Shimizu, Ryotaro +11 · 2 citations
    Arts and Humanities · Business, Management and Accounting · #Computer Vision and Pattern Recognition (cs.CV) #Consumer Perception and Purchasing Behavior #Cultural and Historical Studies #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Machine Learning (cs.LG)
  7. Mid-attribute speaker generation using optimal-transport-based interpolation of Gaussian mixture models
    2022/10/18 by Aya Watanabe, Shinnosuke Takamichi, Watanabe, Aya +7 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  8. Cross-Dialect Text-To-Speech in Pitch-Accent Language Incorporating Multi-Dialect Phoneme-Level BERT
    2024/09/11 by Kazuki Yamauchi, Yuki Saito, Yamauchi, Kazuki +3 · 2 citations
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  9. V2S attack: building DNN-based voice conversion from automatic speaker verification
    2019/08/05 by Taiki Nakamura, Yuki Saito, Nakamura, Taiki +7 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  10. An Empirical Analysis of GPT-4V's Performance on Fashion Aesthetic Evaluation
    2024/10/31 by Yuki Hirakawa, Hirakawa, Yuki, Takashi Wada +10 · 1 citation
    Arts and Humanities · Business, Management and Accounting · Social Sciences · #Computer Vision and Pattern Recognition (cs.CV) #Consumer Perception and Purchasing Behavior #Cultural and Historical Studies #Diverse Topics in Contemporary Research #FOS: Computer and information sciences
  11. Disentangling Likes and Dislikes in Personalized Generative Explainable Recommendation
    2024/10/17 by Ryotaro Shimizu, Shimizu, Ryotaro, Takashi Wada +21 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Machine Learning (cs.LG) #Topic Modeling
  12. Causal Speech Enhancement with Predicting Semantics based on Quantized Self-supervised Learning Features
    2024/12/26 by Emiru Tsunoo, Yuki Saito, Tsunoo, Emiru +5 · 1 citation
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing
  13. DNN-based Speaker Embedding Using Subjective Inter-speaker Similarity for Multi-speaker Modeling in Speech Synthesis
    2019/07/19 by Yuki Saito, Shinnosuke Takamichi, Saito, Yuki +3 · 1 citation
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing