vix.ing · top · new · best · stats · spec

Nanxin Chen

  1. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
    2024/03/08 by Gemini Robotics Team, Gemini Team, Petko Georgiev +2277 · 4 voices · 645 citations
    Computer Science · #Semantic Web and Ontologies
  2. WaveGrad: Estimating Gradients for Waveform Generation
    2020/09/02 by Nanxin Chen, Yu Zhang, Chen, Nanxin +9 · 1 voice · 48 citations
    Computer Science · Engineering · Mathematics · Physics and Astronomy · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Model Reduction and Neural Networks #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #cs.LG #cs.SD #eess.AS #electronic engineering #information engineering #stat.ML
  3. ESPnet: End-to-End Speech Processing Toolkit
    2018/03/30 by Shinji Watanabe, Takaaki Hori, Watanabe, Shinji +21 · 74 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Music and Audio Processing #Natural Language Processing Techniques #Speech Recognition and Synthesis
  4. Gemini: A Family of Highly Capable Multimodal Models
    2023/12/19 by Gemini Robotics Team, Gemini Team, Rohan Anil +2692 · 9 voices · 7 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #cs.AI #cs.CL #cs.CV
  5. Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages
    2023/03/02 by Yu Zhang, Zhang, Yu, Wei Han +51 · 30 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  6. Noise2Music: Text-conditioned Music Generation with Diffusion Models
    2023/02/08 by Qingqing Huang, Huang, Qingqing, Daniel Park +25 · 16 citations
    Computer Science · #Music and Audio Processing #Music Technology and Sound Studies #Speech Recognition and Synthesis
  7. SLM: Bridge the thin gap between speech and text foundation models
    2023/09/30 by Mingqiu Wang, Wang, Mingqiu, Wei Han +33 · 11 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  8. Zero-Shot Multi-Speaker Text-To-Speech with State-of-the-art Neural Speaker Embeddings
    2019/10/23 by Erica Cooper, Cheng-I Lai, Cooper, Erica +11 · 6 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Speech Recognition and Synthesis #electronic engineering #information engineering
  9. Mask CTC: Non-Autoregressive End-to-End ASR with CTC and Mask Predict
    2020/05/18 by Yosuke Higuchi, Higuchi, Yosuke, Shinji Watanabe +7 · 6 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  10. x-vectors meet emotions: A study on dependencies between emotion and\n speaker recognition
    2020/02/12 by Raghavendra Pappagari, Pappagari, Raghavendra, Tianzi Wang +7 · 5 citations
    Computer Science · Psychology · #Speech Recognition and Synthesis #Emotion and Mood Recognition #Speech and Audio Processing
  11. ASSERT: Anti-Spoofing with Squeeze-Excitation and Residual neTworks
    2019/04/01 by Cheng-I Lai, Nanxin Chen, Lai, Cheng-I +5 · 7 citations
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Topic Modeling
  12. E3 TTS: Easy End-to-End Diffusion-based Text to Speech
    2023/11/02 by Yuan Gao, Nobuyuki Morioka, Gao, Yuan +5 · 5 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering