Nanxin Chen
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
2024/03/08 by Gemini Robotics Team, Gemini Team, Petko Georgiev +2277 · 4 voices · 645 citations
Computer Science · #Semantic Web and Ontologies
- WaveGrad: Estimating Gradients for Waveform Generation
2020/09/02 by Nanxin Chen, Yu Zhang, Chen, Nanxin +9 · 1 voice · 48 citations
Computer Science · Engineering · Mathematics · Physics and Astronomy · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Model Reduction and Neural Networks #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #cs.LG #cs.SD #eess.AS #electronic engineering #information engineering #stat.ML
- ESPnet: End-to-End Speech Processing Toolkit
2018/03/30 by Shinji Watanabe, Takaaki Hori, Watanabe, Shinji +21 · 74 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Music and Audio Processing #Natural Language Processing Techniques #Speech Recognition and Synthesis
- Gemini: A Family of Highly Capable Multimodal Models
2023/12/19 by Gemini Robotics Team, Gemini Team, Rohan Anil +2692 · 9 voices · 7 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #cs.AI #cs.CL #cs.CV
- Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages
2023/03/02 by Yu Zhang, Zhang, Yu, Wei Han +51 · 30 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Noise2Music: Text-conditioned Music Generation with Diffusion Models
2023/02/08 by Qingqing Huang, Huang, Qingqing, Daniel Park +25 · 16 citations
Computer Science · #Music and Audio Processing #Music Technology and Sound Studies #Speech Recognition and Synthesis
- SLM: Bridge the thin gap between speech and text foundation models
2023/09/30 by Mingqiu Wang, Wang, Mingqiu, Wei Han +33 · 11 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Zero-Shot Multi-Speaker Text-To-Speech with State-of-the-art Neural Speaker Embeddings
2019/10/23 by Erica Cooper, Cheng-I Lai, Cooper, Erica +11 · 6 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Speech Recognition and Synthesis #electronic engineering #information engineering
- Mask CTC: Non-Autoregressive End-to-End ASR with CTC and Mask Predict
2020/05/18 by Yosuke Higuchi, Higuchi, Yosuke, Shinji Watanabe +7 · 6 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- x-vectors meet emotions: A study on dependencies between emotion and\n speaker recognition
2020/02/12 by Raghavendra Pappagari, Pappagari, Raghavendra, Tianzi Wang +7 · 5 citations
Computer Science · Psychology · #Speech Recognition and Synthesis #Emotion and Mood Recognition #Speech and Audio Processing
- ASSERT: Anti-Spoofing with Squeeze-Excitation and Residual neTworks
2019/04/01 by Cheng-I Lai, Nanxin Chen, Lai, Cheng-I +5 · 7 citations
Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Topic Modeling
- E3 TTS: Easy End-to-End Diffusion-based Text to Speech
2023/11/02 by Yuan Gao, Nobuyuki Morioka, Gao, Yuan +5 · 5 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering