Sheng Zhao
- Scalability in Perception for Autonomous Driving: Waymo Open Dataset
2019/12/10 by Pei Sun, Sun, Pei, Henrik Kretzschmar +47 · 285 citations
Computer Science · Engineering · #Advanced Neural Network Applications #Autonomous Vehicle Technology and Safety #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Video Surveillance and Tracking Methods
- FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
2020/06/08 by Yi Ren, Ren, Yi, Chenxu Hu +10 · 71 citations
Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
- NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
2024/03/05 by Zeqian Ju, Ju, Zeqian, Yuancheng Wang +35 · 69 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
2024/06/26 by Şefik Emre Eskimez, Sefik Emre Eskimez, Eskimez, Sefik Emre +24 · 2 voices · 44 citations
Engineering · Neuroscience · #Brain Tumor Detection and Classification #Industrial Vision Systems and Defect Detection #Ultrasonics and Acoustic Wave Propagation #cs.SD #eess.AS
- Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling
2023/03/07 by Ziqiang Zhang, Long Zhou, Zhang, Ziqiang +23 · 22 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- NaturalSpeech: End-to-End Text to Speech Synthesis with Human-Level Quality
2022/05/09 by Xu Tan, Jiawei Chen, Tan, Xu +25 · 15 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
- Autoregressive Speech Synthesis without Vector Quantization
2024/07/11 by Lingwei Meng, Long Zhou, Meng, Lingwei +21 · 23 citations
Computer Science · #Speech Recognition and Synthesis #Speech and dialogue systems
- AdaSpeech: Adaptive Text to Speech for Custom Voice
2021/03/01 by Mingjian Chen, Chen, Mingjian, Xu Tan +11 · 10 citations
Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
- Almost Unsupervised Text to Speech and Automatic Speech Recognition
2019/05/13 by Yi Ren, Ren, Yi, Xu Tan +9 · 7 citations
Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Speech and Audio Processing
- PromptTTS 2: Describing and Generating Voices with Text Prompt
2023/09/05 by Yichong Leng, Leng, Yichong, Zhifang Guo +27 · 7 citations
Computer Science · #Speech Recognition and Synthesis #Topic Modeling #Natural Language Processing Techniques
- FastSpeech: Fast, Robust and Controllable Text to Speech
2019/05/22 by Yi Ren, Ren, Yi, Yangjun Ruan +12 · 1 voice · 2 citations
Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #cs.CL #cs.LG #cs.SD #eess.AS #electronic engineering #information engineering
- Laugh Now Cry Later: Controlling Time-Varying Emotional States of Flow-Matching-Based Zero-Shot Text-to-Speech
2024/07/17 by Haibin Wu, Wu, Haibin, Xiaofei Wang +19 · 5 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Signal Processing (eess.SP) #Speech Recognition and Synthesis #electronic engineering #information engineering
- A Study of Non-autoregressive Model for Sequence Generation
2020/04/22 by Yi Ren, Jinglin Liu, Ren, Yi +9 · 2 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling #Speech Recognition and Synthesis
- HiFace: High-Fidelity 3D Face Reconstruction by Learning Static and Dynamic Details
2023/03/20 by Zenghao Chai, Tianke Zhang, Chai, Zenghao +17 · 3 citations
Computer Science · Engineering · #Face recognition and analysis #Generative Adversarial Networks and Image Synthesis #3D Shape Modeling and Analysis
- InferGrad: Improving Diffusion Models for Vocoder by Considering Inference in Training
2022/02/08 by Zehua Chen, Xu Tan, Chen, Zehua +11 · 2 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- TransVIP: Speech to Speech Translation System with Voice and Isochrony Preservation
2024/05/28 by Chenyang Le, Le, Chenyang, Yao Qian +19 · 3 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Low-voltage broadband piezoelectric vibration energy harvesting enabled by a highly-coupled harvester and tunable PSSHI circuit
2021/10/27 by Sheng Zhao, Ujwal Radhakrishna, Jeffrey H Lang +3 · 2 citations
Engineering · #Innovative Energy Harvesting Technologies #Energy Harvesting in Wireless Networks #Advanced Sensor and Energy Harvesting Materials
- DenoiSpeech: Denoising Text to Speech with Frame-Level Noise Modeling
2020/12/17 by Chen Zhang, Zhang, Chen, Yi Ren +13 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- DelightfulTTS 2: End-to-End Speech Synthesis with Adversarial Vector-Quantized Auto-Encoders
2022/07/11 by Yanqing Liu, Liu, Yanqing, Ruiqing Xue +7 · 2 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Natural Language Processing Techniques
- A Light-weight contextual spelling correction model for customizing transducer-based speech recognition systems
2021/08/17 by Xiaoqiang Wang, Yanqing Liu, Wang, Xiaoqiang +5 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Towards Contextual Spelling Correction for Customization of End-to-end Speech Recognition Systems
2022/03/02 by Xiaoqiang Wang, Wang, Xiaoqiang, Yanqing Liu +9 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Investigating Neural Audio Codecs for Speech Language Model-Based Speech Generation
2024/09/06 by Jiaqi Li, Li, Jiaqi, Dongmei Wang +29 · 3 citations
Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing
- MeloForm: Generating Melody with Musical Form based on Expert Systems and Neural Networks
2022/08/30 by Peiling Lu, Xu Tan, Lu, Peiling +9 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
- Scarless wound healing programmed by core-shell microneedles
2023/06/10 by Ying Zhang, Yinxian Yang, Sheng Zhao +8 · 1 citation
Pharmacology, Toxicology and Pharmaceutics · Engineering · Medicine · #Advancements in Transdermal Drug Delivery #Optical Coherence Tomography Applications #Wound Healing and Treatments
- Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners
2024/12/06 by Ze Yuan, Yanqing Liu, Yuan, Ze +5 · 3 citations
Computer Science · #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and dialogue systems
- Mixed-Phoneme BERT: Improving BERT with Mixed Phoneme and Sup-Phoneme Representations for Text to Speech
2022/03/31 by Guangyan Zhang, Kaitao Song, Zhang, Guangyan +19 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Magnon blockade in spin-magnon systems with frequency detuning
2025/05/27 by Sheng Zhao, Ya-long Ren, Zhao, Sheng +7 · 2 citations
Materials Science · Physics and Astronomy · #FOS: Physical sciences #Magnetic and transport properties of perovskites and related materials #Mechanical and Optical Resonators #Mesoscale and Nanoscale Physics (cond-mat.mes-hall) #Quantum Physics (quant-ph) #Quantum and electron transport phenomena
- MoBoAligner: a Neural Alignment Model for Non-autoregressive TTS with Monotonic Boundary Search
2020/05/18 by Naihan Li, Shujie Liu, Li, Naihan +9 · 1 citation
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
- An Investigation of Noise Robustness for Flow-Matching-Based Zero-Shot TTS
2024/06/09 by Xiaofei Wang, Şefik Emre Eskimez, Wang, Xiaofei +19 · 1 citation
Engineering · Computer Science · #Advanced Adaptive Filtering Techniques #Speech and Audio Processing #Blind Source Separation Techniques
- Steeringless Drifting: Differential-Torque Control of a Four-Wheel Independently Driven Vehicle
2026/07/26 by Sheng Zhao, Zexin Wu, Dongyang Zhou +2
#cs.RO
- LowPowAR: Power-Constrained Tone Mapping for Augmented Reality
2026/07/21 by Weikai Lin, Sheng Zhao, Ian Ross +3
#cs.GR