vix.ing · top · new · best · stats · spec

Sheng Zhao

  1. Scalability in Perception for Autonomous Driving: Waymo Open Dataset
    2019/12/10 by Pei Sun, Sun, Pei, Henrik Kretzschmar +47 · 285 citations
    Computer Science · Engineering · #Advanced Neural Network Applications #Autonomous Vehicle Technology and Safety #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Video Surveillance and Tracking Methods
  2. FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
    2020/06/08 by Yi Ren, Ren, Yi, Chenxu Hu +10 · 71 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
  3. NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
    2024/03/05 by Zeqian Ju, Ju, Zeqian, Yuancheng Wang +35 · 69 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  4. E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
    2024/06/26 by Şefik Emre Eskimez, Sefik Emre Eskimez, Eskimez, Sefik Emre +24 · 2 voices · 44 citations
    Engineering · Neuroscience · #Brain Tumor Detection and Classification #Industrial Vision Systems and Defect Detection #Ultrasonics and Acoustic Wave Propagation #cs.SD #eess.AS
  5. Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling
    2023/03/07 by Ziqiang Zhang, Long Zhou, Zhang, Ziqiang +23 · 22 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  6. NaturalSpeech: End-to-End Text to Speech Synthesis with Human-Level Quality
    2022/05/09 by Xu Tan, Jiawei Chen, Tan, Xu +25 · 15 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
  7. Autoregressive Speech Synthesis without Vector Quantization
    2024/07/11 by Lingwei Meng, Long Zhou, Meng, Lingwei +21 · 23 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and dialogue systems
  8. AdaSpeech: Adaptive Text to Speech for Custom Voice
    2021/03/01 by Mingjian Chen, Chen, Mingjian, Xu Tan +11 · 10 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
  9. Almost Unsupervised Text to Speech and Automatic Speech Recognition
    2019/05/13 by Yi Ren, Ren, Yi, Xu Tan +9 · 7 citations
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Speech and Audio Processing
  10. PromptTTS 2: Describing and Generating Voices with Text Prompt
    2023/09/05 by Yichong Leng, Leng, Yichong, Zhifang Guo +27 · 7 citations
    Computer Science · #Speech Recognition and Synthesis #Topic Modeling #Natural Language Processing Techniques
  11. FastSpeech: Fast, Robust and Controllable Text to Speech
    2019/05/22 by Yi Ren, Ren, Yi, Yangjun Ruan +12 · 1 voice · 2 citations
    Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #cs.CL #cs.LG #cs.SD #eess.AS #electronic engineering #information engineering
  12. Laugh Now Cry Later: Controlling Time-Varying Emotional States of Flow-Matching-Based Zero-Shot Text-to-Speech
    2024/07/17 by Haibin Wu, Wu, Haibin, Xiaofei Wang +19 · 5 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Signal Processing (eess.SP) #Speech Recognition and Synthesis #electronic engineering #information engineering
  13. A Study of Non-autoregressive Model for Sequence Generation
    2020/04/22 by Yi Ren, Jinglin Liu, Ren, Yi +9 · 2 citations
    Computer Science · #Natural Language Processing Techniques #Topic Modeling #Speech Recognition and Synthesis
  14. HiFace: High-Fidelity 3D Face Reconstruction by Learning Static and Dynamic Details
    2023/03/20 by Zenghao Chai, Tianke Zhang, Chai, Zenghao +17 · 3 citations
    Computer Science · Engineering · #Face recognition and analysis #Generative Adversarial Networks and Image Synthesis #3D Shape Modeling and Analysis
  15. InferGrad: Improving Diffusion Models for Vocoder by Considering Inference in Training
    2022/02/08 by Zehua Chen, Xu Tan, Chen, Zehua +11 · 2 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  16. TransVIP: Speech to Speech Translation System with Voice and Isochrony Preservation
    2024/05/28 by Chenyang Le, Le, Chenyang, Yao Qian +19 · 3 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  17. Low-voltage broadband piezoelectric vibration energy harvesting enabled by a highly-coupled harvester and tunable PSSHI circuit
    2021/10/27 by Sheng Zhao, Ujwal Radhakrishna, Jeffrey H Lang +3 · 2 citations
    Engineering · #Innovative Energy Harvesting Technologies #Energy Harvesting in Wireless Networks #Advanced Sensor and Energy Harvesting Materials
  18. DenoiSpeech: Denoising Text to Speech with Frame-Level Noise Modeling
    2020/12/17 by Chen Zhang, Zhang, Chen, Yi Ren +13 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  19. DelightfulTTS 2: End-to-End Speech Synthesis with Adversarial Vector-Quantized Auto-Encoders
    2022/07/11 by Yanqing Liu, Liu, Yanqing, Ruiqing Xue +7 · 2 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Natural Language Processing Techniques
  20. A Light-weight contextual spelling correction model for customizing transducer-based speech recognition systems
    2021/08/17 by Xiaoqiang Wang, Yanqing Liu, Wang, Xiaoqiang +5 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  21. Towards Contextual Spelling Correction for Customization of End-to-end Speech Recognition Systems
    2022/03/02 by Xiaoqiang Wang, Wang, Xiaoqiang, Yanqing Liu +9 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  22. Investigating Neural Audio Codecs for Speech Language Model-Based Speech Generation
    2024/09/06 by Jiaqi Li, Li, Jiaqi, Dongmei Wang +29 · 3 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing
  23. MeloForm: Generating Melody with Musical Form based on Expert Systems and Neural Networks
    2022/08/30 by Peiling Lu, Xu Tan, Lu, Peiling +9 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  24. Scarless wound healing programmed by core-shell microneedles
    2023/06/10 by Ying Zhang, Yinxian Yang, Sheng Zhao +8 · 1 citation
    Pharmacology, Toxicology and Pharmaceutics · Engineering · Medicine · #Advancements in Transdermal Drug Delivery #Optical Coherence Tomography Applications #Wound Healing and Treatments
  25. Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners
    2024/12/06 by Ze Yuan, Yanqing Liu, Yuan, Ze +5 · 3 citations
    Computer Science · #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and dialogue systems
  26. Mixed-Phoneme BERT: Improving BERT with Mixed Phoneme and Sup-Phoneme Representations for Text to Speech
    2022/03/31 by Guangyan Zhang, Kaitao Song, Zhang, Guangyan +19 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  27. Magnon blockade in spin-magnon systems with frequency detuning
    2025/05/27 by Sheng Zhao, Ya-long Ren, Zhao, Sheng +7 · 2 citations
    Materials Science · Physics and Astronomy · #FOS: Physical sciences #Magnetic and transport properties of perovskites and related materials #Mechanical and Optical Resonators #Mesoscale and Nanoscale Physics (cond-mat.mes-hall) #Quantum Physics (quant-ph) #Quantum and electron transport phenomena
  28. MoBoAligner: a Neural Alignment Model for Non-autoregressive TTS with Monotonic Boundary Search
    2020/05/18 by Naihan Li, Shujie Liu, Li, Naihan +9 · 1 citation
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
  29. An Investigation of Noise Robustness for Flow-Matching-Based Zero-Shot TTS
    2024/06/09 by Xiaofei Wang, Şefik Emre Eskimez, Wang, Xiaofei +19 · 1 citation
    Engineering · Computer Science · #Advanced Adaptive Filtering Techniques #Speech and Audio Processing #Blind Source Separation Techniques
  30. Steeringless Drifting: Differential-Torque Control of a Four-Wheel Independently Driven Vehicle
    2026/07/26 by Sheng Zhao, Zexin Wu, Dongyang Zhou +2
    #cs.RO
  31. LowPowAR: Power-Constrained Tone Mapping for Augmented Reality
    2026/07/21 by Weikai Lin, Sheng Zhao, Ian Ross +3
    #cs.GR