vix.ing · top · new · best · stats · spec

Tan, Xu

  1. HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face
    2023/03/30 by Yongliang Shen, Shen, Yongliang, Kaitao Song +10 · 2 voices · 173 citations
    Computer Science · Engineering · #Topic Modeling #Ferroelectric and Negative Capacitance Devices #Multimodal Machine Learning Applications
  2. MPNet: Masked and Permuted Pre-training for Language Understanding
    2020/04/20 by Kaitao Song, Song, Kaitao, Xu Tan +7 · 124 citations
    Computer Science · #Topic Modeling #Natural Language Processing Techniques #Speech Recognition and Synthesis
  3. FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
    2020/06/08 by Yi Ren, Chenxu Hu, Ren, Yi +10 · 70 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
  4. EvoPrompt: Connecting LLMs with Evolutionary Algorithms Yields Powerful Prompt Optimizers
    2023/09/15 by Qingyan Guo, Guo, Qingyan, Rui Wang +15 · 83 citations
    Computer Science · #Topic Modeling #Natural Language Processing Techniques #Speech and dialogue systems
  5. BRECQ: Pushing the Limit of Post-Training Quantization by Block Reconstruction
    2021/02/10 by Li, Yuhang, Gong, Ruihao, Tan, Xu +6 · 43 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  6. NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
    2024/03/05 by Zeqian Ju, Ju, Zeqian, Yuancheng Wang +35 · 68 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  7. E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
    2024/06/26 by Sefik Emre Eskimez, Şefik Emre Eskimez, Eskimez, Sefik Emre +24 · 2 voices · 44 citations
    Engineering · Neuroscience · #Brain Tumor Detection and Classification #Industrial Vision Systems and Defect Detection #Ultrasonics and Acoustic Wave Propagation #cs.SD #eess.AS
  8. NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
    2023/04/18 by Shen, Kai, Ju, Zeqian, Tan, Xu +6 · 39 citations
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  9. Beyond Language Models: Byte Models are Digital World Simulators
    2024/02/29 by Shangda Wu, Xu Tan, Wu, Shangda +9 · 2 voices · 5 citations
    Computer Science · Social Sciences · #FOS: Computer and information sciences #Language and cultural evolution #Machine Learning (cs.LG) #Natural Language Processing Techniques
  10. Representation Degeneration Problem in Training Natural Language Generation Models
    2019/07/27 by Jun Gao, Di He, Gao, Jun +9 · 20 citations
    Computer Science · #Topic Modeling #Natural Language Processing Techniques #Multimodal Machine Learning Applications
  11. A Survey on Neural Speech Synthesis
    2021/06/29 by Xu Tan, Tan, Xu, Tao Qin +5 · 22 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimedia (cs.MM) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  12. Kimi-Audio Technical Report
    2025/04/25 by KimiTeam, Ding Ding, Zeqian Ju +70 · 83 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
  13. VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
    2024/06/08 by Chen, Sanyuan, Liu, Shujie, Zhou, Long +6 · 39 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  14. UniAudio: An Audio Foundation Model Toward Universal Audio Generation
    2023/10/01 by Yang, Dongchao, Tian, Jinchuan, Tan, Xu +9 · 28 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  15. Achieving Human Parity on Automatic Chinese to English News Translation
    2018/03/15 by Hany Hassan, Anthony Aue, Hassan, Hany +45 · 27 citations
    Computer Science · #Natural Language Processing Techniques #Topic Modeling #Multimodal Machine Learning Applications
  16. AUDIT: Audio Editing by Following Instructions with Latent Diffusion Models
    2023/04/03 by Wang, Yuancheng, Ju, Zeqian, Tan, Xu +4 · 23 citations
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  17. MusicBERT: Symbolic Music Understanding with Large-Scale Pre-Training
    2021/06/10 by Mingliang Zeng, Xu Tan, Zeng, Mingliang +9 · 15 citations
    Computer Science · Neuroscience · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Information Retrieval (cs.IR) #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Neuroscience and Music Perception #Sound (cs.SD) #electronic engineering #information engineering
  18. EASYTOOL: Enhancing LLM-based Agents with Concise Tool Instruction
    2024/01/11 by Yuan, Siyu, Song, Kaitao, Chen, Jiangjie +5 · 24 citations
    #Computation and Language (cs.CL) #FOS: Computer and information sciences
  19. Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language Model
    2024/08/30 by Ye, Zhen, Sun, Peiwen, Lei, Jiahe +9 · 32 citations
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  20. PromptTTS: Controllable Text-to-Speech with Text Descriptions
    2022/11/22 by Guo, Zhifang, Leng, Yichong, Wu, Yihan +2 · 17 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  21. MASS: Masked Sequence to Sequence Pre-training for Language Generation
    2019/05/07 by Kaitao Song, Xu Tan, Song, Kaitao +7 · 1 voice · 9 citations
    #cs.CL #cs.AI #cs.LG
  22. NaturalSpeech: End-to-End Text to Speech Synthesis with Human-Level Quality
    2022/05/09 by Xu Tan, Jiawei Chen, Tan, Xu +25 · 15 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
  23. BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec
    2024/09/09 by Detai Xin, Xin, Detai, Xu Tan +5 · 26 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  24. EvoAgent: Towards Automatic Multi-Agent Generation via Evolutionary Algorithms
    2024/06/20 by Siyu Yuan, Kaitao Song, Yuan, Siyu +9 · 21 citations
    Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Multi-Agent Systems and Negotiation #Reinforcement Learning in Robotics #Robotic Path Planning Algorithms
  25. Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis
    2025/02/06 by Ye, Zhen, Zhu, Xinfa, Chan, Chi-Min +17 · 38 citations
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Sound (cs.SD) #electronic engineering #information engineering
  26. ESPnet-TTS: Unified, Reproducible, and Integratable Open Source End-to-End Text-to-Speech Toolkit
    2019/10/24 by Tomoki Hayashi, Ryuichi Yamamoto, Hayashi, Tomoki +15 · 9 citations
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Music and Audio Processing
  27. TaskBench: Benchmarking Large Language Models for Task Automation
    2023/11/30 by Yongliang Shen, Kaitao Song, Shen, Yongliang +15 · 15 citations
    Computer Science · Medicine · Materials Science · #Topic Modeling #Artificial Intelligence in Healthcare and Education #Machine Learning in Materials Science
  28. XiaoiceSing: A High-Quality and Integrated Singing Voice Synthesis System
    2020/06/11 by Lu, Peiling, Wu, Jie, Luan, Jian +2 · 8 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  29. PopMAG: Pop Music Accompaniment Generation
    2020/08/18 by Ren, Yi, He, Jinzheng, Tan, Xu +3 · 7 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Information Retrieval (cs.IR) #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  30. FastCorrect: Fast Error Correction with Edit Alignment for Automatic Speech Recognition
    2021/05/09 by Yichong Leng, Xu Tan, Leng, Yichong +17 · 7 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  31. AdaSpeech: Adaptive Text to Speech for Custom Voice
    2021/03/01 by Mingjian Chen, Chen, Mingjian, Xu Tan +11 · 10 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
  32. PriorGrad: Improving Conditional Denoising Diffusion Models with Data-Dependent Adaptive Prior
    2021/06/11 by Lee, Sang-gil, Kim, Heeseung, Shin, Chaehun +7 · 7 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Sound (cs.SD) #electronic engineering #information engineering
  33. HiFiSinger: Towards High-Fidelity Neural Singing Voice Synthesis
    2020/09/03 by Jiawei Chen, Chen, Jiawei, Xu Tan +7 · 11 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
  34. MuseCoco: Generating Symbolic Music from Text
    2023/05/31 by Peiling Lu, Lu, Peiling, Xin Xu +11 · 9 citations
    Computer Science · #Music and Audio Processing #Music Technology and Sound Studies #Speech Recognition and Synthesis
  35. YuE: Scaling Open Foundation Models for Long-Form Music Generation
    2025/03/11 by Ruibin Yuan, Yuan, Ruibin, Lin, Hanfeng +110 · 23 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computer Graphics and Visualization Techniques #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  36. BinauralGrad: A Two-Stage Conditional Diffusion Probabilistic Model for Binaural Audio Synthesis
    2022/05/30 by Leng, Yichong, Chen, Zehua, Guo, Junliang +9 · 7 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  37. Museformer: Transformer with Fine- and Coarse-Grained Attention for Music Generation
    2022/10/19 by Botao Yu, Peiling Lu, Yu, Botao +15 · 7 citations
    Computer Science · Neuroscience · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Neuroscience and Music Perception #Sound (cs.SD) #electronic engineering #information engineering
  38. Empowering Diffusion Models on the Embedding Space for Text Generation
    2022/12/19 by Gao, Zhujin, Guo, Junliang, Tan, Xu +4 · 7 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  39. DeepSinger: Singing Voice Synthesis with Data Mined From the Web
    2020/07/09 by Ren, Yi, Tan, Xu, Qin, Tao +3 · 5 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  40. MBNet: MOS Prediction for Synthesized Speech with Mean-Bias Network
    2021/02/27 by Leng, Yichong, Tan, Xu, Zhao, Sheng +3 · 5 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  41. Almost Unsupervised Text to Speech and Automatic Speech Recognition
    2019/05/13 by Yi Ren, Xu Tan, Ren, Yi +9 · 7 citations
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Speech and Audio Processing
  42. MuPT: A Generative Symbolic Music Pretrained Transformer
    2024/04/09 by Xingwei Qu, Yuelin Bai, Qu, Xingwei +53 · 9 citations
    Computer Science · Neuroscience · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Neuroscience and Music Perception #Sound (cs.SD) #electronic engineering #information engineering
  43. UniAudio 1.5: Large Language Model-driven Audio Codec is A Few-shot Audio Task Learner
    2024/06/14 by Dongchao Yang, Haohan Guo, Yang, Dongchao +13 · 10 citations
    Computer Science · #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing
  44. MusicAgent: An AI Agent for Music Understanding and Generation with Large Language Models
    2023/10/18 by Dingyao Yu, Yu, Dingyao, Kaitao Song +13 · 7 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Music and Audio Processing #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  45. GAIA: Zero-shot Talking Avatar Generation
    2023/11/26 by He, Tianyu, Guo, Junliang, Yu, Runyi +10 · 7 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimedia (cs.MM)
  46. PromptTTS 2: Describing and Generating Voices with Text Prompt
    2023/09/05 by Yichong Leng, Zhifang Guo, Leng, Yichong +27 · 7 citations
    Computer Science · #Speech Recognition and Synthesis #Topic Modeling #Natural Language Processing Techniques
  47. FRAGE: Frequency-Agnostic Word Representation
    2018/09/18 by Chengyue Gong, Di He, Gong, Chengyue +9 · 4 citations
    Computer Science · #Topic Modeling #Natural Language Processing Techniques #Speech Recognition and Synthesis
  48. Schrodinger Bridges Beat Diffusion Models on Text-to-Speech Synthesis
    2023/12/06 by Zehua Chen, Guande He, Chen, Zehua +7 · 7 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  49. Foundation Models for Music: A Survey
    2024/08/26 by Yinghao Ma, Anders Øland, Ma, Yinghao +81 · 8 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music Technology and Sound Studies #Sound (cs.SD) #electronic engineering #information engineering
  50. MoonCast: High-Quality Zero-Shot Podcast Generation
    2025/03/18 by Ju, Zeqian, Yang, Dongchao, Yu, Jianwei +7 · 13 citations
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  51. CoMoSpeech: One-Step Speech and Singing Voice Synthesis via Consistency Model
    2023/05/11 by Ye, Zhen, Xue, Wei, Tan, Xu +3 · 5 citations
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimedia (cs.MM) #Sound (cs.SD) #electronic engineering #information engineering
  52. VideoDubber: Machine Translation with Speech-Aware Length Control for Video Dubbing
    2022/11/30 by Wu, Yihan, Guo, Junliang, Tan, Xu +7 · 4 citations
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimedia (cs.MM) #electronic engineering #information engineering
  53. SongMASS: Automatic Song Writing with Pre-training and Alignment Constraint
    2020/12/09 by Zhonghao Sheng, Sheng, Zhonghao, Kaitao Song +11 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Topic Modeling #electronic engineering #information engineering
  54. xTrimoPGLM: Unified 100B-Scale Pre-trained Transformer for Deciphering the Language of Protein
    2024/01/11 by Chen, Bo, Cheng, Xingyi, Li, Pan +12 · 5 citations
    #Artificial Intelligence (cs.AI) #FOS: Biological sciences #FOS: Computer and information sciences #Machine Learning (cs.LG) #Quantitative Methods (q-bio.QM)
  55. DelightfulTTS: The Microsoft Speech Synthesis System for Blizzard Challenge 2021
    2021/10/25 by Liu, Yanqing, Xu, Zhihang, Wang, Gang +6 · 3 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  56. AudioX: A Unified Framework for Anything-to-Audio Generation
    2025/03/13 by Tian, Zeyue, Zhaoyang Liu, Yizhu Jin +10 · 11 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG) #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  57. VidMuse: A Simple Video-to-Music Generation Framework with Long-Short-Term Modeling
    2024/06/06 by Zeyue Tian, Zhaoyang Liu, Tian, Zeyue +15 · 5 citations
    Computer Science · Social Sciences · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimedia (cs.MM) #Multimedia Communication and Technology #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD)
  58. D-CPT Law: Domain-specific Continual Pre-Training Scaling Law for Large Language Models
    2024/06/03 by Haoran Que, Que, Haoran, Jiaheng Liu +29 · 5 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling
  59. Multilingual Neural Machine Translation with Knowledge Distillation
    2019/02/27 by Tan, Xu, Ren, Yi, He, Di +3 · 2 citations
    #Computation and Language (cs.CL) #FOS: Computer and information sciences
  60. FastSpeech: Fast, Robust and Controllable Text to Speech
    2019/05/22 by Yi Ren, Ren, Yi, Yangjun Ruan +12 · 1 voice · 2 citations
    Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #cs.CL #cs.LG #cs.SD #eess.AS #electronic engineering #information engineering
  61. Multilingual Neural Machine Translation with Language Clustering
    2019/08/25 by Tan, Xu, Chen, Jiale, He, Di +3 · 2 citations
    #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  62. A Study of Non-autoregressive Model for Sequence Generation
    2020/04/22 by Yi Ren, Jinglin Liu, Ren, Yi +9 · 2 citations
    Computer Science · #Natural Language Processing Techniques #Topic Modeling #Speech Recognition and Synthesis
  63. UWSpeech: Speech to Speech Translation for Unwritten Languages
    2020/06/14 by Chen Zhang, Xu Tan, Zhang, Chen +9 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  64. LightSpeech: Lightweight and Fast Text to Speech with Neural Architecture Search
    2021/02/08 by Luo, Renqian, Tan, Xu, Wang, Rui +5 · 2 citations
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  65. Adaptive Logit Adjustment Loss for Long-Tailed Visual Recognition
    2021/04/13 by Yan Zhao, Zhao, Yan, Weicong Chen +7 · 2 citations
    Computer Science · #Advanced Image and Video Retrieval Techniques #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences
  66. InstructAvatar: Text-Guided Emotion and Motion Control for Avatar Generation
    2024/05/24 by Yuchi Wang, Wang, Yuchi, Junliang Guo +13 · 4 citations
    Engineering · Psychology · #Human Motion and Animation #Educational Games and Gamification #Social Robot Interaction and HRI
  67. Revisiting Over-Smoothness in Text to Speech
    2022/02/26 by Yi Ren, Xu Tan, Ren, Yi +7 · 2 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
  68. CLaMP 2: Multimodal Music Information Retrieval Across 101 Languages Using Large Language Models
    2024/10/17 by Shangda Wu, Wu, Shangda, Yashan Wang +27 · 4 citations
    Computer Science · Arts and Humanities · #Music and Audio Processing #Diverse Musicological Studies #Natural Language Processing Techniques
  69. InferGrad: Improving Diffusion Models for Vocoder by Considering Inference in Training
    2022/02/08 by Zehua Chen, Xu Tan, Chen, Zehua +11 · 2 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  70. Mitigating Reversal Curse in Large Language Models via Semantic-aware Permutation Training
    2024/03/01 by Qingyan Guo, Guo, Qingyan, Rui Wang +9 · 3 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Law, AI, and Intellectual Property #Machine Learning (cs.LG)
  71. Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey
    2024/12/16 by Liang Chen, Chen, Liang, Zekun Wang +49 · 5 citations
    Computer Science · #Natural Language Processing Techniques #Topic Modeling
  72. Codec-SUPERB @ SLT 2024: A lightweight benchmark for neural audio codec models
    2024/09/21 by Wu, Haibin, Chen, Xuanjun, Lin, Yi-Cheng +13 · 4 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  73. RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis
    2024/04/04 by Xin, Detai, Tan, Xu, Shen, Kai +8 · 3 citations
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  74. ALMTokenizer: A Low-bitrate and Semantic-rich Audio Codec Tokenizer for Audio Language Modeling
    2025/04/14 by Yang, Dongchao, Liu, Songxiang, Guo, Haohan +9 · 8 citations
    #FOS: Computer and information sciences #Sound (cs.SD)
  75. A Study on ReLU and Softmax in Transformer
    2023/02/13 by Shen, Kai, Guo, Junliang, Tan, Xu +3 · 2 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  76. ResiDual: Transformer with Dual Residual Connections
    2023/04/28 by Xie, Shufang, Zhang, Huishuai, Guo, Junliang +6 · 2 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural and Evolutionary Computing (cs.NE)
  77. ResGrad: Residual Denoising Diffusion Probabilistic Models for Text to Speech
    2022/12/30 by Chen, Zehua, Wu, Yihan, Leng, Yichong +9 · 2 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Signal Processing (eess.SP) #Sound (cs.SD) #electronic engineering #information engineering
  78. GETMusic: Generating Any Music Tracks with a Unified Representation and Diffusion Framework
    2023/05/18 by Ang Lv, Xu Tan, Lv, Ang +11 · 2 citations
    Computer Science · Neuroscience · #Music and Audio Processing #Music Technology and Sound Studies #Neuroscience and Music Perception
  79. Context-Aware Talking-Head Video Editing
    2023/08/01 by Songlin Yang, Yang, Songlin, Wei Wang +9 · 2 citations
    Computer Science · #FOS: Computer and information sciences #Face recognition and analysis #Multimedia (cs.MM) #Speech and Audio Processing #Video Analysis and Summarization
  80. FlashSpeech: Efficient Zero-Shot Speech Synthesis
    2024/04/23 by Zhen Ye, Zeqian Ju, Ye, Zhen +23 · 3 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing
  81. MelodyGLM: Multi-task Pre-training for Symbolic Melody Generation
    2023/09/19 by Wu, Xinda, Huang, Zhijie, Zhang, Kejun +5 · 2 citations
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Information Retrieval (cs.IR) #Multimedia (cs.MM) #Sound (cs.SD) #electronic engineering #information engineering
  82. Beyond Error Propagation in Neural Machine Translation: Characteristics of Language Also Matter
    2018/09/01 by Wu, Lijun, Tan, Xu, He, Di +4 · 1 citation
    #Computation and Language (cs.CL) #FOS: Computer and information sciences
  83. VAST: Vivify Your Talking Avatar via Zero-Shot Expressive Facial Style Transfer
    2023/08/09 by Chen, Liyang, Wu, Zhiyong, Li, Runnan +4 · 2 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  84. Non-Autoregressive Neural Machine Translation with Enhanced Decoder Input
    2018/12/23 by Guo, Junliang, Tan, Xu, He, Di +3 · 1 citation
    #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  85. LightPAFF: A Two-Stage Distillation Framework for Pre-training and Fine-tuning
    2020/04/27 by Kaitao Song, Hao Sun, Song, Kaitao +11 · 1 citation
    Computer Science · #Topic Modeling #Natural Language Processing Techniques #Multimodal Machine Learning Applications
  86. A Survey on Low-Resource Neural Machine Translation
    2021/07/09 by Wang, Rui, Tan, Xu, Luo, Renqian +2 · 1 citation
    #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  87. DenoiSpeech: Denoising Text to Speech with Frame-Level Noise Modeling
    2020/12/17 by Chen Zhang, Yi Ren, Zhang, Chen +13 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  88. DelightfulTTS 2: End-to-End Speech Synthesis with Adversarial Vector-Quantized Auto-Encoders
    2022/07/11 by Yanqing Liu, Ruiqing Xue, Liu, Yanqing +7 · 2 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Natural Language Processing Techniques
  89. FastSAG: Towards Fast Non-Autoregressive Singing Accompaniment Generation
    2024/05/13 by Jianyi Chen, Wei Xue, Chen, Jianyi +9 · 2 citations
    Computer Science · #Speech and Audio Processing #Music and Audio Processing #Speech Recognition and Synthesis
  90. TeleMelody: Lyric-to-Melody Generation with a Template-Based Two-Stage Method
    2021/09/20 by Zeqian Ju, Ju, Zeqian, Peiling Lu +17 · 1 citation
    Computer Science · #Music and Audio Processing #Topic Modeling #Natural Language Processing Techniques
  91. Transformer-S2A: Robust and Efficient Speech-to-Animation
    2021/11/18 by Chen, Liyang, Wu, Zhiyong, Ling, Jun +3 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Graphics (cs.GR) #Multimedia (cs.MM) #Sound (cs.SD) #electronic engineering #information engineering
  92. PDAugment: Data Augmentation by Pitch and Duration Adjustments for Automatic Lyrics Transcription
    2021/09/16 by Zhang, Chen, Yu, Jiaxing, Chang, LuChin +4 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  93. FastCorrect 2: Fast Error Correction on Multiple Candidates for Automatic Speech Recognition
    2021/09/29 by Yichong Leng, Xu Tan, Leng, Yichong +18 · 1 citation
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Music and Audio Processing
  94. MeloForm: Generating Melody with Musical Form based on Expert Systems and Neural Networks
    2022/08/30 by Peiling Lu, Lu, Peiling, Xu Tan +9 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  95. WuYun: Exploring hierarchical skeleton-guided melody generation using knowledge-enhanced deep learning
    2023/01/11 by Kejun Zhang, Zhang, Kejun, Xinda Wu +13 · 1 citation
    Computer Science · Neuroscience · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG) #Music Technology and Sound Studies #Music and Audio Processing #Neuroscience and Music Perception #Sound (cs.SD) #electronic engineering #information engineering
  96. HiFace: High-Fidelity 3D Face Reconstruction by Learning Static and Dynamic Details
    2023/03/20 by Zenghao Chai, Tianke Zhang, Chai, Zenghao +17 · 1 citation
    Computer Science · Engineering · #Face recognition and analysis #Generative Adversarial Networks and Image Synthesis #3D Shape Modeling and Analysis
  97. Mask the Correct Tokens: An Embarrassingly Simple Approach for Error Correction
    2022/11/23 by Shen, Kai, Leng, Yichong, Tan, Xu +4 · 1 citation
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences
  98. SoftCorrect: Error Correction with Soft Detection for Automatic Speech Recognition
    2022/12/02 by Yichong Leng, Leng, Yichong, Xu Tan +15 · 1 citation
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Music and Audio Processing
  99. ERA-Solver: Error-Robust Adams Solver for Fast Sampling of Diffusion Probabilistic Models
    2023/01/30 by Shengmeng Li, Li, Shengming, Luping Liu +5 · 1 citation
    Computer Science · Medicine · #Advanced Neuroimaging Techniques and Applications #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG)
  100. Improving Few-Shot Learning for Talking Face System with TTS Data Augmentation
    2023/03/09 by Chen, Qi, Ma, Ziyang, Liu, Tao +4 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Sound (cs.SD) #electronic engineering #information engineering
  101. DiffusionNER: Boundary Diffusion for Named Entity Recognition
    2023/05/22 by Shen, Yongliang, Song, Kaitao, Tan, Xu +3 · 1 citation
    #Computation and Language (cs.CL) #FOS: Computer and information sciences
  102. Regeneration Learning: A Learning Paradigm for Data Generation
    2023/01/21 by Tan, Xu, Qin, Tao, Bian, Jiang +2 · 1 citation
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #electronic engineering #information engineering
  103. Deliberate then Generate: Enhanced Prompting Framework for Text Generation
    2023/05/31 by Bei Li, Li, Bei, Rui Wang +17 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech and dialogue systems #Topic Modeling
  104. Predictor-Corrector Enhanced Transformers with Exponential Moving Average Coefficient Learning
    2024/11/05 by Li, Bei, Zheng, Tong, Wang, Rui +8 · 2 citations
    #Computation and Language (cs.CL) #FOS: Computer and information sciences
  105. CoMoSVC: Consistency Model-based Singing Voice Conversion
    2024/01/03 by Yiwen Lu, Lu, Yiwen, Zhen Ye +9 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  106. Mixed-Phoneme BERT: Improving BERT with Mixed Phoneme and Sup-Phoneme Representations for Text to Speech
    2022/03/31 by Guangyan Zhang, Kaitao Song, Zhang, Guangyan +19 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  107. Sentence-Level or Token-Level? A Comprehensive Study on Knowledge Distillation
    2024/04/23 by Wei, Jingxuan, Sun, Linzhuang, Leng, Yichong +3 · 1 citation
    #Computation and Language (cs.CL) #FOS: Computer and information sciences
  108. Joint Selective State Space Model and Detrending for Robust Time Series Anomaly Detection
    2024/05/30 by Chen, Junqi, Tan, Xu, Rahardja, Sylwan +2 · 1 citation
    #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  109. FoundationTTS: Text-to-Speech for ASR Customization with Generative Language Model
    2023/03/06 by Xue, Ruiqing, Liu, Yanqing, He, Lei +4 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering