Xu Tan
- HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face
2023/03/30 by Yongliang Shen, Kaitao Song, Shen, Yongliang +10 · 2 voices · 177 citations
Computer Science · Engineering · #Topic Modeling #Ferroelectric and Negative Capacitance Devices #Multimodal Machine Learning Applications
- MPNet: Masked and Permuted Pre-training for Language Understanding
2020/04/20 by Kaitao Song, Xu Tan, Song, Kaitao +7 · 126 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Speech Recognition and Synthesis
- EvoPrompt: Connecting LLMs with Evolutionary Algorithms Yields Powerful Prompt Optimizers
2023/09/15 by Qingyan Guo, Guo, Qingyan, Rui Wang +15 · 85 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Speech and dialogue systems
- NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
2024/03/05 by Zeqian Ju, Ju, Zeqian, Yuancheng Wang +35 · 69 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
2024/06/26 by Şefik Emre Eskimez, Sefik Emre Eskimez, Eskimez, Sefik Emre +24 · 2 voices · 44 citations
Engineering · Neuroscience · #Brain Tumor Detection and Classification #Industrial Vision Systems and Defect Detection #Ultrasonics and Acoustic Wave Propagation #cs.SD #eess.AS
- Beyond Language Models: Byte Models are Digital World Simulators
2024/02/29 by Shangda Wu, Xu Tan, Wu, Shangda +9 · 2 voices · 5 citations
Computer Science · Social Sciences · #FOS: Computer and information sciences #Language and cultural evolution #Machine Learning (cs.LG) #Natural Language Processing Techniques
- Representation Degeneration Problem in Training Natural Language Generation Models
2019/07/27 by Jun Gao, Gao, Jun, Di He +9 · 20 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Multimodal Machine Learning Applications
- A Survey on Neural Speech Synthesis
2021/06/29 by Xu Tan, Tao Qin, Tan, Xu +5 · 22 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimedia (cs.MM) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Achieving Human Parity on Automatic Chinese to English News Translation
2018/03/15 by Hany Hassan, Hassan, Hany, Anthony Aue +45 · 29 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling #Multimodal Machine Learning Applications
- MusicBERT: Symbolic Music Understanding with Large-Scale Pre-Training
2021/06/10 by Mingliang Zeng, Xu Tan, Zeng, Mingliang +9 · 15 citations
Computer Science · Neuroscience · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Information Retrieval (cs.IR) #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Neuroscience and Music Perception #Sound (cs.SD) #electronic engineering #information engineering
- MASS: Masked Sequence to Sequence Pre-training for Language Generation
2019/05/07 by Kaitao Song, Song, Kaitao, Xu Tan +7 · 1 voice · 9 citations
#cs.CL #cs.AI #cs.LG
- NaturalSpeech: End-to-End Text to Speech Synthesis with Human-Level Quality
2022/05/09 by Xu Tan, Tan, Xu, Jiawei Chen +25 · 15 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
- BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec
2024/09/09 by Detai Xin, Xin, Detai, Xu Tan +5 · 27 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- EvoAgent: Towards Automatic Multi-Agent Generation via Evolutionary Algorithms
2024/06/20 by Siyu Yuan, Yuan, Siyu, Kaitao Song +9 · 21 citations
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Multi-Agent Systems and Negotiation #Reinforcement Learning in Robotics #Robotic Path Planning Algorithms
- TaskBench: Benchmarking Large Language Models for Task Automation
2023/11/30 by Yongliang Shen, Shen, Yongliang, Kaitao Song +15 · 17 citations
Computer Science · Medicine · Materials Science · #Topic Modeling #Artificial Intelligence in Healthcare and Education #Machine Learning in Materials Science
- ESPnet-TTS: Unified, Reproducible, and Integratable Open Source End-to-End Text-to-Speech Toolkit
2019/10/24 by Tomoki Hayashi, Ryuichi Yamamoto, Hayashi, Tomoki +15 · 9 citations
Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Music and Audio Processing
- FastCorrect: Fast Error Correction with Edit Alignment for Automatic Speech Recognition
2021/05/09 by Yichong Leng, Leng, Yichong, Xu Tan +17 · 7 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- AdaSpeech: Adaptive Text to Speech for Custom Voice
2021/03/01 by Mingjian Chen, Chen, Mingjian, Xu Tan +11 · 10 citations
Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
- HiFiSinger: Towards High-Fidelity Neural Singing Voice Synthesis
2020/09/03 by Jiawei Chen, Xu Tan, Chen, Jiawei +7 · 11 citations
Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
- MuseCoco: Generating Symbolic Music from Text
2023/05/31 by Peiling Lu, Lu, Peiling, Xin Xu +11 · 9 citations
Computer Science · #Music and Audio Processing #Music Technology and Sound Studies #Speech Recognition and Synthesis
- YuE: Scaling Open Foundation Models for Long-Form Music Generation
2025/03/11 by Ruibin Yuan, Yuan, Ruibin, Shuyue Guo +110 · 23 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computer Graphics and Visualization Techniques #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
- Museformer: Transformer with Fine- and Coarse-Grained Attention for Music Generation
2022/10/19 by Botao Yu, Peiling Lu, Yu, Botao +15 · 7 citations
Computer Science · Neuroscience · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Neuroscience and Music Perception #Sound (cs.SD) #electronic engineering #information engineering
- Almost Unsupervised Text to Speech and Automatic Speech Recognition
2019/05/13 by Yi Ren, Ren, Yi, Xu Tan +9 · 7 citations
Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Speech and Audio Processing
- MuPT: A Generative Symbolic Music Pretrained Transformer
2024/04/09 by Xingwei Qu, Qu, Xingwei, Yuelin Bai +53 · 9 citations
Computer Science · Neuroscience · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Neuroscience and Music Perception #Sound (cs.SD) #electronic engineering #information engineering
- UniAudio 1.5: Large Language Model-driven Audio Codec is A Few-shot Audio Task Learner
2024/06/14 by Dongchao Yang, Haohan Guo, Yang, Dongchao +13 · 10 citations
Computer Science · #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing
- MusicAgent: An AI Agent for Music Understanding and Generation with Large Language Models
2023/10/18 by Dingyao Yu, Kaitao Song, Yu, Dingyao +13 · 7 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Music and Audio Processing #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- PromptTTS 2: Describing and Generating Voices with Text Prompt
2023/09/05 by Yichong Leng, Zhifang Guo, Leng, Yichong +27 · 7 citations
Computer Science · #Speech Recognition and Synthesis #Topic Modeling #Natural Language Processing Techniques
- FRAGE: Frequency-Agnostic Word Representation
2018/09/18 by Chengyue Gong, Di He, Gong, Chengyue +9 · 4 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Speech Recognition and Synthesis
- Schrodinger Bridges Beat Diffusion Models on Text-to-Speech Synthesis
2023/12/06 by Zehua Chen, Guande He, Chen, Zehua +7 · 7 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Foundation Models for Music: A Survey
2024/08/26 by Yinghao Ma, Ma, Yinghao, Anders Øland +81 · 8 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music Technology and Sound Studies #Sound (cs.SD) #electronic engineering #information engineering
- SongMASS: Automatic Song Writing with Pre-training and Alignment Constraint
2020/12/09 by Zhonghao Sheng, Kaitao Song, Sheng, Zhonghao +11 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Topic Modeling #electronic engineering #information engineering
- VidMuse: A Simple Video-to-Music Generation Framework with Long-Short-Term Modeling
2024/06/06 by Zeyue Tian, Zhaoyang Liu, Tian, Zeyue +15 · 6 citations
Computer Science · Social Sciences · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimedia (cs.MM) #Multimedia Communication and Technology #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD)
- D-CPT Law: Domain-specific Continual Pre-Training Scaling Law for Large Language Models
2024/06/03 by Haoran Que, Que, Haoran, Jiaheng Liu +29 · 5 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling
- FastSpeech: Fast, Robust and Controllable Text to Speech
2019/05/22 by Yi Ren, Ren, Yi, Yangjun Ruan +12 · 1 voice · 2 citations
Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #cs.CL #cs.LG #cs.SD #eess.AS #electronic engineering #information engineering
- A Study of Non-autoregressive Model for Sequence Generation
2020/04/22 by Yi Ren, Ren, Yi, Jinglin Liu +9 · 2 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling #Speech Recognition and Synthesis
- HiFace: High-Fidelity 3D Face Reconstruction by Learning Static and Dynamic Details
2023/03/20 by Zenghao Chai, Chai, Zenghao, Tianke Zhang +17 · 3 citations
Computer Science · Engineering · #Face recognition and analysis #Generative Adversarial Networks and Image Synthesis #3D Shape Modeling and Analysis
- UWSpeech: Speech to Speech Translation for Unwritten Languages
2020/06/14 by Chen Zhang, Xu Tan, Zhang, Chen +9 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Adaptive Logit Adjustment Loss for Long-Tailed Visual Recognition
2021/04/13 by Yan Zhao, Zhao, Yan, Weicong Chen +7 · 2 citations
Computer Science · #Advanced Image and Video Retrieval Techniques #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences
- InstructAvatar: Text-Guided Emotion and Motion Control for Avatar Generation
2024/05/24 by Yuchi Wang, Junliang Guo, Wang, Yuchi +13 · 4 citations
Engineering · Psychology · #Human Motion and Animation #Educational Games and Gamification #Social Robot Interaction and HRI
- Revisiting Over-Smoothness in Text to Speech
2022/02/26 by Yi Ren, Ren, Yi, Xu Tan +7 · 2 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
- CLaMP 2: Multimodal Music Information Retrieval Across 101 Languages Using Large Language Models
2024/10/17 by Shangda Wu, Yashan Wang, Wu, Shangda +27 · 4 citations
Computer Science · Arts and Humanities · #Music and Audio Processing #Diverse Musicological Studies #Natural Language Processing Techniques
- InferGrad: Improving Diffusion Models for Vocoder by Considering Inference in Training
2022/02/08 by Zehua Chen, Xu Tan, Chen, Zehua +11 · 2 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Mitigating Reversal Curse in Large Language Models via Semantic-aware Permutation Training
2024/03/01 by Qingyan Guo, Rui Wang, Guo, Qingyan +9 · 3 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Law, AI, and Intellectual Property #Machine Learning (cs.LG)
- Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey
2024/12/16 by Liang Chen, Zekun Wang, Chen, Liang +49 · 5 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling
- WuYun: Exploring hierarchical skeleton-guided melody generation using knowledge-enhanced deep learning
2023/01/11 by Kejun Zhang, Xinda Wu, Zhang, Kejun +13 · 2 citations
Computer Science · Neuroscience · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG) #Music Technology and Sound Studies #Music and Audio Processing #Neuroscience and Music Perception #Sound (cs.SD) #electronic engineering #information engineering
- GETMusic: Generating Any Music Tracks with a Unified Representation and Diffusion Framework
2023/05/18 by Ang Lv, Lv, Ang, Xu Tan +11 · 2 citations
Computer Science · Neuroscience · #Music and Audio Processing #Music Technology and Sound Studies #Neuroscience and Music Perception
- Context-Aware Talking-Head Video Editing
2023/08/01 by Songlin Yang, Wei Wang, Yang, Songlin +9 · 2 citations
Computer Science · #FOS: Computer and information sciences #Face recognition and analysis #Multimedia (cs.MM) #Speech and Audio Processing #Video Analysis and Summarization
- FlashSpeech: Efficient Zero-Shot Speech Synthesis
2024/04/23 by Zhen Ye, Ye, Zhen, Zeqian Ju +23 · 3 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing
- Migration and the Value of Social Networks
2023/12/14 by Joshua Blumenstock, Joshua E Blumenstock, Guanghua Chi +1 · 2 citations
Social Sciences · #Urban, Neighborhood, and Segregation Studies #Migration and Labor Dynamics #Human Mobility and Location-Based Analysis
- LightPAFF: A Two-Stage Distillation Framework for Pre-training and Fine-tuning
2020/04/27 by Kaitao Song, Hao Sun, Song, Kaitao +11 · 1 citation
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Multimodal Machine Learning Applications
- DenoiSpeech: Denoising Text to Speech with Frame-Level Noise Modeling
2020/12/17 by Chen Zhang, Yi Ren, Zhang, Chen +13 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- DelightfulTTS 2: End-to-End Speech Synthesis with Adversarial Vector-Quantized Auto-Encoders
2022/07/11 by Yanqing Liu, Ruiqing Xue, Liu, Yanqing +7 · 2 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Natural Language Processing Techniques
- FastSAG: Towards Fast Non-Autoregressive Singing Accompaniment Generation
2024/05/13 by Jianyi Chen, Wei Xue, Chen, Jianyi +9 · 2 citations
Computer Science · #Speech and Audio Processing #Music and Audio Processing #Speech Recognition and Synthesis
- TeleMelody: Lyric-to-Melody Generation with a Template-Based Two-Stage Method
2021/09/20 by Zeqian Ju, Ju, Zeqian, Peiling Lu +17 · 1 citation
Computer Science · #Music and Audio Processing #Topic Modeling #Natural Language Processing Techniques
- FastCorrect 2: Fast Error Correction on Multiple Candidates for Automatic Speech Recognition
2021/09/29 by Yichong Leng, Leng, Yichong, Xu Tan +18 · 1 citation
Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Music and Audio Processing
- MeloForm: Generating Melody with Musical Form based on Expert Systems and Neural Networks
2022/08/30 by Peiling Lu, Lu, Peiling, Xu Tan +9 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
- SoftCorrect: Error Correction with Soft Detection for Automatic Speech Recognition
2022/12/02 by Yichong Leng, Leng, Yichong, Xu Tan +15 · 1 citation
Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Music and Audio Processing
- ERA-Solver: Error-Robust Adams Solver for Fast Sampling of Diffusion Probabilistic Models
2023/01/30 by Shengmeng Li, Li, Shengming, Luping Liu +5 · 1 citation
Computer Science · Medicine · #Advanced Neuroimaging Techniques and Applications #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG)
- Deliberate then Generate: Enhanced Prompting Framework for Text Generation
2023/05/31 by Bei Li, Rui Wang, Li, Bei +17 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech and dialogue systems #Topic Modeling
- CoMoSVC: Consistency Model-based Singing Voice Conversion
2024/01/03 by Yiwen Lu, Lu, Yiwen, Zhen Ye +9 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Mixed-Phoneme BERT: Improving BERT with Mixed Phoneme and Sup-Phoneme Representations for Text to Speech
2022/03/31 by Guangyan Zhang, Kaitao Song, Zhang, Guangyan +19 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Synergistic Effects of Cepharanthine and Amphotericin B Against Candida albicans by Targeting PMP3-Mediated Sphingolipid Biosynthesis and Membrane Disruption
2026/07/23 by Song Wang, Qiya Zhang, Qinlan Li +4