vix.ing · top · new · best · stats · spec

Xu Tan

  1. HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face
    2023/03/30 by Yongliang Shen, Kaitao Song, Shen, Yongliang +10 · 2 voices · 177 citations
    Computer Science · Engineering · #Topic Modeling #Ferroelectric and Negative Capacitance Devices #Multimodal Machine Learning Applications
  2. MPNet: Masked and Permuted Pre-training for Language Understanding
    2020/04/20 by Kaitao Song, Xu Tan, Song, Kaitao +7 · 126 citations
    Computer Science · #Topic Modeling #Natural Language Processing Techniques #Speech Recognition and Synthesis
  3. EvoPrompt: Connecting LLMs with Evolutionary Algorithms Yields Powerful Prompt Optimizers
    2023/09/15 by Qingyan Guo, Guo, Qingyan, Rui Wang +15 · 85 citations
    Computer Science · #Topic Modeling #Natural Language Processing Techniques #Speech and dialogue systems
  4. NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
    2024/03/05 by Zeqian Ju, Ju, Zeqian, Yuancheng Wang +35 · 69 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  5. E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
    2024/06/26 by Şefik Emre Eskimez, Sefik Emre Eskimez, Eskimez, Sefik Emre +24 · 2 voices · 44 citations
    Engineering · Neuroscience · #Brain Tumor Detection and Classification #Industrial Vision Systems and Defect Detection #Ultrasonics and Acoustic Wave Propagation #cs.SD #eess.AS
  6. Beyond Language Models: Byte Models are Digital World Simulators
    2024/02/29 by Shangda Wu, Xu Tan, Wu, Shangda +9 · 2 voices · 5 citations
    Computer Science · Social Sciences · #FOS: Computer and information sciences #Language and cultural evolution #Machine Learning (cs.LG) #Natural Language Processing Techniques
  7. Representation Degeneration Problem in Training Natural Language Generation Models
    2019/07/27 by Jun Gao, Gao, Jun, Di He +9 · 20 citations
    Computer Science · #Topic Modeling #Natural Language Processing Techniques #Multimodal Machine Learning Applications
  8. A Survey on Neural Speech Synthesis
    2021/06/29 by Xu Tan, Tao Qin, Tan, Xu +5 · 22 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimedia (cs.MM) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  9. Achieving Human Parity on Automatic Chinese to English News Translation
    2018/03/15 by Hany Hassan, Hassan, Hany, Anthony Aue +45 · 29 citations
    Computer Science · #Natural Language Processing Techniques #Topic Modeling #Multimodal Machine Learning Applications
  10. MusicBERT: Symbolic Music Understanding with Large-Scale Pre-Training
    2021/06/10 by Mingliang Zeng, Xu Tan, Zeng, Mingliang +9 · 15 citations
    Computer Science · Neuroscience · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Information Retrieval (cs.IR) #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Neuroscience and Music Perception #Sound (cs.SD) #electronic engineering #information engineering
  11. MASS: Masked Sequence to Sequence Pre-training for Language Generation
    2019/05/07 by Kaitao Song, Song, Kaitao, Xu Tan +7 · 1 voice · 9 citations
    #cs.CL #cs.AI #cs.LG
  12. NaturalSpeech: End-to-End Text to Speech Synthesis with Human-Level Quality
    2022/05/09 by Xu Tan, Tan, Xu, Jiawei Chen +25 · 15 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
  13. BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec
    2024/09/09 by Detai Xin, Xin, Detai, Xu Tan +5 · 27 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  14. EvoAgent: Towards Automatic Multi-Agent Generation via Evolutionary Algorithms
    2024/06/20 by Siyu Yuan, Yuan, Siyu, Kaitao Song +9 · 21 citations
    Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Multi-Agent Systems and Negotiation #Reinforcement Learning in Robotics #Robotic Path Planning Algorithms
  15. TaskBench: Benchmarking Large Language Models for Task Automation
    2023/11/30 by Yongliang Shen, Shen, Yongliang, Kaitao Song +15 · 17 citations
    Computer Science · Medicine · Materials Science · #Topic Modeling #Artificial Intelligence in Healthcare and Education #Machine Learning in Materials Science
  16. ESPnet-TTS: Unified, Reproducible, and Integratable Open Source End-to-End Text-to-Speech Toolkit
    2019/10/24 by Tomoki Hayashi, Ryuichi Yamamoto, Hayashi, Tomoki +15 · 9 citations
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Music and Audio Processing
  17. FastCorrect: Fast Error Correction with Edit Alignment for Automatic Speech Recognition
    2021/05/09 by Yichong Leng, Leng, Yichong, Xu Tan +17 · 7 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  18. AdaSpeech: Adaptive Text to Speech for Custom Voice
    2021/03/01 by Mingjian Chen, Chen, Mingjian, Xu Tan +11 · 10 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
  19. HiFiSinger: Towards High-Fidelity Neural Singing Voice Synthesis
    2020/09/03 by Jiawei Chen, Xu Tan, Chen, Jiawei +7 · 11 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
  20. MuseCoco: Generating Symbolic Music from Text
    2023/05/31 by Peiling Lu, Lu, Peiling, Xin Xu +11 · 9 citations
    Computer Science · #Music and Audio Processing #Music Technology and Sound Studies #Speech Recognition and Synthesis
  21. YuE: Scaling Open Foundation Models for Long-Form Music Generation
    2025/03/11 by Ruibin Yuan, Yuan, Ruibin, Shuyue Guo +110 · 23 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computer Graphics and Visualization Techniques #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  22. Museformer: Transformer with Fine- and Coarse-Grained Attention for Music Generation
    2022/10/19 by Botao Yu, Peiling Lu, Yu, Botao +15 · 7 citations
    Computer Science · Neuroscience · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Neuroscience and Music Perception #Sound (cs.SD) #electronic engineering #information engineering
  23. Almost Unsupervised Text to Speech and Automatic Speech Recognition
    2019/05/13 by Yi Ren, Ren, Yi, Xu Tan +9 · 7 citations
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Speech and Audio Processing
  24. MuPT: A Generative Symbolic Music Pretrained Transformer
    2024/04/09 by Xingwei Qu, Qu, Xingwei, Yuelin Bai +53 · 9 citations
    Computer Science · Neuroscience · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Neuroscience and Music Perception #Sound (cs.SD) #electronic engineering #information engineering
  25. UniAudio 1.5: Large Language Model-driven Audio Codec is A Few-shot Audio Task Learner
    2024/06/14 by Dongchao Yang, Haohan Guo, Yang, Dongchao +13 · 10 citations
    Computer Science · #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing
  26. MusicAgent: An AI Agent for Music Understanding and Generation with Large Language Models
    2023/10/18 by Dingyao Yu, Kaitao Song, Yu, Dingyao +13 · 7 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Music and Audio Processing #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  27. PromptTTS 2: Describing and Generating Voices with Text Prompt
    2023/09/05 by Yichong Leng, Zhifang Guo, Leng, Yichong +27 · 7 citations
    Computer Science · #Speech Recognition and Synthesis #Topic Modeling #Natural Language Processing Techniques
  28. FRAGE: Frequency-Agnostic Word Representation
    2018/09/18 by Chengyue Gong, Di He, Gong, Chengyue +9 · 4 citations
    Computer Science · #Topic Modeling #Natural Language Processing Techniques #Speech Recognition and Synthesis
  29. Schrodinger Bridges Beat Diffusion Models on Text-to-Speech Synthesis
    2023/12/06 by Zehua Chen, Guande He, Chen, Zehua +7 · 7 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  30. Foundation Models for Music: A Survey
    2024/08/26 by Yinghao Ma, Ma, Yinghao, Anders Øland +81 · 8 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music Technology and Sound Studies #Sound (cs.SD) #electronic engineering #information engineering
  31. SongMASS: Automatic Song Writing with Pre-training and Alignment Constraint
    2020/12/09 by Zhonghao Sheng, Kaitao Song, Sheng, Zhonghao +11 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Topic Modeling #electronic engineering #information engineering
  32. VidMuse: A Simple Video-to-Music Generation Framework with Long-Short-Term Modeling
    2024/06/06 by Zeyue Tian, Zhaoyang Liu, Tian, Zeyue +15 · 6 citations
    Computer Science · Social Sciences · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimedia (cs.MM) #Multimedia Communication and Technology #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD)
  33. D-CPT Law: Domain-specific Continual Pre-Training Scaling Law for Large Language Models
    2024/06/03 by Haoran Que, Que, Haoran, Jiaheng Liu +29 · 5 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling
  34. FastSpeech: Fast, Robust and Controllable Text to Speech
    2019/05/22 by Yi Ren, Ren, Yi, Yangjun Ruan +12 · 1 voice · 2 citations
    Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #cs.CL #cs.LG #cs.SD #eess.AS #electronic engineering #information engineering
  35. A Study of Non-autoregressive Model for Sequence Generation
    2020/04/22 by Yi Ren, Ren, Yi, Jinglin Liu +9 · 2 citations
    Computer Science · #Natural Language Processing Techniques #Topic Modeling #Speech Recognition and Synthesis
  36. HiFace: High-Fidelity 3D Face Reconstruction by Learning Static and Dynamic Details
    2023/03/20 by Zenghao Chai, Chai, Zenghao, Tianke Zhang +17 · 3 citations
    Computer Science · Engineering · #Face recognition and analysis #Generative Adversarial Networks and Image Synthesis #3D Shape Modeling and Analysis
  37. UWSpeech: Speech to Speech Translation for Unwritten Languages
    2020/06/14 by Chen Zhang, Xu Tan, Zhang, Chen +9 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  38. Adaptive Logit Adjustment Loss for Long-Tailed Visual Recognition
    2021/04/13 by Yan Zhao, Zhao, Yan, Weicong Chen +7 · 2 citations
    Computer Science · #Advanced Image and Video Retrieval Techniques #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences
  39. InstructAvatar: Text-Guided Emotion and Motion Control for Avatar Generation
    2024/05/24 by Yuchi Wang, Junliang Guo, Wang, Yuchi +13 · 4 citations
    Engineering · Psychology · #Human Motion and Animation #Educational Games and Gamification #Social Robot Interaction and HRI
  40. Revisiting Over-Smoothness in Text to Speech
    2022/02/26 by Yi Ren, Ren, Yi, Xu Tan +7 · 2 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
  41. CLaMP 2: Multimodal Music Information Retrieval Across 101 Languages Using Large Language Models
    2024/10/17 by Shangda Wu, Yashan Wang, Wu, Shangda +27 · 4 citations
    Computer Science · Arts and Humanities · #Music and Audio Processing #Diverse Musicological Studies #Natural Language Processing Techniques
  42. InferGrad: Improving Diffusion Models for Vocoder by Considering Inference in Training
    2022/02/08 by Zehua Chen, Xu Tan, Chen, Zehua +11 · 2 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  43. Mitigating Reversal Curse in Large Language Models via Semantic-aware Permutation Training
    2024/03/01 by Qingyan Guo, Rui Wang, Guo, Qingyan +9 · 3 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Law, AI, and Intellectual Property #Machine Learning (cs.LG)
  44. Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey
    2024/12/16 by Liang Chen, Zekun Wang, Chen, Liang +49 · 5 citations
    Computer Science · #Natural Language Processing Techniques #Topic Modeling
  45. WuYun: Exploring hierarchical skeleton-guided melody generation using knowledge-enhanced deep learning
    2023/01/11 by Kejun Zhang, Xinda Wu, Zhang, Kejun +13 · 2 citations
    Computer Science · Neuroscience · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG) #Music Technology and Sound Studies #Music and Audio Processing #Neuroscience and Music Perception #Sound (cs.SD) #electronic engineering #information engineering
  46. GETMusic: Generating Any Music Tracks with a Unified Representation and Diffusion Framework
    2023/05/18 by Ang Lv, Lv, Ang, Xu Tan +11 · 2 citations
    Computer Science · Neuroscience · #Music and Audio Processing #Music Technology and Sound Studies #Neuroscience and Music Perception
  47. Context-Aware Talking-Head Video Editing
    2023/08/01 by Songlin Yang, Wei Wang, Yang, Songlin +9 · 2 citations
    Computer Science · #FOS: Computer and information sciences #Face recognition and analysis #Multimedia (cs.MM) #Speech and Audio Processing #Video Analysis and Summarization
  48. FlashSpeech: Efficient Zero-Shot Speech Synthesis
    2024/04/23 by Zhen Ye, Ye, Zhen, Zeqian Ju +23 · 3 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing
  49. Migration and the Value of Social Networks
    2023/12/14 by Joshua Blumenstock, Joshua E Blumenstock, Guanghua Chi +1 · 2 citations
    Social Sciences · #Urban, Neighborhood, and Segregation Studies #Migration and Labor Dynamics #Human Mobility and Location-Based Analysis
  50. LightPAFF: A Two-Stage Distillation Framework for Pre-training and Fine-tuning
    2020/04/27 by Kaitao Song, Hao Sun, Song, Kaitao +11 · 1 citation
    Computer Science · #Topic Modeling #Natural Language Processing Techniques #Multimodal Machine Learning Applications
  51. DenoiSpeech: Denoising Text to Speech with Frame-Level Noise Modeling
    2020/12/17 by Chen Zhang, Yi Ren, Zhang, Chen +13 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  52. DelightfulTTS 2: End-to-End Speech Synthesis with Adversarial Vector-Quantized Auto-Encoders
    2022/07/11 by Yanqing Liu, Ruiqing Xue, Liu, Yanqing +7 · 2 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Natural Language Processing Techniques
  53. FastSAG: Towards Fast Non-Autoregressive Singing Accompaniment Generation
    2024/05/13 by Jianyi Chen, Wei Xue, Chen, Jianyi +9 · 2 citations
    Computer Science · #Speech and Audio Processing #Music and Audio Processing #Speech Recognition and Synthesis
  54. TeleMelody: Lyric-to-Melody Generation with a Template-Based Two-Stage Method
    2021/09/20 by Zeqian Ju, Ju, Zeqian, Peiling Lu +17 · 1 citation
    Computer Science · #Music and Audio Processing #Topic Modeling #Natural Language Processing Techniques
  55. FastCorrect 2: Fast Error Correction on Multiple Candidates for Automatic Speech Recognition
    2021/09/29 by Yichong Leng, Leng, Yichong, Xu Tan +18 · 1 citation
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Music and Audio Processing
  56. MeloForm: Generating Melody with Musical Form based on Expert Systems and Neural Networks
    2022/08/30 by Peiling Lu, Lu, Peiling, Xu Tan +9 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  57. SoftCorrect: Error Correction with Soft Detection for Automatic Speech Recognition
    2022/12/02 by Yichong Leng, Leng, Yichong, Xu Tan +15 · 1 citation
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Music and Audio Processing
  58. ERA-Solver: Error-Robust Adams Solver for Fast Sampling of Diffusion Probabilistic Models
    2023/01/30 by Shengmeng Li, Li, Shengming, Luping Liu +5 · 1 citation
    Computer Science · Medicine · #Advanced Neuroimaging Techniques and Applications #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG)
  59. Deliberate then Generate: Enhanced Prompting Framework for Text Generation
    2023/05/31 by Bei Li, Rui Wang, Li, Bei +17 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech and dialogue systems #Topic Modeling
  60. CoMoSVC: Consistency Model-based Singing Voice Conversion
    2024/01/03 by Yiwen Lu, Lu, Yiwen, Zhen Ye +9 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  61. Mixed-Phoneme BERT: Improving BERT with Mixed Phoneme and Sup-Phoneme Representations for Text to Speech
    2022/03/31 by Guangyan Zhang, Kaitao Song, Zhang, Guangyan +19 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  62. Synergistic Effects of Cepharanthine and Amphotericin B Against Candida albicans by Targeting PMP3-Mediated Sphingolipid Biosynthesis and Membrane Disruption
    2026/07/23 by Song Wang, Qiya Zhang, Qinlan Li +4