vix.ing · top · new · best · stats · spec

Jiatong Shi

  1. AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head
    2023/04/25 by Rongjie Huang, Mingze Li, Huang, Rongjie +23 · 1 voice · 47 citations
    Computer Science · #Music and Audio Processing #Speech Recognition and Synthesis #Topic Modeling
  2. SUPERB: Speech processing Universal PERformance Benchmark
    2021/05/03 by Shu-Wen Yang, Po-Han Chi, Yang, Shu-wen +37 · 69 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  3. Dynamic-SUPERB: Towards A Dynamic, Collaborative, and Comprehensive Instruction-Tuning Benchmark for Speech
    2023/09/18 by Chien‐Yu Huang, Huang, Chien-yu, Ke-Han Lu +27 · 20 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  4. Towards Robust Speech Representation Learning for Thousands of Languages
    2024/06/30 by William Chen, Chen, William, Wangyou Zhang +17 · 15 citations
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Speech and dialogue systems
  5. Reproducing Whisper-Style Training Using an Open-Source Toolkit and Publicly Available Data
    2023/09/25 by Yifan Peng, Peng, Yifan, Jinchuan Tian +29 · 9 citations
    Computer Science · #Speech Recognition and Synthesis #Topic Modeling #Natural Language Processing Techniques
  6. The Singing Voice Conversion Challenge 2023
    2023/06/26 by Wen-Chin Huang, Huang, Wen-Chin, Lester Phillip Violeta +7 · 9 citations
    Computer Science · Medicine · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Voice and Speech Disorders #electronic engineering #information engineering
  7. VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music
    2024/12/23 by Jiatong Shi, Shi, Jiatong, Hye-jin Shim +31 · 17 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  8. ESPnet-SPK: full pipeline speaker embedding toolkit with reproducible recipes, self-supervised front-ends, and off-the-shelf models
    2024/01/30 by Jee-weon Jung, Wangyou Zhang, Jung, Jee-weon +13 · 9 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  9. SingMOS: An extensive Open-Source Singing Voice Dataset for MOS Prediction
    2024/06/16 by Yuxun Tang, Jiatong Shi, Tang, Yuxun +5 · 10 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  10. Muskits: an End-to-End Music Processing Toolkit for Singing Voice Synthesis
    2022/05/09 by Jiatong Shi, Shuai Guo, Shi, Jiatong +21 · 6 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
  11. Sequence-to-sequence Singing Voice Synthesis with Perceptual Entropy Loss
    2020/10/22 by Jiatong Shi, Shuai Guo, Shi, Jiatong +7 · 4 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  12. Multi-resolution HuBERT: Multi-resolution Speech Self-Supervised Learning with Masked Unit Prediction
    2023/10/04 by Jiatong Shi, Shi, Jiatong, Hirofumi Inaguma +7 · 4 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  13. Discrete Audio Tokens: More Than a Survey!
    2025/06/12 by Pooneh Mousavi, Gallil Maimon, Mousavi, Pooneh +39 · 13 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Digital Media Forensic Detection #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  14. The Interspeech 2024 Challenge on Speech Processing Using Discrete Units
    2024/06/11 by Xuankai Chang, Jiatong Shi, Chang, Xuankai +17 · 5 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  15. Findings of the IWSLT 2024 Evaluation Campaign
    2024/11/07 by Ibrahim Said Ahmad, Antonios Anastasopoulos, Ahmad, Ibrahim Said +87 · 6 citations
    Computer Science · Engineering · #Advanced Data Processing Techniques #Computation and Language (cs.CL) #FOS: Computer and information sciences #GNSS positioning and interference #Seismology and Earthquake Studies
  16. ESPnet2-TTS: Extending the Edge of TTS Research
    2021/10/15 by Tomoki Hayashi, Hayashi, Tomoki, Ryuichi Yamamoto +17 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
  17. Aligning Text-to-Music Evaluation with Human Preferences
    2025/03/20 by Yichen Huang, Huang, Yichen, Zachary Novack +13 · 7 citations
    Computer Science · #Music and Audio Processing #Music Technology and Sound Studies #Speech and Audio Processing
  18. Self-supervised Speech Representations Still Struggle with African American Vernacular English
    2024/08/26 by Kalvin Chang, Chang, Kalvin, Yi-Hui Chou +11 · 4 citations
    Computer Science · #Speech Recognition and Synthesis
  19. SVDD 2024: The Inaugural Singing Voice Deepfake Detection Challenge
    2024/08/28 by You Zhang, Yongyi Zang, Zhang, You +9 · 3 citations
    Computer Science · #Music and Audio Processing #Speech and Audio Processing #Speech Recognition and Synthesis
  20. Understanding the Tradeoffs in Client-side Privacy for Downstream Speech Tasks
    2021/01/22 by Peter Wu, Paul Pu Liang, Wu, Peter +9 · 1 citation
    Computer Science · #Adversarial Robustness in Machine Learning #Audio and Speech Processing (eess.AS) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #FOS: Electrical engineering #Privacy-Preserving Technologies in Data #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  21. MMM: Multi-Layer Multi-Residual Multi-Stream Discrete Speech Representation from Self-supervised Learning Model
    2024/06/14 by Jiatong Shi, Shi, Jiatong, Xutai Ma +7 · 2 citations
    Computer Science · #Speech and Audio Processing #Speech Recognition and Synthesis
  22. OpusLM: A Family of Open Unified Speech Language Models
    2025/06/21 by Jinchuan Tian, William Chen, Tian, Jinchuan +21 · 6 citations
    Computer Science · #Speech Recognition and Synthesis #Topic Modeling #Natural Language Processing Techniques
  23. Combining Spectral and Self-Supervised Features for Low Resource Speech Recognition and Translation
    2022/04/05 by Dan Berrebbi, Berrebbi, Dan, Jiatong Shi +9 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  24. ESPnet-SDS: Unified Toolkit and Demo for Spoken Dialogue Systems
    2025/03/11 by Siddhant Arora, Arora, Siddhant, Yifan Peng +21 · 1 voice · 3 citations
    Computer Science · #Speech and dialogue systems #Multi-Agent Systems and Negotiation #Natural Language Processing Techniques
  25. Bridging Speech and Textual Pre-trained Models with Unsupervised ASR
    2022/11/06 by Jiatong Shi, Shi, Jiatong, Chan-Jan Hsu +13 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  26. Chain-of-Thought Training for Open E2E Spoken Dialogue Systems
    2025/05/31 by Siddhant Arora, Jinchuan Tian, Arora, Siddhant +13 · 5 citations
    Computer Science · #Speech and dialogue systems #Intelligent Tutoring Systems and Adaptive Learning #Topic Modeling
  27. ESPnet-SpeechLM: An Open Speech Language Model Toolkit
    2025/02/21 by Jinchuan Tian, Jiatong Shi, Tian, Jinchuan +29 · 5 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech and dialogue systems #electronic engineering #information engineering
  28. Exploration on HuBERT with Multiple Resolutions
    2023/06/01 by Jiatong Shi, Shi, Jiatong, Yun Tang +9 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
  29. EFFUSE: Efficient Self-Supervised Feature Fusion for E2E ASR in Low Resource and Multilingual Scenarios
    2023/10/05 by Tejes Srivastava, Srivastava, Tejes, Jiatong Shi +5 · 1 citation
    Computer Science · Engineering · #Advanced Chemical Sensor Technologies #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Fault Detection and Control Systems #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  30. Joint Beam Search Integrating CTC, Attention, and Transducer Decoders
    2024/06/05 by Yui Sudo, Sudo, Yui, Muhammad Shakeel +11 · 2 citations
    Engineering · Physics and Astronomy · #Advanced Radiotherapy Techniques #Advancements in Photolithography Techniques #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  31. SUPERB @ SLT 2022: Challenge on Generalization and Efficiency of Self-Supervised Speech Representation Learning
    2022/10/16 by Tzu-hsun Feng, Feng, Tzu-hsun, Annie Dong +25 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  32. SVDD Challenge 2024: A Singing Voice Deepfake Detection Challenge Evaluation Plan
    2024/05/08 by You Zhang, Yongyi Zang, Zhang, You +13 · 1 citation
    Computer Science · #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing
  33. FoodPuzzle: Developing Large Language Model Agents as Flavor Scientists
    2024/09/19 by Tenghao Huang, Dong‐Hee Lee, Huang, Tenghao +12 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Topic Modeling
  34. Findings of the 2023 ML-SUPERB Challenge: Pre-Training and Evaluation over More Languages and Beyond
    2023/10/09 by Jiatong Shi, Shi, Jiatong, William Chen +23 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering