vix.ing · top · new · best · stats · spec

Qin, Yong

  1. Better Zero-Shot Reasoning with Role-Play Prompting
    2023/08/15 by Aobo Kong, Shiwan Zhao, Kong, Aobo +15 · 1 voice · 88 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Logic, Reasoning, and Knowledge #Multi-Agent Systems and Negotiation #Topic Modeling #cs.CL
  2. LAFMA: A Latent Flow Matching Model for Text-to-Audio Generation
    2024/06/12 by Wenhao Guan, Guan, Wenhao, Kaidi Wang +15 · 13 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  3. SDPO: Segment-Level Direct Preference Optimization for Social Agents
    2025/01/03 by Aobo Kong, Kong, Aobo, Wentao Ma +17 · 14 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Data Management and Algorithms #FOS: Computer and information sciences #Multi-Agent Systems and Negotiation #Reinforcement Learning in Robotics
  4. MusicEval: A Generative Music Dataset with Expert Ratings for Automatic Text-to-Music Evaluation
    2025/01/18 by Cheng Liu, Hui Wang, Liu, Cheng +15 · 17 citations
    Arts and Humanities · Computer Science · #Audio and Speech Processing (eess.AS) #Diverse Musicological Studies #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  5. PromptRank: Unsupervised Keyphrase Extraction Using Prompt
    2023/05/08 by Aobo Kong, Shiwan Zhao, Kong, Aobo +10 · 5 citations
    Computer Science · #Advanced Text Analysis Techniques #FOS: Computer and information sciences #Information Retrieval (cs.IR)
  6. AudioEditor: A Training-Free Diffusion-Based Audio Editing Framework
    2024/09/19 by Jia, Yuhang, Chen, Yang, Zhao, Jinghua +4 · 8 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  7. Self-Prompt Tuning: Enable Autonomous Role-Playing in LLMs
    2024/07/12 by Aobo Kong, Shiwan Zhao, Kong, Aobo +15 · 7 citations
    Business, Management and Accounting · Computer Science · #Business Process Modeling and Analysis #Computation and Language (cs.CL) #FOS: Computer and information sciences #Model-Driven Software Engineering Techniques #Multi-Agent Systems and Negotiation
  8. Fine-grained Disentangled Representation Learning for Multimodal Emotion Recognition
    2023/12/21 by Sun, Haoqin, Zhao, Shiwan, Wang, Xuechen +3 · 5 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Sound (cs.SD) #electronic engineering #information engineering
  9. FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow Matching
    2025/02/16 by Wang, Hui, Liu, Shujie, Meng, Lingwei +9 · 10 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  10. kNN-CTC: Enhancing ASR via Retrieval of CTC Pseudo Labels
    2023/12/21 by Jiaming Zhou, Shiwan Zhao, Zhou, Jiaming +9 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  11. CS-Dialogue: A 104-Hour Dataset of Spontaneous Mandarin-English Code-Switching Dialogues for Speech Recognition
    2025/02/26 by Jiaming Zhou, Zhou, Jiaming, Yujie Guo +21 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
  12. Enhancing Emotion Recognition in Incomplete Data: A Novel Cross-Modal Alignment, Reconstruction, and Refinement Framework
    2024/07/12 by Haoqin Sun, Shiwan Zhao, Sun, Haoqin +17 · 4 citations
    Computer Science · Psychology · #Anomaly Detection Techniques and Applications #Emotion and Mood Recognition
  13. ChildMandarin: A Comprehensive Mandarin Speech Dataset for Young Children Aged 3-5
    2024/09/27 by Jiaming Zhou, Zhou, Jiaming, Shiyao Wang +19 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  14. Uncertainty-Aware Mean Opinion Score Prediction
    2024/08/23 by Hui Wang, Wang, Hui, Shiwan Zhao +11 · 4 citations
    Computer Science · Physics and Astronomy · #Sentiment Analysis and Opinion Mining #Opinion Dynamics and Social Influence #Advanced Text Analysis Techniques
  15. M2R-Whisper: Multi-stage and Multi-scale Retrieval Augmentation for Enhancing Whisper
    2024/09/18 by Jiaming Zhou, Zhou, Jiaming, Shiwan Zhao +12 · 3 citations
    Computer Science · #Sentiment Analysis and Opinion Mining
  16. Improving Zero-Shot Chinese-English Code-Switching ASR with kNN-CTC and Gated Monolingual Datastores
    2024/06/06 by Zhou, Jiaming, Zhao, Shiwan, Wang, Hui +4 · 2 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  17. AISHELL-Stammertalk 中文口吃数据库 A Mandarin stuttered speech dataset
    2024/06/11 by Rong Gong, Gong, Rong, Hongfei Xue +25 · 2 citations
    Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Stuttering Research and Treatment #electronic engineering #information engineering
  18. EmotionTalk: An Interactive Chinese Multimodal Emotion Dataset With Rich Annotations
    2025/05/29 by Sun, Haoqin, Wang, Xuechen, Zhao, Jinghua +9 · 3 citations
    #FOS: Computer and information sciences #Multimedia (cs.MM)
  19. SeniorTalk: A Chinese Conversation Dataset with Rich Annotations for Super-Aged Seniors
    2025/03/20 by Yang Chen, Chen, Yang, Hui Wang +16 · 3 citations
    Computer Science · Medicine · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Face recognition and analysis #Sound (cs.SD) #Speech Recognition and Synthesis #Voice and Speech Disorders #electronic engineering #information engineering
  20. Enhancing Multimodal Emotion Recognition through Multi-Granularity Cross-Modal Alignment
    2024/12/30 by Xuechen Wang, Shiwan Zhao, Wang, Xuechen +9 · 3 citations
    Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Emotion and Mood Recognition #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  21. DIFFA: Large Language Diffusion Models Can Listen and Understand
    2025/07/24 by Jiaming Zhou, Zhou, Jiaming, H. Y. Chen +17 · 4 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  22. UB-Mesh: a Hierarchically Localized nD-FullMesh Datacenter Network Architecture
    2025/03/26 by Heng Liao, Bingyang Liu, Liao, Heng +60 · 2 citations
    Computer Science · #Graph Theory and Algorithms #Cloud Computing and Resource Management #Distributed and Parallel Computing Systems
  23. A Self-Training Approach for Whisper to Enhance Long Dysarthric Speech Recognition
    2025/06/28 by Shiyao Wang, Wang, Shiyao, Jiaming Zhou +5 · 2 citations
    Computer Science · Medicine · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonocardiography and Auscultation Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Voice and Speech Disorders #electronic engineering #information engineering
  24. Towards Automatic Evaluation and High-Quality Pseudo-Parallel Dataset Construction for Audio Editing: A Human-in-the-Loop Method
    2025/08/16 by Yuhang Jia, Jia, Yuhang, Wang, Hui +8 · 2 citations
    Computer Science · #FOS: Computer and information sciences #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing
  25. StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling
    2025/06/14 by Hui Wang, Wang, Hui, Shujie Liu +16 · 4 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  26. AudioEval: Automatic Dual-Perspective and Multi-Dimensional Evaluation of Text-to-Audio-Generation
    2025/10/16 by Wang, Hui, Zhao, Jinghua, Liu, Cheng +4 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  27. Findings of the 2024 Mandarin Stuttering Event Detection and Automatic Speech Recognition Challenge
    2024/09/09 by Hongfei Xue, Xue, Hongfei, Gong, Rong +20 · 1 citation
    Computer Science · #Speech Recognition and Synthesis
  28. WildElder: A Chinese Elderly Speech Dataset from the Wild with Fine-Grained Manual Annotations
    2025/10/10 by Hui Wang, Wang, Hui, Jiaming Zhou +7 · 1 citation
    Computer Science · Engineering · #cs.SD #eess.AS
  29. RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval
    2025/05/26 by Sun, Haoqin, Jingguang Tian, Tian, Jingguang +18 · 2 citations
    Computer Science · #Music and Audio Processing #Speech and Audio Processing #Speech Recognition and Synthesis
  30. MADI: Inter-domain Matching and Intra-domain Discrimination for Cross-domain Speech Recognition
    2023/02/22 by Zhou, Jiaming, Zhao, Shiwan, Jiang, Ning +2 · 1 citation
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  31. Iterative Prototype Refinement for Ambiguous Speech Emotion Recognition
    2024/08/01 by Haoqin Sun, Sun, Haoqin, Shiwan Zhao +11 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering