Qin, Yong
- Better Zero-Shot Reasoning with Role-Play Prompting
2023/08/15 by Aobo Kong, Shiwan Zhao, Kong, Aobo +15 · 1 voice · 88 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Logic, Reasoning, and Knowledge #Multi-Agent Systems and Negotiation #Topic Modeling #cs.CL
- LAFMA: A Latent Flow Matching Model for Text-to-Audio Generation
2024/06/12 by Wenhao Guan, Guan, Wenhao, Kaidi Wang +15 · 13 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- SDPO: Segment-Level Direct Preference Optimization for Social Agents
2025/01/03 by Aobo Kong, Kong, Aobo, Wentao Ma +17 · 14 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Data Management and Algorithms #FOS: Computer and information sciences #Multi-Agent Systems and Negotiation #Reinforcement Learning in Robotics
- MusicEval: A Generative Music Dataset with Expert Ratings for Automatic Text-to-Music Evaluation
2025/01/18 by Cheng Liu, Hui Wang, Liu, Cheng +15 · 17 citations
Arts and Humanities · Computer Science · #Audio and Speech Processing (eess.AS) #Diverse Musicological Studies #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
- PromptRank: Unsupervised Keyphrase Extraction Using Prompt
2023/05/08 by Aobo Kong, Shiwan Zhao, Kong, Aobo +10 · 5 citations
Computer Science · #Advanced Text Analysis Techniques #FOS: Computer and information sciences #Information Retrieval (cs.IR)
- AudioEditor: A Training-Free Diffusion-Based Audio Editing Framework
2024/09/19 by Jia, Yuhang, Chen, Yang, Zhao, Jinghua +4 · 8 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Self-Prompt Tuning: Enable Autonomous Role-Playing in LLMs
2024/07/12 by Aobo Kong, Shiwan Zhao, Kong, Aobo +15 · 7 citations
Business, Management and Accounting · Computer Science · #Business Process Modeling and Analysis #Computation and Language (cs.CL) #FOS: Computer and information sciences #Model-Driven Software Engineering Techniques #Multi-Agent Systems and Negotiation
- Fine-grained Disentangled Representation Learning for Multimodal Emotion Recognition
2023/12/21 by Sun, Haoqin, Zhao, Shiwan, Wang, Xuechen +3 · 5 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Sound (cs.SD) #electronic engineering #information engineering
- FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow Matching
2025/02/16 by Wang, Hui, Liu, Shujie, Meng, Lingwei +9 · 10 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- kNN-CTC: Enhancing ASR via Retrieval of CTC Pseudo Labels
2023/12/21 by Jiaming Zhou, Shiwan Zhao, Zhou, Jiaming +9 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- CS-Dialogue: A 104-Hour Dataset of Spontaneous Mandarin-English Code-Switching Dialogues for Speech Recognition
2025/02/26 by Jiaming Zhou, Zhou, Jiaming, Yujie Guo +21 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
- Enhancing Emotion Recognition in Incomplete Data: A Novel Cross-Modal Alignment, Reconstruction, and Refinement Framework
2024/07/12 by Haoqin Sun, Shiwan Zhao, Sun, Haoqin +17 · 4 citations
Computer Science · Psychology · #Anomaly Detection Techniques and Applications #Emotion and Mood Recognition
- ChildMandarin: A Comprehensive Mandarin Speech Dataset for Young Children Aged 3-5
2024/09/27 by Jiaming Zhou, Zhou, Jiaming, Shiyao Wang +19 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Uncertainty-Aware Mean Opinion Score Prediction
2024/08/23 by Hui Wang, Wang, Hui, Shiwan Zhao +11 · 4 citations
Computer Science · Physics and Astronomy · #Sentiment Analysis and Opinion Mining #Opinion Dynamics and Social Influence #Advanced Text Analysis Techniques
- M2R-Whisper: Multi-stage and Multi-scale Retrieval Augmentation for Enhancing Whisper
2024/09/18 by Jiaming Zhou, Zhou, Jiaming, Shiwan Zhao +12 · 3 citations
Computer Science · #Sentiment Analysis and Opinion Mining
- Improving Zero-Shot Chinese-English Code-Switching ASR with kNN-CTC and Gated Monolingual Datastores
2024/06/06 by Zhou, Jiaming, Zhao, Shiwan, Wang, Hui +4 · 2 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- AISHELL-Stammertalk 中文口吃数据库 A Mandarin stuttered speech dataset
2024/06/11 by Rong Gong, Gong, Rong, Hongfei Xue +25 · 2 citations
Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Stuttering Research and Treatment #electronic engineering #information engineering
- EmotionTalk: An Interactive Chinese Multimodal Emotion Dataset With Rich Annotations
2025/05/29 by Sun, Haoqin, Wang, Xuechen, Zhao, Jinghua +9 · 3 citations
#FOS: Computer and information sciences #Multimedia (cs.MM)
- SeniorTalk: A Chinese Conversation Dataset with Rich Annotations for Super-Aged Seniors
2025/03/20 by Yang Chen, Chen, Yang, Hui Wang +16 · 3 citations
Computer Science · Medicine · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Face recognition and analysis #Sound (cs.SD) #Speech Recognition and Synthesis #Voice and Speech Disorders #electronic engineering #information engineering
- Enhancing Multimodal Emotion Recognition through Multi-Granularity Cross-Modal Alignment
2024/12/30 by Xuechen Wang, Shiwan Zhao, Wang, Xuechen +9 · 3 citations
Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Emotion and Mood Recognition #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- DIFFA: Large Language Diffusion Models Can Listen and Understand
2025/07/24 by Jiaming Zhou, Zhou, Jiaming, H. Y. Chen +17 · 4 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- UB-Mesh: a Hierarchically Localized nD-FullMesh Datacenter Network Architecture
2025/03/26 by Heng Liao, Bingyang Liu, Liao, Heng +60 · 2 citations
Computer Science · #Graph Theory and Algorithms #Cloud Computing and Resource Management #Distributed and Parallel Computing Systems
- A Self-Training Approach for Whisper to Enhance Long Dysarthric Speech Recognition
2025/06/28 by Shiyao Wang, Wang, Shiyao, Jiaming Zhou +5 · 2 citations
Computer Science · Medicine · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonocardiography and Auscultation Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Voice and Speech Disorders #electronic engineering #information engineering
- Towards Automatic Evaluation and High-Quality Pseudo-Parallel Dataset Construction for Audio Editing: A Human-in-the-Loop Method
2025/08/16 by Yuhang Jia, Jia, Yuhang, Wang, Hui +8 · 2 citations
Computer Science · #FOS: Computer and information sciences #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing
- StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling
2025/06/14 by Hui Wang, Wang, Hui, Shujie Liu +16 · 4 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- AudioEval: Automatic Dual-Perspective and Multi-Dimensional Evaluation of Text-to-Audio-Generation
2025/10/16 by Wang, Hui, Zhao, Jinghua, Liu, Cheng +4 · 1 citation
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Findings of the 2024 Mandarin Stuttering Event Detection and Automatic Speech Recognition Challenge
2024/09/09 by Hongfei Xue, Xue, Hongfei, Gong, Rong +20 · 1 citation
Computer Science · #Speech Recognition and Synthesis
- WildElder: A Chinese Elderly Speech Dataset from the Wild with Fine-Grained Manual Annotations
2025/10/10 by Hui Wang, Wang, Hui, Jiaming Zhou +7 · 1 citation
Computer Science · Engineering · #cs.SD #eess.AS
- RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval
2025/05/26 by Sun, Haoqin, Jingguang Tian, Tian, Jingguang +18 · 2 citations
Computer Science · #Music and Audio Processing #Speech and Audio Processing #Speech Recognition and Synthesis
- MADI: Inter-domain Matching and Intra-domain Discrimination for Cross-domain Speech Recognition
2023/02/22 by Zhou, Jiaming, Zhao, Shiwan, Jiang, Ning +2 · 1 citation
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Iterative Prototype Refinement for Ambiguous Speech Emotion Recognition
2024/08/01 by Haoqin Sun, Sun, Haoqin, Shiwan Zhao +11 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering