Kai Yu
- F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
2024/10/09 by Yushen Chen, Zhikang Niu, Chen, Yushen +14 · 2 voices · 125 citations
Computer Science · #Music and Audio Processing #cs.SD #eess.AS
- Diverse Data Augmentation with Diffusions for Effective Test-time Prompt Tuning
2023/08/11 by Chun-Mei Feng, Kai Yu, Feng, Chun-Mei +7 · 30 citations
Biochemistry, Genetics and Molecular Biology · Computer Science · #Cancer-related molecular mechanisms research #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications
- RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
2025/04/24 by Zihan Wang, Kangrui Wang, Wang, Zihan +33 · 76 citations
Computer Science · #Reinforcement Learning in Robotics #Multimodal Machine Learning Applications #Domain Adaptation and Few-Shot Learning
- HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer
2025/05/28 by Cai Qi, Jingwen Chen, Cai, Qi +40 · 61 citations
Computer Science · Arts and Humanities · #Generative Adversarial Networks and Image Synthesis #Computer Graphics and Visualization Techniques #Digital Humanities and Scholarship
- WebSRC: A Dataset for Web-Based Structural Reading Comprehension
2021/01/23 by Xingyu Chen, Chen, Xingyu, Zihan Zhao +13 · 11 citations
Computer Science · Social Sciences · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Misinformation and Its Impacts #Topic Modeling #Web Data Mining and Analysis
- SciEval: A Multi-Level Large Language Model Evaluation Benchmark for Scientific Research
2023/08/25 by Liangtai Sun, Yang Han, Sun, Liangtai +12 · 16 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
- MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
2025/05/19 by Ziyang Ma, Yinghao Ma, Ma, Ziyang +62 · 49 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Multimodal Machine Learning Applications #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- AniTalker: Animate Vivid and Diverse Talking Faces through Identity-Decoupled Facial Motion Encoding
2024/05/06 by Tao Liu, Feilong Chen, Liu, Tao +11 · 18 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Face recognition and analysis
- LGESQL: Line Graph Enhanced Text-to-SQL Model with Mixed Local and Non-Local Relations
2021/06/02 by Ruisheng Cao, Cao, Ruisheng, Lu Chen +9 · 9 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling #Web Data Mining and Analysis
- Recent Advances in Discrete Speech Tokens: A Review
2025/02/10 by Yiwei Guo, Zhihan Li, Guo, Yiwei +16 · 20 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Internet Traffic Analysis and Secure E-voting #Multimedia (cs.MM) #Signal Processing (eess.SP) #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering #semigroups and automata theory
- Developing ChemDFM as a large language foundation model for chemistry
2024/01/26 by Zihan Zhao, Da Ma, Zhao, Zihan +23 · 10 citations
Computer Science · Materials Science · #Topic Modeling #Machine Learning in Materials Science #Advanced Text Analysis Techniques
- Towards Universal Speech Discrete Tokens: A Case Study for ASR and TTS
2023/09/14 by Yifan Yang, Feiyu Shen, Yang, Yifan +11 · 8 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
- SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training
2024/12/20 by Wenxi Chen, Chen, Wenxi, Ziyang Ma +27 · 16 citations
Computer Science · #Speech and dialogue systems #Speech Recognition and Synthesis
- Spider2-V: How Far Are Multimodal Agents From Automating Data Science and Engineering Workflows?
2024/07/15 by Ruisheng Cao, Cao, Ruisheng, Fangyu Lei +43 · 10 citations
Computer Science · #Semantic Web and Ontologies
- D4: a Chinese Dialogue Dataset for Depression-Diagnosis-Oriented Chat
2022/05/24 by Binwei Yao, Yao, Binwei, Chao Shi +13 · 5 citations
Psychology · Computer Science · #Mental Health via Writing #Digital Mental Health Interventions #Machine Learning in Healthcare
- LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec
2024/10/21 by Yiwei Guo, Guo, Yiwei, Zhihan Li +9 · 8 citations
Computer Science · #Advanced Data Compression Techniques #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Reducing Tool Hallucination via Reliability Alignment
2024/12/05 by Hongshen Xu, Xu, Hongshen, Zhu, Zichen +13 · 8 citations
Computer Science · Engineering · #Anomaly Detection Techniques and Applications #Computation and Language (cs.CL) #FOS: Computer and information sciences #Safety Systems Engineering in Autonomy #Software Reliability and Analysis Research
- URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models
2025/02/25 by Yan, Ruiqi, Xiquan Li, Wenxi Chen +12 · 9 citations
Computer Science · #Topic Modeling #Speech and dialogue systems #Multimodal Machine Learning Applications
- IBSEN: Director-Actor Agent Collaboration for Controllable and Interactive Drama Script Generation
2024/07/01 by Senyu Han, Han, Senyu, Lu Chen +7 · 5 citations
Engineering · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Human Motion and Animation #Multiagent Systems (cs.MA)
- DiffusionGAN3D: Boosting Text-guided 3D Generation and Domain Adaptation by Combining 3D GANs and Diffusion Priors
2023/12/28 by Biwen Lei, Lei, Biwen, Kai Yu +7 · 4 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications
- Quantum Federated Learning for Distributed Quantum Networks
2022/12/25 by Kai Yu, Yu, Kai, Gao, Fei +2 · 3 citations
Computer Science · #FOS: Physical sciences #Quantum Computing Algorithms and Architecture #Quantum Information and Cryptography #Quantum Physics (quant-ph) #Stochastic Gradient Optimization Techniques
- DiffDub: Person-generic Visual Dubbing Using Inpainting Renderer with Diffusion Auto-encoder
2023/11/03 by Tao Liu, Chenpeng Du, Liu, Tao +7 · 3 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Face recognition and analysis #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications
- Planning Jerk-Optimized Trajectory with Discrete-Time Constraints for Redundant Robots
2019/09/14 by Chengkai Dai, Sylvain Lefèbvre, Dai, Chengkai +7 · 2 citations
Computer Science · Engineering · #FOS: Computer and information sciences #Robot Manipulation and Learning #Robotic Mechanisms and Dynamics #Robotic Path Planning Algorithms #Robotics (cs.RO)
- Infinite Hidden Relational Models
2012/06/27 by Zhao Xu, Xu, Zhao, Volker Tresp +5 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Bayesian Modeling and Causal Inference #Databases (cs.DB) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Large Scale Strongly Supervised Ensemble Metric Learning, with Applications to Face Verification and Retrieval
2012/12/25 by Chang Huang, Shenghuo Zhu, Huang, Chang +3 · 2 citations
Computer Science · #Face recognition and analysis #Face and Expression Recognition #Advanced Image and Video Retrieval Techniques
- Anticancer potential of Hericium erinaceus extracts against human gastrointestinal cancers
2014/04/01 by Guang Li, Kai Yu, Fushuang Li +5 · 1 citation
- VQTalker: Towards Multilingual Talking Avatars through Facial Motion Tokenization
2024/12/13 by Tao Liu, Ziyang Ma, Liu, Tao +11 · 4 citations
Computer Science · Engineering · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Face recognition and analysis #Human Motion and Animation #Human Pose and Action Recognition
- UrFound: Towards Universal Retinal Foundation Models via Knowledge-Guided Masked Modeling
2024/08/10 by Kai Yu, Yu, Kai, Yang Zhou +13 · 3 citations
Medicine · Decision Sciences · Engineering · #Retinal Imaging and Analysis #Scientific Computing and Data Management #Robotics and Automated Systems
- Contrastive Learning With Audio Discrimination For Customizable Keyword Spotting In Continuous Speech
2024/01/12 by Yu Xi, Xi, Yu, Baochen Yang +7 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Text and Document Classification Technologies #electronic engineering #information engineering
- Semantic Parsing with Dual Learning
2019/07/10 by Ruisheng Cao, Cao, Ruisheng, Su Zhu +7 · 1 citation
Computer Science · #Natural Language Processing Techniques #Multimodal Machine Learning Applications #Topic Modeling
- Attention-Constrained Inference for Robust Decoder-Only Text-to-Speech
2024/04/30 by Hankun Wang, Wang, Hankun, Chenpeng Du +9 · 2 citations
Computer Science · #Speech Recognition and Synthesis
- Neural Directed Speech Enhancement with Dual Microphone Array in High Noise Scenario
2024/12/24 by Wen Wen, Wen, Wen, Qiang Zhou +8 · 3 citations
Computer Science · Engineering · #Speech and Audio Processing #Speech Recognition and Synthesis #Advanced Adaptive Filtering Techniques
- Modular End-to-end Automatic Speech Recognition Framework for Acoustic-to-word Model
2020/07/31 by Qi Liu, Zhehuai Chen, Liu, Qi +9 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators
2025/10/01 by Hanyang Li, Li, Hengtao, Pengxiang Ding +19 · 6 citations
Computer Science · Engineering · #Multimodal Machine Learning Applications #Human Pose and Action Recognition #Robot Manipulation and Learning
- SLAM-AAC: Enhancing Audio Captioning with Paraphrasing Augmentation and CLAP-Refine through LLMs
2024/10/12 by Wenxi Chen, Chen, Wenxi, Ziyang Ma +13 · 2 citations
Arts and Humanities · Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimodal Machine Learning Applications #Music and Audio Processing #Sound (cs.SD) #Subtitles and Audiovisual Media #electronic engineering #information engineering
- Towards Instance-adaptive Inference for Federated Learning
2023/08/11 by Chun-Mei Feng, Kai Yu, Feng, Chun-Mei +9 · 1 citation
Computer Science · Medicine · #COVID-19 diagnosis using AI #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning in Healthcare #Privacy-Preserving Technologies in Data
- AdaEAGLE: Optimizing Speculative Decoding via Explicit Modeling of Adaptive Draft Structures
2024/12/25 by Situo Zhang, Hankun Wang, Zhang, Situo +11 · 2 citations
Computer Science · #Natural Language Processing Techniques
- Weakly-supervised Learning of Mid-level Features for Pedestrian Attribute Recognition and Localization
2016/11/17 by Kai Yu, Yu, Kai, Biao Leng +7 · 1 citation
Computer Science · #Advanced Neural Network Applications #Video Surveillance and Tracking Methods #Domain Adaptation and Few-Shot Learning
- DSE-TTS: Dual Speaker Embedding for Cross-Lingual Text-to-Speech
2023/06/25 by Sen Liu, Liu, Sen, Yiwei Guo +7 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Enhancing LLM Reliability via Explicit Knowledge Boundary Modeling
2025/03/04 by Hongshen Xu, Zheng, Hang, Xu, Hongshen +8 · 3 citations
Computer Science · Decision Sciences · #Topic Modeling #Advanced Graph Neural Networks #Data Quality and Management
- Acoustic BPE for Speech Generation with Discrete Tokens
2023/10/23 by Feiyu Shen, Yiwei Guo, Shen, Feiyu +7 · 1 citation
Computer Science · #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling
- A BiRGAT Model for Multi-intent Spoken Language Understanding with Hierarchical Semantic Frames
2024/02/28 by Hongshen Xu, Ruisheng Cao, Xu, Hongshen +11 · 1 citation
Computer Science · #Natural Language Processing Techniques #Speech and dialogue systems #Topic Modeling
- MULTI: Multimodal Understanding Leaderboard with Text and Images
2024/02/05 by Zichen Zhu, Yang Xu, Zhu, Zichen +25 · 1 citation
Psychology · Social Sciences · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Education and Technology Integration #FOS: Computer and information sciences #Language, Metaphor, and Cognition
- TDT-KWS: Fast And Accurate Keyword Spotting Using Token-and-duration Transducer
2024/03/20 by Yu Xi, Hao Li, Xi, Yu +9 · 1 citation
Computer Science · #Advanced Text Analysis Techniques #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #electronic engineering #information engineering
- Is Self-knowledge and Action Consistent or Not: Investigating Large Language Model's Personality
2024/02/22 by Yiming Ai, Ai, Yiming, Zhiwei He +13 · 1 citation
Computer Science · #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling
- Multilingual Brain Surgeon: Large Language Models Can be Compressed Leaving No Language Behind
2024/04/06 by Hongchuan Zeng, Zeng, Hongchuan, Hongshen Xu +5 · 1 citation
Medicine · #Artificial Intelligence in Healthcare and Education #Computation and Language (cs.CL) #FOS: Computer and information sciences #Radiomics and Machine Learning in Medical Imaging
- AlignSum: Data Pyramid Hierarchical Fine-tuning for Aligning with Human Summarization Preference
2024/10/01 by Yang Han, Han, Yang, Yiming Wang +7 · 2 citations
Computer Science · #Data Management and Algorithms
- Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis
2025/04/14 by Yifan Yang, Shujie Liu, Yang, Yifan +22 · 4 citations
Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Topic Modeling
- Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling
2025/05/26 by Yushen Chen, Zheng, Qixi, Zhikang Niu +10 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- vec2wav 2.0: Advancing Voice Conversion via Discrete Token Vocoders
2024/09/03 by Yiwei Guo, Zhihan Li, Guo, Yiwei +13 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- NTC-KWS: Noise-aware CTC for Robust Keyword Spotting
2024/12/17 by Yu Xi, Xi, Yu, Haoyu Li +11 · 1 citation
Computer Science · #Advanced Text Analysis Techniques #Text and Document Classification Technologies
- Phased One-Step Adversarial Equilibrium for Video Diffusion Models
2025/08/28 by Jiaxiang Cheng, Cheng, Jiaxiang, Bing Ma +17 · 3 citations
Physics and Astronomy · Computer Science · #Model Reduction and Neural Networks #Advanced Image Processing Techniques #Adversarial Robustness in Machine Learning
- Reasoning-Driven Retrosynthesis Prediction with Large Language Models via Reinforcement Learning
2025/07/23 by Situo Zhang, Zhang, Situo, Hanqi Li +14 · 3 citations
Computer Science · #Topic Modeling
- Semi-Supervised Text Simplification with Back-Translation and Asymmetric Denoising Autoencoders
2020/04/30 by Yanbin Zhao, Lu Chen, Zhao, Yanbin +5 · 1 citation
Computer Science · #Text Readability and Simplification #Natural Language Processing Techniques #Topic Modeling
- X-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System
2026/07/20 by Yuxiang Zhao, Yichi Zhang, Yanjie An +10
#eess.AS
- Towards High-Level Semantic Intelligence
2026/07/27 by Xiujie Song, Gefei Yang, Yining You +6
#cs.AI
- AgenticASR: Refining Speech Recognition in Real-World Scenarios via an Agentic Approach
2026/07/30 by Zixuan Jiang, Binghao Qiang, Jiaying Chi +3
Computer Science · #cs.AI
- SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings
2026/07/16 by Shuai Wang, Zihan Qian, Ke Zhang +9
#eess.AS #cs.SD