vix.ing · top · new · best · stats · spec

Kai Yu

  1. F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
    2024/10/09 by Yushen Chen, Zhikang Niu, Chen, Yushen +14 · 2 voices · 125 citations
    Computer Science · #Music and Audio Processing #cs.SD #eess.AS
  2. Diverse Data Augmentation with Diffusions for Effective Test-time Prompt Tuning
    2023/08/11 by Chun-Mei Feng, Kai Yu, Feng, Chun-Mei +7 · 30 citations
    Biochemistry, Genetics and Molecular Biology · Computer Science · #Cancer-related molecular mechanisms research #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications
  3. RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
    2025/04/24 by Zihan Wang, Kangrui Wang, Wang, Zihan +33 · 76 citations
    Computer Science · #Reinforcement Learning in Robotics #Multimodal Machine Learning Applications #Domain Adaptation and Few-Shot Learning
  4. HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer
    2025/05/28 by Cai Qi, Jingwen Chen, Cai, Qi +40 · 61 citations
    Computer Science · Arts and Humanities · #Generative Adversarial Networks and Image Synthesis #Computer Graphics and Visualization Techniques #Digital Humanities and Scholarship
  5. WebSRC: A Dataset for Web-Based Structural Reading Comprehension
    2021/01/23 by Xingyu Chen, Chen, Xingyu, Zihan Zhao +13 · 11 citations
    Computer Science · Social Sciences · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Misinformation and Its Impacts #Topic Modeling #Web Data Mining and Analysis
  6. SciEval: A Multi-Level Large Language Model Evaluation Benchmark for Scientific Research
    2023/08/25 by Liangtai Sun, Yang Han, Sun, Liangtai +12 · 16 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
  7. MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
    2025/05/19 by Ziyang Ma, Yinghao Ma, Ma, Ziyang +62 · 49 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Multimodal Machine Learning Applications #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  8. AniTalker: Animate Vivid and Diverse Talking Faces through Identity-Decoupled Facial Motion Encoding
    2024/05/06 by Tao Liu, Feilong Chen, Liu, Tao +11 · 18 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Face recognition and analysis
  9. LGESQL: Line Graph Enhanced Text-to-SQL Model with Mixed Local and Non-Local Relations
    2021/06/02 by Ruisheng Cao, Cao, Ruisheng, Lu Chen +9 · 9 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling #Web Data Mining and Analysis
  10. Recent Advances in Discrete Speech Tokens: A Review
    2025/02/10 by Yiwei Guo, Zhihan Li, Guo, Yiwei +16 · 20 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Internet Traffic Analysis and Secure E-voting #Multimedia (cs.MM) #Signal Processing (eess.SP) #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering #semigroups and automata theory
  11. Developing ChemDFM as a large language foundation model for chemistry
    2024/01/26 by Zihan Zhao, Da Ma, Zhao, Zihan +23 · 10 citations
    Computer Science · Materials Science · #Topic Modeling #Machine Learning in Materials Science #Advanced Text Analysis Techniques
  12. Towards Universal Speech Discrete Tokens: A Case Study for ASR and TTS
    2023/09/14 by Yifan Yang, Feiyu Shen, Yang, Yifan +11 · 8 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
  13. SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training
    2024/12/20 by Wenxi Chen, Chen, Wenxi, Ziyang Ma +27 · 16 citations
    Computer Science · #Speech and dialogue systems #Speech Recognition and Synthesis
  14. Spider2-V: How Far Are Multimodal Agents From Automating Data Science and Engineering Workflows?
    2024/07/15 by Ruisheng Cao, Cao, Ruisheng, Fangyu Lei +43 · 10 citations
    Computer Science · #Semantic Web and Ontologies
  15. D4: a Chinese Dialogue Dataset for Depression-Diagnosis-Oriented Chat
    2022/05/24 by Binwei Yao, Yao, Binwei, Chao Shi +13 · 5 citations
    Psychology · Computer Science · #Mental Health via Writing #Digital Mental Health Interventions #Machine Learning in Healthcare
  16. LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec
    2024/10/21 by Yiwei Guo, Guo, Yiwei, Zhihan Li +9 · 8 citations
    Computer Science · #Advanced Data Compression Techniques #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  17. Reducing Tool Hallucination via Reliability Alignment
    2024/12/05 by Hongshen Xu, Xu, Hongshen, Zhu, Zichen +13 · 8 citations
    Computer Science · Engineering · #Anomaly Detection Techniques and Applications #Computation and Language (cs.CL) #FOS: Computer and information sciences #Safety Systems Engineering in Autonomy #Software Reliability and Analysis Research
  18. URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models
    2025/02/25 by Yan, Ruiqi, Xiquan Li, Wenxi Chen +12 · 9 citations
    Computer Science · #Topic Modeling #Speech and dialogue systems #Multimodal Machine Learning Applications
  19. IBSEN: Director-Actor Agent Collaboration for Controllable and Interactive Drama Script Generation
    2024/07/01 by Senyu Han, Han, Senyu, Lu Chen +7 · 5 citations
    Engineering · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Human Motion and Animation #Multiagent Systems (cs.MA)
  20. DiffusionGAN3D: Boosting Text-guided 3D Generation and Domain Adaptation by Combining 3D GANs and Diffusion Priors
    2023/12/28 by Biwen Lei, Lei, Biwen, Kai Yu +7 · 4 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications
  21. Quantum Federated Learning for Distributed Quantum Networks
    2022/12/25 by Kai Yu, Yu, Kai, Gao, Fei +2 · 3 citations
    Computer Science · #FOS: Physical sciences #Quantum Computing Algorithms and Architecture #Quantum Information and Cryptography #Quantum Physics (quant-ph) #Stochastic Gradient Optimization Techniques
  22. DiffDub: Person-generic Visual Dubbing Using Inpainting Renderer with Diffusion Auto-encoder
    2023/11/03 by Tao Liu, Chenpeng Du, Liu, Tao +7 · 3 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Face recognition and analysis #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications
  23. Planning Jerk-Optimized Trajectory with Discrete-Time Constraints for Redundant Robots
    2019/09/14 by Chengkai Dai, Sylvain Lefèbvre, Dai, Chengkai +7 · 2 citations
    Computer Science · Engineering · #FOS: Computer and information sciences #Robot Manipulation and Learning #Robotic Mechanisms and Dynamics #Robotic Path Planning Algorithms #Robotics (cs.RO)
  24. Infinite Hidden Relational Models
    2012/06/27 by Zhao Xu, Xu, Zhao, Volker Tresp +5 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Bayesian Modeling and Causal Inference #Databases (cs.DB) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  25. Large Scale Strongly Supervised Ensemble Metric Learning, with Applications to Face Verification and Retrieval
    2012/12/25 by Chang Huang, Shenghuo Zhu, Huang, Chang +3 · 2 citations
    Computer Science · #Face recognition and analysis #Face and Expression Recognition #Advanced Image and Video Retrieval Techniques
  26. Anticancer potential of Hericium erinaceus extracts against human gastrointestinal cancers
    2014/04/01 by Guang Li, Kai Yu, Fushuang Li +5 · 1 citation
  27. VQTalker: Towards Multilingual Talking Avatars through Facial Motion Tokenization
    2024/12/13 by Tao Liu, Ziyang Ma, Liu, Tao +11 · 4 citations
    Computer Science · Engineering · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Face recognition and analysis #Human Motion and Animation #Human Pose and Action Recognition
  28. UrFound: Towards Universal Retinal Foundation Models via Knowledge-Guided Masked Modeling
    2024/08/10 by Kai Yu, Yu, Kai, Yang Zhou +13 · 3 citations
    Medicine · Decision Sciences · Engineering · #Retinal Imaging and Analysis #Scientific Computing and Data Management #Robotics and Automated Systems
  29. Contrastive Learning With Audio Discrimination For Customizable Keyword Spotting In Continuous Speech
    2024/01/12 by Yu Xi, Xi, Yu, Baochen Yang +7 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Text and Document Classification Technologies #electronic engineering #information engineering
  30. Semantic Parsing with Dual Learning
    2019/07/10 by Ruisheng Cao, Cao, Ruisheng, Su Zhu +7 · 1 citation
    Computer Science · #Natural Language Processing Techniques #Multimodal Machine Learning Applications #Topic Modeling
  31. Attention-Constrained Inference for Robust Decoder-Only Text-to-Speech
    2024/04/30 by Hankun Wang, Wang, Hankun, Chenpeng Du +9 · 2 citations
    Computer Science · #Speech Recognition and Synthesis
  32. Neural Directed Speech Enhancement with Dual Microphone Array in High Noise Scenario
    2024/12/24 by Wen Wen, Wen, Wen, Qiang Zhou +8 · 3 citations
    Computer Science · Engineering · #Speech and Audio Processing #Speech Recognition and Synthesis #Advanced Adaptive Filtering Techniques
  33. Modular End-to-end Automatic Speech Recognition Framework for Acoustic-to-word Model
    2020/07/31 by Qi Liu, Zhehuai Chen, Liu, Qi +9 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  34. VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators
    2025/10/01 by Hanyang Li, Li, Hengtao, Pengxiang Ding +19 · 6 citations
    Computer Science · Engineering · #Multimodal Machine Learning Applications #Human Pose and Action Recognition #Robot Manipulation and Learning
  35. SLAM-AAC: Enhancing Audio Captioning with Paraphrasing Augmentation and CLAP-Refine through LLMs
    2024/10/12 by Wenxi Chen, Chen, Wenxi, Ziyang Ma +13 · 2 citations
    Arts and Humanities · Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimodal Machine Learning Applications #Music and Audio Processing #Sound (cs.SD) #Subtitles and Audiovisual Media #electronic engineering #information engineering
  36. Towards Instance-adaptive Inference for Federated Learning
    2023/08/11 by Chun-Mei Feng, Kai Yu, Feng, Chun-Mei +9 · 1 citation
    Computer Science · Medicine · #COVID-19 diagnosis using AI #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning in Healthcare #Privacy-Preserving Technologies in Data
  37. AdaEAGLE: Optimizing Speculative Decoding via Explicit Modeling of Adaptive Draft Structures
    2024/12/25 by Situo Zhang, Hankun Wang, Zhang, Situo +11 · 2 citations
    Computer Science · #Natural Language Processing Techniques
  38. Weakly-supervised Learning of Mid-level Features for Pedestrian Attribute Recognition and Localization
    2016/11/17 by Kai Yu, Yu, Kai, Biao Leng +7 · 1 citation
    Computer Science · #Advanced Neural Network Applications #Video Surveillance and Tracking Methods #Domain Adaptation and Few-Shot Learning
  39. DSE-TTS: Dual Speaker Embedding for Cross-Lingual Text-to-Speech
    2023/06/25 by Sen Liu, Liu, Sen, Yiwei Guo +7 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  40. Enhancing LLM Reliability via Explicit Knowledge Boundary Modeling
    2025/03/04 by Hongshen Xu, Zheng, Hang, Xu, Hongshen +8 · 3 citations
    Computer Science · Decision Sciences · #Topic Modeling #Advanced Graph Neural Networks #Data Quality and Management
  41. Acoustic BPE for Speech Generation with Discrete Tokens
    2023/10/23 by Feiyu Shen, Yiwei Guo, Shen, Feiyu +7 · 1 citation
    Computer Science · #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling
  42. A BiRGAT Model for Multi-intent Spoken Language Understanding with Hierarchical Semantic Frames
    2024/02/28 by Hongshen Xu, Ruisheng Cao, Xu, Hongshen +11 · 1 citation
    Computer Science · #Natural Language Processing Techniques #Speech and dialogue systems #Topic Modeling
  43. MULTI: Multimodal Understanding Leaderboard with Text and Images
    2024/02/05 by Zichen Zhu, Yang Xu, Zhu, Zichen +25 · 1 citation
    Psychology · Social Sciences · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Education and Technology Integration #FOS: Computer and information sciences #Language, Metaphor, and Cognition
  44. TDT-KWS: Fast And Accurate Keyword Spotting Using Token-and-duration Transducer
    2024/03/20 by Yu Xi, Hao Li, Xi, Yu +9 · 1 citation
    Computer Science · #Advanced Text Analysis Techniques #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #electronic engineering #information engineering
  45. Is Self-knowledge and Action Consistent or Not: Investigating Large Language Model's Personality
    2024/02/22 by Yiming Ai, Ai, Yiming, Zhiwei He +13 · 1 citation
    Computer Science · #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling
  46. Multilingual Brain Surgeon: Large Language Models Can be Compressed Leaving No Language Behind
    2024/04/06 by Hongchuan Zeng, Zeng, Hongchuan, Hongshen Xu +5 · 1 citation
    Medicine · #Artificial Intelligence in Healthcare and Education #Computation and Language (cs.CL) #FOS: Computer and information sciences #Radiomics and Machine Learning in Medical Imaging
  47. AlignSum: Data Pyramid Hierarchical Fine-tuning for Aligning with Human Summarization Preference
    2024/10/01 by Yang Han, Han, Yang, Yiming Wang +7 · 2 citations
    Computer Science · #Data Management and Algorithms
  48. Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis
    2025/04/14 by Yifan Yang, Shujie Liu, Yang, Yifan +22 · 4 citations
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Topic Modeling
  49. Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling
    2025/05/26 by Yushen Chen, Zheng, Qixi, Zhikang Niu +10 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  50. vec2wav 2.0: Advancing Voice Conversion via Discrete Token Vocoders
    2024/09/03 by Yiwei Guo, Zhihan Li, Guo, Yiwei +13 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  51. NTC-KWS: Noise-aware CTC for Robust Keyword Spotting
    2024/12/17 by Yu Xi, Xi, Yu, Haoyu Li +11 · 1 citation
    Computer Science · #Advanced Text Analysis Techniques #Text and Document Classification Technologies
  52. Phased One-Step Adversarial Equilibrium for Video Diffusion Models
    2025/08/28 by Jiaxiang Cheng, Cheng, Jiaxiang, Bing Ma +17 · 3 citations
    Physics and Astronomy · Computer Science · #Model Reduction and Neural Networks #Advanced Image Processing Techniques #Adversarial Robustness in Machine Learning
  53. Reasoning-Driven Retrosynthesis Prediction with Large Language Models via Reinforcement Learning
    2025/07/23 by Situo Zhang, Zhang, Situo, Hanqi Li +14 · 3 citations
    Computer Science · #Topic Modeling
  54. Semi-Supervised Text Simplification with Back-Translation and Asymmetric Denoising Autoencoders
    2020/04/30 by Yanbin Zhao, Lu Chen, Zhao, Yanbin +5 · 1 citation
    Computer Science · #Text Readability and Simplification #Natural Language Processing Techniques #Topic Modeling
  55. X-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System
    2026/07/20 by Yuxiang Zhao, Yichi Zhang, Yanjie An +10
    #eess.AS
  56. Towards High-Level Semantic Intelligence
    2026/07/27 by Xiujie Song, Gefei Yang, Yining You +6
    #cs.AI
  57. AgenticASR: Refining Speech Recognition in Real-World Scenarios via an Agentic Approach
    2026/07/30 by Zixuan Jiang, Binghao Qiang, Jiaying Chi +3
    Computer Science · #cs.AI
  58. SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings
    2026/07/16 by Shuai Wang, Zihan Qian, Ke Zhang +9
    #eess.AS #cs.SD