vix.ing · top · new · best · stats · spec

Yu, Kai

  1. F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
    2024/10/09 by Yushen Chen, Zhikang Niu, Chen, Yushen +14 · 2 voices · 141 citations
    Computer Science · #Music and Audio Processing #cs.SD #eess.AS
  2. Bidirectional LSTM-CRF Models for Sequence Tagging
    2015/08/09 by Huang, Zhiheng, Xu, Wei, Yu, Kai · 28 citations
    #Computation and Language (cs.CL) #FOS: Computer and information sciences
  3. Diverse Data Augmentation with Diffusions for Effective Test-time Prompt Tuning
    2023/08/11 by Chun-Mei Feng, Kai Yu, Feng, Chun-Mei +7 · 39 citations
    Biochemistry, Genetics and Molecular Biology · Computer Science · #Cancer-related molecular mechanisms research #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications
  4. PointGPT: Auto-regressively Generative Pre-training from Point Clouds
    2023/05/19 by Chen, Guangyan, Wang, Meiling, Yang, Yi +3 · 23 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  5. WebSRC: A Dataset for Web-Based Structural Reading Comprehension
    2021/01/23 by Xingyu Chen, Zihan Zhao, Chen, Xingyu +13 · 12 citations
    Computer Science · Social Sciences · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Misinformation and Its Impacts #Topic Modeling #Web Data Mining and Analysis
  6. META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI
    2022/05/23 by Sun, Liangtai, Chen, Xingyu, Chen, Lu +3 · 14 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences
  7. SciEval: A Multi-Level Large Language Model Evaluation Benchmark for Scientific Research
    2023/08/25 by Liangtai Sun, Sun, Liangtai, Yang Han +12 · 18 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
  8. HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer
    2025/05/28 by Cai Qi, Cai, Qi, Jingwen Chen +40 · 62 citations
    Computer Science · Arts and Humanities · #Generative Adversarial Networks and Image Synthesis #Computer Graphics and Visualization Techniques #Digital Humanities and Scholarship
  9. AniTalker: Animate Vivid and Diverse Talking Faces through Identity-Decoupled Facial Motion Encoding
    2024/05/06 by Tao Liu, Feilong Chen, Liu, Tao +11 · 21 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Face recognition and analysis
  10. LGESQL: Line Graph Enhanced Text-to-SQL Model with Mixed Local and Non-Local Relations
    2021/06/02 by Ruisheng Cao, Lu Chen, Cao, Ruisheng +9 · 11 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling #Web Data Mining and Analysis
  11. MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
    2025/05/19 by Ziyang Ma, Yinghao Ma, Ma, Ziyang +62 · 49 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Multimodal Machine Learning Applications #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  12. VoiceFlow: Efficient Text-to-Speech with Rectified Flow Matching
    2023/09/10 by Guo, Yiwei, Du, Chenpeng, Ma, Ziyang +2 · 13 citations
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Sound (cs.SD) #electronic engineering #information engineering
  13. A Survey on Speech Large Language Models for Understanding
    2024/10/24 by Peng, Jing, Wang, Yucheng, Li, Bohan +9 · 23 citations
    #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #electronic engineering #information engineering
  14. Mobile-Env: Building Qualified Evaluation Benchmarks for LLM-GUI Interaction
    2023/05/14 by Danyang Zhang, Zhang, Danyang, Shen, Zhennan +9 · 10 citations
    Computer Science · #AI in Service Interactions #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling
  15. ACT-SQL: In-Context Learning for Text-to-SQL with Automatically-Generated Chain-of-Thought
    2023/10/26 by Zhang, Hanchong, Cao, Ruisheng, Chen, Lu +2 · 11 citations
    #Computation and Language (cs.CL) #FOS: Computer and information sciences
  16. Recent Advances in Discrete Speech Tokens: A Review
    2025/02/10 by Yiwei Guo, Zhihan Li, Guo, Yiwei +16 · 23 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Internet Traffic Analysis and Secure E-voting #Multimedia (cs.MM) #Signal Processing (eess.SP) #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering #semigroups and automata theory
  17. Developing ChemDFM as a large language foundation model for chemistry
    2024/01/26 by Zihan Zhao, Zhao, Zihan, Da Ma +23 · 12 citations
    Computer Science · Materials Science · #Topic Modeling #Machine Learning in Materials Science #Advanced Text Analysis Techniques
  18. Spider2-V: How Far Are Multimodal Agents From Automating Data Science and Engineering Workflows?
    2024/07/15 by Ruisheng Cao, Fangyu Lei, Cao, Ruisheng +43 · 14 citations
    Computer Science · #Semantic Web and Ontologies
  19. Margin Matters: Towards More Discriminative Deep Neural Network Embeddings for Speaker Recognition
    2019/06/18 by Xiang, Xu, Wang, Shuai, Huang, Houjun +2 · 6 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  20. SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training
    2024/12/20 by Wenxi Chen, Ziyang Ma, Chen, Wenxi +27 · 21 citations
    Computer Science · #Speech and dialogue systems #Speech Recognition and Synthesis
  21. ChatCite: LLM Agent with Human Workflow Guidance for Comparative Literature Summary
    2024/03/05 by Yutong Li, Lu Chen, Li, Yutong +7 · 11 citations
    Biochemistry, Genetics and Molecular Biology · Computer Science · #68T50 #Artificial Intelligence (cs.AI) #Biomedical Text Mining and Ontologies #Computation and Language (cs.CL) #FOS: Computer and information sciences #I.2.7 #Information Retrieval (cs.IR) #Topic Modeling
  22. GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement
    2024/06/17 by Yifan Yang, Yang, Yifan, Zheshu Song +29 · 13 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #electronic engineering #information engineering
  23. EmoDiff: Intensity Controllable Emotional Text-to-Speech with Soft-Label Guidance
    2022/11/17 by Guo, Yiwei, Du, Chenpeng, Chen, Xie +1 · 8 citations
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  24. Towards Universal Speech Discrete Tokens: A Case Study for ASR and TTS
    2023/09/14 by Yifan Yang, Feiyu Shen, Yang, Yifan +11 · 9 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
  25. Semantic Alignment-Enhanced Code Translation via an LLM-Based Multi-Agent System
    2024/09/30 by Yuan, Zhiqiang, Chen, Weitong, Wang, Hanlin +3 · 11 citations
    #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Software Engineering (cs.SE)
  26. D4: a Chinese Dialogue Dataset for Depression-Diagnosis-Oriented Chat
    2022/05/24 by Binwei Yao, Chao Shi, Yao, Binwei +13 · 5 citations
    Psychology · Computer Science · #Mental Health via Writing #Digital Mental Health Interventions #Machine Learning in Healthcare
  27. Rejection Improves Reliability: Training LLMs to Refuse Unknown Questions Using RL from Knowledge Feedback
    2024/03/27 by Hongshen Xu, Xu, Hongshen, Zichen Zhu +11 · 8 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Intelligent Tutoring Systems and Adaptive Learning #Topic Modeling
  28. IBSEN: Director-Actor Agent Collaboration for Controllable and Interactive Drama Script Generation
    2024/07/01 by Senyu Han, Han, Senyu, Lu Chen +7 · 8 citations
    Engineering · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Human Motion and Animation #Multiagent Systems (cs.MA)
  29. Enhancing Diagnostic Accuracy in Rare and Common Fundus Diseases with a Knowledge-Rich Vision-Language Model
    2024/06/13 by Wang, Meng, Lin, Tian, Lin, Aidi +46 · 8 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Image and Video Processing (eess.IV) #electronic engineering #information engineering
  30. Towards Weakly Supervised Text-to-Audio Grounding
    2024/01/05 by Xu, Xuenan, Ma, Ziyang, Wu, Mengyue +1 · 7 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  31. Reducing Tool Hallucination via Reliability Alignment
    2024/12/05 by Hongshen Xu, Xu, Hongshen, Zhu, Zichen +13 · 9 citations
    Computer Science · Engineering · #Anomaly Detection Techniques and Applications #Computation and Language (cs.CL) #FOS: Computer and information sciences #Safety Systems Engineering in Autonomy #Software Reliability and Analysis Research
  32. LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec
    2024/10/21 by Yiwei Guo, Guo, Yiwei, Zhihan Li +9 · 8 citations
    Computer Science · #Advanced Data Compression Techniques #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  33. ShadowGNN: Graph Projection Neural Network for Text-to-SQL Parser
    2021/04/10 by Zhi Chen, Chen, Zhi, Lu Chen +11 · 3 citations
    Computer Science · #Advanced Graph Neural Networks #Computation and Language (cs.CL) #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Natural Language Processing Techniques #Topic Modeling
  34. URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models
    2025/02/25 by Xiquan Li, Yan, Ruiqi, Li, Xiquan +12 · 9 citations
    Computer Science · #Topic Modeling #Speech and dialogue systems #Multimodal Machine Learning Applications
  35. In-Context Learning for MIMO Equalization Using Transformer-Based Sequence Models
    2023/11/10 by Zecchin, Matteo, Yu, Kai, Simeone, Osvaldo · 5 citations
    #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Information Theory (cs.IT) #Machine Learning (cs.LG)
  36. DiffusionGAN3D: Boosting Text-guided 3D Generation and Domain Adaptation by Combining 3D GANs and Diffusion Priors
    2023/12/28 by Biwen Lei, Lei, Biwen, Kai Yu +7 · 4 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications
  37. Quantum Federated Learning for Distributed Quantum Networks
    2022/12/25 by Kai Yu, Yu, Kai, Song Lin +2 · 3 citations
    Computer Science · #FOS: Physical sciences #Quantum Computing Algorithms and Architecture #Quantum Information and Cryptography #Quantum Physics (quant-ph) #Stochastic Gradient Optimization Techniques
  38. Text-to-Audio Grounding: Building Correspondence Between Captions and Sound Events
    2021/02/23 by Xu, Xuenan, Dinkel, Heinrich, Wu, Mengyue +1 · 2 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  39. Large Language Models Are Semi-Parametric Reinforcement Learning Agents
    2023/06/09 by Zhang, Danyang, Chen, Lu, Zhang, Situo +3 · 3 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences
  40. Exploring Separable Attention for Multi-Contrast MR Image Super-Resolution
    2021/09/03 by Chun-Mei Feng, Yunlu Yan, Feng, Chun-Mei +8 · 2 citations
    Computer Science · Engineering · #Advanced Image Processing Techniques #Photoacoustic and Ultrasonic Imaging #Image and Signal Denoising Methods
  41. DiffDub: Person-generic Visual Dubbing Using Inpainting Renderer with Diffusion Auto-encoder
    2023/11/03 by Tao Liu, Liu, Tao, Chenpeng Du +7 · 3 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Face recognition and analysis #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications
  42. VQTalker: Towards Multilingual Talking Avatars through Facial Motion Tokenization
    2024/12/13 by Tao Liu, Liu, Tao, Ziyang Ma +11 · 6 citations
    Computer Science · Engineering · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Face recognition and analysis #Human Motion and Animation #Human Pose and Action Recognition
  43. SEF-VC: Speaker Embedding Free Zero-Shot Voice Conversion with Cross Attention
    2023/12/14 by Li, Junjie, Guo, Yiwei, Chen, Xie +1 · 3 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  44. Beyond the Status Quo: A Contemporary Survey of Advances and Challenges in Audio Captioning
    2022/05/11 by Xuenan Xu, Xu, Xuenan, Mengyue Wu +4 · 2 citations
    Arts and Humanities · Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #Subtitles and Audiovisual Media #electronic engineering #information engineering
  45. CoE-SQL: In-Context Learning for Multi-Turn Text-to-SQL with Chain-of-Editions
    2024/05/04 by Zhang, Hanchong, Cao, Ruisheng, Xu, Hongshen +2 · 3 citations
    #Computation and Language (cs.CL) #FOS: Computer and information sciences
  46. Infinite Hidden Relational Models
    2012/06/27 by Zhao Xu, Volker Tresp, Xu, Zhao +5 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Bayesian Modeling and Causal Inference #Databases (cs.DB) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  47. Large Scale Strongly Supervised Ensemble Metric Learning, with Applications to Face Verification and Retrieval
    2012/12/25 by Chang Huang, Huang, Chang, Shenghuo Zhu +3 · 2 citations
    Computer Science · #Face recognition and analysis #Face and Expression Recognition #Advanced Image and Video Retrieval Techniques
  48. Cell-Free Multi-User MIMO Equalization via In-Context Learning
    2024/04/08 by Zecchin, Matteo, Yu, Kai, Simeone, Osvaldo · 3 citations
    #FOS: Computer and information sciences #FOS: Electrical engineering #Information Theory (cs.IT) #Machine Learning (cs.LG) #Signal Processing (eess.SP) #electronic engineering #information engineering
  49. Converging to a Lingua Franca: Evolution of Linguistic Regions and Semantics Alignment in Multilingual Large Language Models
    2024/10/15 by Zeng, Hongchuan, Han, Senyu, Chen, Lu +1 · 4 citations
    #Computation and Language (cs.CL) #FOS: Computer and information sciences
  50. UrFound: Towards Universal Retinal Foundation Models via Knowledge-Guided Masked Modeling
    2024/08/10 by Kai Yu, Yang Zhou, Yu, Kai +13 · 3 citations
    Medicine · Decision Sciences · Engineering · #Retinal Imaging and Analysis #Scientific Computing and Data Management #Robotics and Automated Systems
  51. On Modular Training of Neural Acoustics-to-Word Model for LVCSR
    2018/03/03 by Chen, Zhehuai, Liu, Qi, Li, Hao +1 · 1 citation
    #68T10 #Computation and Language (cs.CL) #FOS: Computer and information sciences #I.2.7
  52. Contrastive Learning With Audio Discrimination For Customizable Keyword Spotting In Continuous Speech
    2024/01/12 by Yu Xi, Xi, Yu, Baochen Yang +7 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Text and Document Classification Technologies #electronic engineering #information engineering
  53. End-to-End Monaural Multi-speaker ASR System without Pretraining
    2018/11/05 by Chang, Xuankai, Qian, Yanmin, Yu, Kai +1 · 1 citation
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  54. ChemDFM-R: A Chemical Reasoning LLM Enhanced with Atomized Chemical Knowledge
    2025/07/29 by Zhao, Zihan, Chen, Bo, Wan, Ziping +12 · 5 citations
    #Artificial Intelligence (cs.AI) #Computational Engineering #FOS: Computer and information sciences #Finance #and Science (cs.CE)
  55. Semantic Parsing with Dual Learning
    2019/07/10 by Ruisheng Cao, Su Zhu, Cao, Ruisheng +7 · 1 citation
    Computer Science · #Natural Language Processing Techniques #Multimodal Machine Learning Applications #Topic Modeling
  56. MobA: Multifaceted Memory-Enhanced Adaptive Planning for Efficient Mobile Task Automation
    2024/10/17 by Zichen Zhu, Zhu, Zichen, Hao Tang +24 · 3 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Mobile Agent-Based Network Management #Multi-Agent Systems and Negotiation #Multiagent Systems (cs.MA) #Optimization and Search Problems
  57. Hierarchical Multimodal Pre-training for Visually Rich Webpage Understanding
    2024/02/28 by Xu, Hongshen, Chen, Lu, Zhao, Zihan +4 · 2 citations
    #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  58. Unsupervised Dual Paraphrasing for Two-stage Semantic Parsing
    2020/05/27 by Cao, Ruisheng, Zhu, Su, Yang, Chenyu +5 · 1 citation
    #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  59. Attention-Constrained Inference for Robust Decoder-Only Text-to-Speech
    2024/04/30 by Hankun Wang, Chenpeng Du, Wang, Hankun +9 · 2 citations
    Computer Science · #Speech Recognition and Synthesis
  60. AdaEAGLE: Optimizing Speculative Decoding via Explicit Modeling of Adaptive Draft Structures
    2024/12/25 by Situo Zhang, Zhang, Situo, Hankun Wang +11 · 3 citations
    Computer Science · #Natural Language Processing Techniques
  61. Neural Directed Speech Enhancement with Dual Microphone Array in High Noise Scenario
    2024/12/24 by Wen Wen, Qiang Zhou, Wen, Wen +8 · 3 citations
    Computer Science · Engineering · #Speech and Audio Processing #Speech Recognition and Synthesis #Advanced Adaptive Filtering Techniques
  62. Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity
    2024/12/03 by Ma, Da, Chen, Lu, Zhang, Situo +8 · 3 citations
    #Computation and Language (cs.CL) #FOS: Computer and information sciences
  63. Modular End-to-end Automatic Speech Recognition Framework for Acoustic-to-word Model
    2020/07/31 by Qi Liu, Zhehuai Chen, Liu, Qi +9 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  64. Alignment for Efficient Tool Calling of Large Language Models
    2025/03/09 by Xu, Hongshen, Wang, Zihan, Zhu, Zichen +4 · 4 citations
    #Computation and Language (cs.CL) #FOS: Computer and information sciences
  65. From Generalist to Specialist: A Survey of Large Language Models for Chemistry
    2024/12/28 by Han, Yang, Wan, Ziping, Chen, Lu +2 · 3 citations
    #Artificial Intelligence (cs.AI) #Chemical Physics (physics.chem-ph) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Physical sciences #Machine Learning (cs.LG)
  66. Unsupervised word-level prosody tagging for controllable speech synthesis
    2022/02/15 by Guo, Yiwei, Du, Chenpeng, Yu, Kai · 1 citation
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  67. ProgRM: Build Better GUI Agents with Progress Rewards
    2025/05/23 by Zhang, Danyang, Zhang, Situo, Yang, Ziyue +5 · 5 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  68. TIE: Topological Information Enhanced Structural Reading Comprehension on Web Pages
    2022/05/13 by Zhao, Zihan, Chen, Lu, Cao, Ruisheng +3 · 1 citation
    #Computation and Language (cs.CL) #FOS: Computer and information sciences
  69. UniDU: Towards A Unified Generative Dialogue Understanding Framework
    2022/04/10 by Chen, Zhi, Chen, Lu, Chen, Bei +5 · 1 citation
    #Computation and Language (cs.CL) #FOS: Computer and information sciences
  70. Audio-text Retrieval in Context
    2022/03/25 by Lou, Siyu, Xu, Xuenan, Wu, Mengyue +1 · 1 citation
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  71. SLAM-AAC: Enhancing Audio Captioning with Paraphrasing Augmentation and CLAP-Refine through LLMs
    2024/10/12 by Wenxi Chen, Chen, Wenxi, Ziyang Ma +13 · 2 citations
    Arts and Humanities · Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimodal Machine Learning Applications #Music and Audio Processing #Sound (cs.SD) #Subtitles and Audiovisual Media #electronic engineering #information engineering
  72. Towards Instance-adaptive Inference for Federated Learning
    2023/08/11 by Chun-Mei Feng, Feng, Chun-Mei, Kai Yu +9 · 1 citation
    Computer Science · Medicine · #COVID-19 diagnosis using AI #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning in Healthcare #Privacy-Preserving Technologies in Data
  73. Enhancing LLM Reliability via Explicit Knowledge Boundary Modeling
    2025/03/04 by Zheng, Hang, Hongshen Xu, Yuncong Liu +8 · 4 citations
    Computer Science · Decision Sciences · #Topic Modeling #Advanced Graph Neural Networks #Data Quality and Management
  74. FakeSound: Deepfake General Audio Detection
    2024/06/12 by Zeyu Xie, Baihan Li, Xie, Zeyu +9 · 2 citations
    Computer Science · #68Txx #Audio and Speech Processing (eess.AS) #Digital Media Forensic Detection #FOS: Computer and information sciences #FOS: Electrical engineering #I.2 #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  75. Improving Few-Shot Learning for Talking Face System with TTS Data Augmentation
    2023/03/09 by Chen, Qi, Ma, Ziyang, Liu, Tao +4 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Sound (cs.SD) #electronic engineering #information engineering
  76. Streaming Keyword Spotting Boosted by Cross-layer Discrimination Consistency
    2024/12/17 by Yu Xi, Haoyu Li, Xi, Yu +9 · 2 citations
    Computer Science · #Advanced Text Analysis Techniques #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  77. Weakly-supervised Learning of Mid-level Features for Pedestrian Attribute Recognition and Localization
    2016/11/17 by Kai Yu, Biao Leng, Yu, Kai +7 · 1 citation
    Computer Science · #Advanced Neural Network Applications #Video Surveillance and Tracking Methods #Domain Adaptation and Few-Shot Learning
  78. DSE-TTS: Dual Speaker Embedding for Cross-Lingual Text-to-Speech
    2023/06/25 by Sen Liu, Yiwei Guo, Liu, Sen +7 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  79. Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis
    2025/04/14 by Yifan Yang, Yang, Yifan, Shujie Liu +22 · 5 citations
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Topic Modeling
  80. Acoustic BPE for Speech Generation with Discrete Tokens
    2023/10/23 by Feiyu Shen, Yiwei Guo, Shen, Feiyu +7 · 1 citation
    Computer Science · #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling
  81. A BiRGAT Model for Multi-intent Spoken Language Understanding with Hierarchical Semantic Frames
    2024/02/28 by Hongshen Xu, Ruisheng Cao, Xu, Hongshen +11 · 1 citation
    Computer Science · #Natural Language Processing Techniques #Speech and dialogue systems #Topic Modeling
  82. MULTI: Multimodal Understanding Leaderboard with Text and Images
    2024/02/05 by Zichen Zhu, Zhu, Zichen, Yang Xu +25 · 1 citation
    Psychology · Social Sciences · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Education and Technology Integration #FOS: Computer and information sciences #Language, Metaphor, and Cognition
  83. TDT-KWS: Fast And Accurate Keyword Spotting Using Token-and-duration Transducer
    2024/03/20 by Yu Xi, Hao Li, Xi, Yu +9 · 1 citation
    Computer Science · #Advanced Text Analysis Techniques #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #electronic engineering #information engineering
  84. Is Self-knowledge and Action Consistent or Not: Investigating Large Language Model's Personality
    2024/02/22 by Yiming Ai, Ai, Yiming, Zhiwei He +13 · 1 citation
    Computer Science · #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling
  85. Multilingual Brain Surgeon: Large Language Models Can be Compressed Leaving No Language Behind
    2024/04/06 by Hongchuan Zeng, Zeng, Hongchuan, Hongshen Xu +5 · 1 citation
    Medicine · #Artificial Intelligence in Healthcare and Education #Computation and Language (cs.CL) #FOS: Computer and information sciences #Radiomics and Machine Learning in Medical Imaging
  86. Sparsity-Accelerated Training for Large Language Models
    2024/06/03 by Da Ma, Lu Chen, Ma, Da +15 · 1 citation
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
  87. Evolving Subnetwork Training for Large Language Models
    2024/06/11 by Li, Hanqi, Chen, Lu, Ma, Da +3 · 1 citation
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences
  88. AlignSum: Data Pyramid Hierarchical Fine-tuning for Aligning with Human Summarization Preference
    2024/10/01 by Yang Han, Han, Yang, Yiming Wang +7 · 2 citations
    Computer Science · #Data Management and Algorithms
  89. A Detailed Audio-Text Data Simulation Pipeline using Single-Event Sounds
    2024/03/07 by Xu, Xuenan, Xu, Xiaohang, Xie, Zeyu +3 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  90. Minimizing False Positives in Static Bug Detection via LLM-Enhanced Path Feasibility Analysis
    2025/06/12 by Xueying Du, Du, Xueying, Kaiping Yu +15 · 2 citations
    Computer Science · #Software Engineering Research #Software Testing and Debugging Techniques #Software Engineering Techniques and Practices
  91. Disentangling Foreground and Background Motion for Enhanced Realism in Human Video Generation
    2024/05/26 by Jinlin Liu, Liu, Jinlin, Kai Yu +7 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #Digital Media Forensic Detection #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Human Pose and Action Recognition
  92. Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling
    2025/05/26 by Zheng, Qixi, Yushen Chen, Zhikang Niu +10 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  93. Fast and High-Quality Auto-Regressive Speech Synthesis via Speculative Decoding
    2024/10/29 by Bohan Li, Hankun Wang, Li, Bohan +7 · 1 citation
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Speech and Audio Processing
  94. Investigating Local and Global Information for Automated Audio Captioning with Transfer Learning
    2021/02/23 by Xu, Xuenan, Dinkel, Heinrich, Wu, Mengyue +2 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  95. vec2wav 2.0: Advancing Voice Conversion via Discrete Token Vocoders
    2024/09/03 by Yiwei Guo, Guo, Yiwei, Zhihan Li +13 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  96. SciDFM: A Large Language Model with Mixture-of-Experts for Science
    2024/09/27 by Sun, Liangtai, Luo, Danyu, Ma, Da +7 · 1 citation
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences
  97. NTC-KWS: Noise-aware CTC for Robust Keyword Spotting
    2024/12/17 by Yu Xi, Xi, Yu, Haoyu Li +11 · 1 citation
    Computer Science · #Advanced Text Analysis Techniques #Text and Document Classification Technologies
  98. Phased One-Step Adversarial Equilibrium for Video Diffusion Models
    2025/08/28 by Jiaxiang Cheng, Cheng, Jiaxiang, Bing Ma +17 · 3 citations
    Physics and Astronomy · Computer Science · #Model Reduction and Neural Networks #Advanced Image Processing Techniques #Adversarial Robustness in Machine Learning
  99. Enhancing Speech-to-Speech Dialogue Modeling with End-to-End Retrieval-Augmented Generation
    2025/04/27 by Feng, Pengchao, Ma, Ziyang, Chen, Wenxi +4 · 2 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Information Retrieval (cs.IR)
  100. Why Do Speech Language Models Fail to Generate Semantically Coherent Outputs? A Modality Evolving Perspective
    2024/12/22 by Wang, Hankun, Wang, Haoran, Guo, Yiwei +4 · 1 citation
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  101. Quantum Graph Convolutional Networks Based on Spectral Methods
    2025/03/09 by Ye, Zi, Yu, Kai, Lin, Song · 1 citation
    #FOS: Physical sciences #Quantum Physics (quant-ph)
  102. Unlocking Temporal Flexibility: Neural Speech Codec with Variable Frame Rate
    2025/05/22 by Zhang, Hanglei, Guo, Yiwei, Li, Zhihan +3 · 2 citations
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  103. Efficient Context and Schema Fusion Networks for Multi-Domain Dialogue State Tracking
    2020/04/07 by Su Zhu, Jieyu Li, Zhu, Su +5 · 2 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech and dialogue systems #Topic Modeling
  104. Omni-Captioner: Data Pipeline, Models, and Benchmark for Omni Detailed Perception
    2025/10/14 by Ma, Ziyang, Xu, Ruiyang, Xing, Zhenghao +9 · 2 citations
    #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimedia (cs.MM) #Sound (cs.SD)
  105. Reasoning-Driven Retrosynthesis Prediction with Large Language Models via Reinforcement Learning
    2025/07/23 by Situo Zhang, Hanqi Li, Zhang, Situo +14 · 3 citations
    Computer Science · #Topic Modeling
  106. Semi-Supervised Text Simplification with Back-Translation and Asymmetric Denoising Autoencoders
    2020/04/30 by Yanbin Zhao, Zhao, Yanbin, Lu Chen +5 · 1 citation
    Computer Science · #Text Readability and Simplification #Natural Language Processing Techniques #Topic Modeling