vix.ing · top · new · best · stats · spec

Zhiyong Wu

  1. A Survey on In-context Learning
    2022/12/31 by Qingxiu Dong, Lei Li, Dong, Qingxiu +23 · 4 voices · 185 citations
    Computer Science · #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL
  2. Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
    2024/12/06 by Zhe Chen, Weiyun Wang, Chen, Zhe +86 · 3 voices · 704 citations
    Computer Science · Decision Sciences · #Topic Modeling #Scientific Computing and Data Management #Machine Learning and Data Classification
  3. SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
    2024/01/17 by Kanzhi Cheng, Qiushi Sun, Cheng, Kanzhi +11 · 182 citations
    Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Interactive and Immersive Displays #Virtual Reality Applications and Impacts
  4. DiffuSeq: Sequence to Sequence Text Generation with Diffusion Models
    2022/10/17 by Shansan Gong, Mukai Li, Gong, Shansan +7 · 97 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Music and Audio Processing #Natural Language Processing Techniques #Topic Modeling
  5. OS-ATLAS: A Foundation Action Model for Generalist GUI Agents
    2024/10/30 by Zhiyong Wu, Wu, Zhiyong, Zhenyu Wu +19 · 150 citations
    Computer Science · Psychology · #Reinforcement Learning in Robotics #Multi-Agent Systems and Negotiation #Social Robot Interaction and HRI
  6. Can We Edit Factual Knowledge by In-Context Learning?
    2023/05/22 by Ce Zheng, Lei Li, Zheng, Ce +11 · 68 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning in Healthcare #Natural Language Processing Techniques #Topic Modeling
  7. OS-Copilot: Towards Generalist Computer Agents with Self-Improvement
    2024/02/12 by Zhiyong Wu, Chengcheng Han, Wu, Zhiyong +13 · 55 citations
    Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Multi-Agent Systems and Negotiation #Semantic Web and Ontologies
  8. Self-Adaptive In-Context Learning: An Information Compression Perspective for In-Context Example Selection and Ordering
    2022/12/20 by Zhiyong Wu, Wu, Zhiyong, Yaoxiang Wang +5 · 37 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Music and Audio Processing #Text and Document Classification Technologies
  9. Compositional Exemplars for In-context Learning
    2023/02/11 by Jiacheng Ye, Zhiyong Wu, Ye, Jiacheng +7 · 32 citations
    Computer Science · #Topic Modeling #Natural Language Processing Techniques #Multimodal Machine Learning Applications
  10. OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis
    2024/12/27 by Qiushi Sun, Sun, Qiushi, Kanzhi Cheng +27 · 58 citations
    Computer Science · Engineering · Psychology · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Context-Aware Activity Recognition Systems #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Robotics and Automated Systems #Social Robot Interaction and HRI
  11. ZeroGen: Efficient Zero-shot Learning via Dataset Generation
    2022/02/16 by Jiacheng Ye, Jiahui Gao, Ye, Jiacheng +13 · 20 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
  12. DiffuseStyleGesture: Stylized Audio-Driven Co-Speech Gesture Generation with Diffusion Models
    2023/05/08 by Sicheng Yang, Zhiyong Wu, Yang, Sicheng +13 · 21 citations
    Engineering · Computer Science · #Human Motion and Animation #Hand Gesture Recognition Systems #Human Pose and Action Recognition
  13. Corex: Pushing the Boundaries of Complex Reasoning through Multi-Model Collaboration
    2023/09/30 by Qiushi Sun, Zhangyue Yin, Sun, Qiushi +9 · 19 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling
  14. TDAG: A Multi-Agent Framework based on Dynamic Task Decomposition and Agent Generation
    2024/02/15 by Yaoxiang Wang, Zhiyong Wu, Wang, Yaoxiang +5 · 18 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multi-Agent Systems and Negotiation
  15. MFA-Conformer: Multi-scale Feature Aggregation Conformer for Automatic Speaker Verification
    2022/03/29 by Yang Zhang, Zhiqiang Lv, Zhang, Yang +13 · 11 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  16. AV-SepFormer: Cross-Attention SepFormer for Audio-Visual Target Speaker Extraction
    2023/06/25 by Jiuxin Lin, Xinyu Cai, Lin, Jiuxin +17 · 14 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  17. Perturbed Masking: Parameter-free Probing for Analyzing and Interpreting BERT
    2020/04/30 by Zhiyong Wu, Wu, Zhiyong, Yun Chen +4 · 8 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
  18. SECap: Speech Emotion Captioning with Large Language Model
    2023/12/16 by Yaoxun Xu, Xu, Yaoxun, Hangting Chen +15 · 14 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  19. DiffuSeq-v2: Bridging Discrete and Continuous Text Spaces for Accelerated Seq2Seq Diffusion Models
    2023/10/09 by Shansan Gong, Gong, Shansan, Mukai Li +7 · 12 citations
    Biochemistry, Genetics and Molecular Biology · Computer Science · #Cancer-related molecular mechanisms research #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Speech Recognition and Synthesis #Speech and Audio Processing
  20. SCNet: Sparse Compression Network for Music Source Separation
    2024/01/24 by Weinan Tong, Jiaxu Zhu, Tong, Weinan +13 · 12 citations
    Computer Science · Engineering · #Speech and Audio Processing #Advanced Adaptive Filtering Techniques #Blind Source Separation Techniques
  21. Symbol-LLM: Towards Foundational Symbol-centric Interface For Large Language Models
    2023/11/15 by Fangzhi Xu, Xu, Fangzhi, Zhiyong Wu +15 · 11 citations
    Computer Science · Materials Science · #Topic Modeling #Advanced Text Analysis Techniques #Machine Learning in Materials Science
  22. Self-Guided Noise-Free Data Generation for Efficient Zero-Shot Learning
    2022/05/25 by Jiahui Gao, Gao, Jiahui, Renjie Pi +16 · 7 citations
    Computer Science · Medicine · #COVID-19 diagnosis using AI #Computation and Language (cs.CL) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Topic Modeling
  23. Automated Peer Reviewing in Paper SEA: Standardization, Evaluation, and Analysis
    2024/07/09 by Jianxiang Yu, Yu, Jianxiang, Ding, Zichen +20 · 12 citations
    Computer Science · #Computation and Language (cs.CL) #Digital Libraries (cs.DL) #Expert finding and Q&A systems #FOS: Computer and information sciences #Information Retrieval (cs.IR)
  24. Foundation Models for Music: A Survey
    2024/08/26 by Yinghao Ma, Ma, Yinghao, Anders Øland +81 · 14 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music Technology and Sound Studies #Sound (cs.SD) #electronic engineering #information engineering
  25. UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
    2025/09/02 by Haoming Wang, Wang, Haoming, Haoyang Zou +201 · 50 citations
    Computer Science · Psychology · Engineering · #Context-Aware Activity Recognition Systems #Social Robot Interaction and HRI #Robotics and Automated Systems
  26. Co-Speech Gesture Video Generation via Motion-Decoupled Diffusion Model
    2024/04/02 by He Xu, He, Xu, Qiaochu Huang +17 · 9 citations
    Computer Science · Engineering · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Hand Gesture Recognition Systems #Human Motion and Animation #Human-Computer Interaction (cs.HC) #Multimedia (cs.MM) #Simulation and Modeling Applications
  27. φ-Decoding: Adaptive Foresight Sampling for Balanced Inference-Time Exploration and Exploitation
    2025/03/17 by Fangzhi Xu, Hang Yan, Xu, Fangzhi +11 · 1 voice · 6 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.AI #cs.CL #cs.LG
  28. In-Context Learning with Many Demonstration Examples
    2023/02/09 by Mukai Li, Li, Mukai, Shansan Gong +11 · 6 citations
    Computer Science · #Topic Modeling #Natural Language Processing Techniques #Multimodal Machine Learning Applications
  29. SongCreator: Lyrics-based Universal Song Generation
    2024/09/09 by Shun Lei, Lei, Shun, Yixuan Zhou +15 · 11 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  30. Good for Misconceived Reasons: An Empirical Revisiting on the Need for Visual Context in Multimodal Machine Translation
    2021/05/30 by Zhiyong Wu, Lingpeng Kong, Wu, Zhiyong +7 · 4 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
  31. An Approach to Mispronunciation Detection and Diagnosis with Acoustic, Phonetic and Linguistic (APL) Embeddings
    2021/10/14 by Wenxuan Ye, Ye, Wenxuan, Shaoguang Mao +11 · 4 citations
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  32. OpenICL: An Open-Source Framework for In-context Learning
    2023/03/06 by Zhenyu Wu, YaoXiang Wang, Wu, Zhenyu +11 · 5 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
  33. Non-Autoregressive Transformer ASR with CTC-Enhanced Decoder Input
    2020/10/28 by Xingchen Song, Zhiyong Wu, Song, Xingchen +9 · 4 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
  34. MC-SpEx: Towards Effective Speaker Extraction with Multi-Scale Interfusion and Conditional Speaker Modulation
    2023/06/28 by Jun Chen, Wei Rao, Chen, Jun +13 · 5 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  35. Enhancing Speaking Styles in Conversational Text-to-Speech Synthesis with Graph-based Multi-modal Context Modeling
    2021/06/11 by Jingbei Li, Li, Jingbei, Meng Yi +11 · 3 citations
    Computer Science · #Speech and dialogue systems #Topic Modeling #Speech Recognition and Synthesis
  36. Improving the Adversarial Robustness for Speaker Verification by Self-Supervised Learning
    2021/06/01 by Haibin Wu, Xu Li, Wu, Haibin +9 · 3 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Hate Speech and Cyberbullying Detection #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  37. Adversarial Sample Detection for Speaker Verification by Neural Vocoders
    2021/07/01 by Haibin Wu, Wu, Haibin, Po‐Chun Hsu +15 · 3 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Audio and Speech Processing (eess.AS) #Digital Media Forensic Detection #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  38. Content-Dependent Fine-Grained Speaker Embedding for Zero-Shot Speaker Adaptation in Text-to-Speech Synthesis
    2022/04/03 by Yixuan Zhou, Changhe Song, Zhou, Yixuan +13 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  39. Freetalker: Controllable Speech and Text-Driven Gesture Generation Based on Diffusion Models for Enhanced Speaker Naturalness
    2024/01/07 by Sicheng Yang, Yang, Sicheng, Zunnan Xu +11 · 4 citations
    Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Face recognition and analysis #Human Motion and Animation #Human-Computer Interaction (cs.HC) #Multimedia (cs.MM) #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  40. Inter-SubNet: Speech Enhancement with Subband Interaction
    2023/05/09 by Jun Chen, Chen, Jun, Wei Rao +13 · 3 citations
    Computer Science · Engineering · #Speech and Audio Processing #Speech Recognition and Synthesis #Acoustic Wave Phenomena Research
  41. Adversarially learning disentangled speech representations for robust multi-factor voice conversion
    2021/01/30 by Jie Wang, Jingbei Li, Wang, Jie +7 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  42. Implicit Search via Discrete Diffusion: A Study on Chess
    2025/02/27 by Jiacheng Ye, Zhenyu Wu, Ye, Jiacheng +11 · 9 citations
    Computer Science · Decision Sciences · #Artificial Intelligence in Games #Reinforcement Learning in Robotics #Advanced Bandit Algorithms Research
  43. Emotion controllable speech synthesis using emotion-unlabeled dataset with the assistance of cross-domain speech emotion recognition
    2020/10/26 by Xiong Cai, Dongyang Dai, Cai, Xiong +9 · 2 citations
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Emotion and Mood Recognition #FOS: Computer and information sciences #FOS: Electrical engineering #I.2 #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  44. AudioComposer: Towards Fine-grained Audio Generation with Natural Language Descriptions
    2024/09/19 by Yuanyuan Wang, Wang, Yuanyuan, Hangting Chen +7 · 5 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  45. Focus on the Sound around You: Monaural Target Speaker Extraction via Distance and Speaker Information
    2023/06/28 by Jiuxin Lin, Peng Wang, Lin, Jiuxin +15 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  46. Genius: A Generalizable and Purely Unsupervised Self-Training Framework For Advanced Reasoning
    2025/04/11 by Fangzhi Xu, Hang Yan, Xu, Fangzhi +15 · 12 citations
    Neuroscience · Social Sciences · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Education and Critical Thinking Development #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neuroscience, Education and Cognitive Function
  47. A Controlled Study on Long Context Extension and Generalization in LLMs
    2024/09/18 by Yi Lü, Jing Nathan Yan, Lu, Yi +15 · 5 citations
    Computer Science · #Computation and Language (cs.CL) #Data Mining Algorithms and Applications #FOS: Computer and information sciences #Machine Learning (cs.LG) #Semantic Web and Ontologies
  48. MagicMan: Generative Novel View Synthesis of Humans with 3D-Aware Diffusion and Iterative Refinement
    2024/08/26 by He Xu, He, Xu, Xiaoyu Li +17 · 4 citations
    Computer Science · #Advanced Vision and Imaging #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Video Surveillance and Tracking Methods
  49. SimCalib: Graph Neural Network Calibration based on Similarity between Nodes
    2023/12/19 by Boshi Tang, Tang, Boshi, Zhiyong Wu +11 · 3 citations
    Computer Science · Neuroscience · #Advanced Graph Neural Networks #Brain Tumor Detection and Classification #FOS: Computer and information sciences #Graph Theory and Algorithms #Machine Learning (cs.LG) #Social and Information Networks (cs.SI)
  50. TrimTail: Low-Latency Streaming ASR with Simple but Effective Spectrogram-Level Length Penalty
    2022/11/01 by Xingchen Song, Di Wu, Song, Xingchen +15 · 2 citations
    Engineering · Social Sciences · #Advanced Chemical Sensor Technologies #Advanced Computing and Algorithms #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #I.2.7 #Sound (cs.SD) #Underwater Vehicles and Communication Systems #electronic engineering #information engineering
  51. Speech Representation Disentanglement with Adversarial Mutual Information Learning for One-shot Voice Conversion
    2022/08/18 by SiCheng Yang, Yang, SiCheng, Methawee Tantrawenith +19 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  52. HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
    2025/09/10 by Liyang Chen, Chen, Liyang, Tianxiang Ma +17 · 9 citations
    Computer Science · #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications #Face recognition and analysis
  53. Leveraging Chain of Thought towards Empathetic Spoken Dialogue without Corresponding Question-Answering Data
    2025/01/19 by Jingran Xie, Xie, Jingran, Shun Lei +11 · 6 citations
    Computer Science · Psychology · #Speech and dialogue systems #Innovative Teaching and Learning Methods #Intelligent Tutoring Systems and Adaptive Learning
  54. When Fuzzing Meets LLMs: Challenges and Opportunities
    2024/04/25 by Yu Jiang, Jiang, Yu, Jie Liang +19 · 3 citations
    Business, Management and Accounting · Engineering · #Artificial Intelligence (cs.AI) #Big Data and Business Intelligence #Collaboration in agile enterprises #Digital Transformation in Industry #FOS: Computer and information sciences #Software Engineering (cs.SE)
  55. StyleSpeech: Self-supervised Style Enhancing with VQ-VAE-based Pre-training for Expressive Audiobook Speech Synthesis
    2023/12/19 by Xueyuan Chen, Xi Wang, Chen, Xueyuan +11 · 3 citations
    Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  56. TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch
    2024/12/11 by Xingchen Song, Song, Xingchen, Mengtao Xing +20 · 5 citations
    Computer Science · Decision Sciences · #Context-Aware Activity Recognition Systems #Personal Information Management and User Behavior
  57. Spontaneous Style Text-to-Speech Synthesis with Controllable Spontaneous Behaviors Based on Language Models
    2024/07/18 by Weiqin Li, Peiji Yang, Li, Weiqin +13 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
  58. LeVo: High-Quality Song Generation with Multi-Preference Alignment
    2025/06/09 by Shun Lei, Yanying Xu, Lei, Shun +22 · 9 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Generative Adversarial Networks and Image Synthesis #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  59. A Discourse-level Multi-scale Prosodic Model for Fine-grained Emotion Analysis
    2023/09/21 by Xianhao Wei, Jia Jia, Wei, Xianhao +7 · 2 citations
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Emotion and Mood Recognition #FOS: Computer and information sciences #FOS: Electrical engineering #Sentiment Analysis and Opinion Mining #Sound (cs.SD) #electronic engineering #information engineering
  60. Explore 3D Dance Generation via Reward Model from Automatically-Ranked Demonstrations
    2023/12/18 by Zilin Wang, Haolin Zhuang, Wang, Zilin +15 · 2 citations
    Computer Science · Engineering · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Human Motion and Animation #Human Pose and Action Recognition #Human-Computer Interaction (cs.HC) #I.3.7 #Reinforcement Learning in Robotics
  61. VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
    2025/09/29 by Yixuan Zhou, Zhou, Yixuan, Guoqing Zeng +21 · 4 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and dialogue systems #Speech and Audio Processing
  62. How Vocabulary Sharing Facilitates Multilingualism in LLaMA?
    2023/11/15 by Fei Yuan, Yuan, Fei, Shuai Yuan +5 · 2 citations
    Computer Science · #Topic Modeling #Natural Language Processing Techniques #Text Readability and Simplification
  63. Speaker Independent and Multilingual/Mixlingual Speech-Driven Talking Head Generation Using Phonetic Posteriorgrams
    2020/06/20 by Huirong Huang, Zhiyong Wu, Huang, Huirong +21 · 1 citation
    Computer Science · #Speech and Audio Processing #Speech Recognition and Synthesis #Face recognition and analysis
  64. The Multi-speaker Multi-style Voice Cloning Challenge 2021
    2021/04/05 by Qicong Xie, Xie, Qicong, Xiaohai Tian +21 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  65. PTeacher: a Computer-Aided Personalized Pronunciation Training System with Exaggerated Audio-Visual Corrective Feedback
    2021/05/11 by Yaohua Bu, Tianyi Ma, Bu, Yaohua +27 · 1 citation
    Computer Science · Psychology · #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Phonetics and Phonology Research #Speech Recognition and Synthesis #Speech and dialogue systems
  66. AutoStyle-TTS: Retrieval-Augmented Generation based Automatic Style Matching Text-to-Speech Synthesis
    2025/04/14 by Dan Luo, Chengyuan Ma, Luo, Dan +9 · 4 citations
    Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling
  67. The Mars Orbiter Magnetometer of Tianwen-1: In-flight Performance and First Science Results
    2023/01/02 by Yuming Wang, Tielong Zhang, Wang, Yuming +37 · 1 citation
    Biochemistry, Genetics and Molecular Biology · Physics and Astronomy · #Astro and Planetary Science #Earth and Planetary Astrophysics (astro-ph.EP) #FOS: Physical sciences #Geomagnetism and Paleomagnetism Studies #Planetary Science and Exploration #Space Physics (physics.space-ph)
  68. Exploiting Audio-Visual Features with Pretrained AV-HuBERT for Multi-Modal Dysarthric Speech Reconstruction
    2024/01/31 by Xueyuan Chen, Chen, Xueyuan, Yuejiao Wang +11 · 2 citations
    Computer Science · Medicine · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Voice and Speech Disorders #electronic engineering #information engineering
  69. Speech Enhancement with Fullband-Subband Cross-Attention Network
    2022/11/10 by Jun Chen, Wei Rao, Chen, Jun +13 · 1 citation
    Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Hand Gesture Recognition Systems #Indoor and Outdoor Localization Technologies #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  70. Disentangled Speech Representation Learning for One-Shot Cross-lingual Voice Conversion Using β-VAE
    2022/10/25 by Hui Lü, Lu, Hui, Disong Wang +9 · 1 citation
    Computer Science · Medicine · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Voice and Speech Disorders #electronic engineering #information engineering
  71. Adversarial Speaker Disentanglement Using Unannotated External Data for Self-supervised Representation Based Voice Conversion
    2023/05/16 by Xintao Zhao, Shuai Wang, Zhao, Xintao +7 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  72. Context-aware Coherent Speaking Style Prediction with Hierarchical Transformers for Audiobook Speech Synthesis
    2023/04/13 by Shun Lei, Yixuan Zhou, Lei, Shun +9 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  73. SnakeGAN: A Universal Vocoder Leveraging DDSP Prior Knowledge and Periodic Inductive Bias
    2023/09/14 by Sipan Li, Songxiang Liu, Li, Sipan +13 · 1 citation
    Computer Science · #Speech and Audio Processing #Music and Audio Processing #Music Technology and Sound Studies
  74. EMO: Earth Mover Distance Optimization for Auto-Regressive Language Modeling
    2023/10/07 by Siyu Ren, Ren, Siyu, Zhiyong Wu +3 · 1 citation
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
  75. Interactive Evolution: A Neural-Symbolic Self-Training Framework For Large Language Models
    2024/06/17 by Fangzhi Xu, Qiushi Sun, Xu, Fangzhi +9 · 2 citations
    Computer Science · #Topic Modeling
  76. Towards Improving the Expressiveness of Singing Voice Synthesis with BERT Derived Semantic Information
    2023/08/31 by Shaohuan Zhou, Zhou, Shaohuan, Shun Lei +13 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  77. Rhythmic Foley: A Framework For Seamless Audio-Visual Alignment In Video-to-Audio Synthesis
    2024/09/13 by Zhiqi Huang, Dan Luo, Huang, Zhiqi +9 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  78. MuCodec: Ultra Low-Bitrate Music Codec
    2024/09/20 by Yaoxun Xu, Xu, Yaoxun, Hangting Chen +12 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  79. CALM: Contrastive Cross-modal Speaking Style Modeling for Expressive Text-to-Speech Synthesis
    2023/08/30 by Yi Meng, Xiang Li, Meng, Yi +15 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  80. An End-to-End Approach for Chord-Conditioned Song Generation
    2024/09/10 by Shuochen Gao, Shun Lei, Gao, Shuochen +15 · 2 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Cellular Automata and Applications #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  81. RobustSVC: HuBERT-based Melody Extractor and Adversarial Learning for Robust Singing Voice Conversion
    2024/09/10 by Wei Chen, Chen, Wei, Xintao Zhao +9 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  82. OS-MAP: How Far Can Computer-Using Agents Go in Breadth and Depth?
    2025/07/25 by X. R. Chen, Chen, Xuetian, Yuquan Chen +27 · 4 citations
    Decision Sciences · Psychology · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human-Automation Interaction and Safety #Human-Computer Interaction (cs.HC) #Personal Information Management and User Behavior #Social Robot Interaction and HRI
  83. AdaMesh: Personalized Facial Expressions and Head Poses for Adaptive Speech-Driven 3D Facial Animation
    2023/10/11 by Liyang Chen, Weihong Bao, Chen, Liyang +12 · 1 citation
    Computer Science · Engineering · Medicine · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Face recognition and analysis #Facial Nerve Paralysis Treatment and Research #Human Motion and Animation #Multimedia (cs.MM)
  84. Singing Voice Conversion with Accompaniment Using Self-Supervised Representation-Based Melody Features
    2025/02/07 by Wei Chen, Binzhu Sha, Chen, Wei +9 · 1 citation
    Computer Science · #Speech and Audio Processing #Music and Audio Processing #Speech Recognition and Synthesis
  85. OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
    2026/07/30 by Qiushi Sun, Kanzhi Cheng, Yian Wang +20 · 1 voice
    Computer Science · #cs.AI #cs.CL #cs.CV
  86. Qwen-Music Technical Report
    2026/07/27 by Jin Xu, Kangdi Wang, Ruibin Yuan +24
    #cs.SD