Zhiyong Wu
- A Survey on In-context Learning
2022/12/31 by Qingxiu Dong, Lei Li, Dong, Qingxiu +23 · 4 voices · 185 citations
Computer Science · #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
2024/12/06 by Zhe Chen, Weiyun Wang, Chen, Zhe +86 · 3 voices · 704 citations
Computer Science · Decision Sciences · #Topic Modeling #Scientific Computing and Data Management #Machine Learning and Data Classification
- SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
2024/01/17 by Kanzhi Cheng, Qiushi Sun, Cheng, Kanzhi +11 · 182 citations
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Interactive and Immersive Displays #Virtual Reality Applications and Impacts
- DiffuSeq: Sequence to Sequence Text Generation with Diffusion Models
2022/10/17 by Shansan Gong, Mukai Li, Gong, Shansan +7 · 97 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Music and Audio Processing #Natural Language Processing Techniques #Topic Modeling
- OS-ATLAS: A Foundation Action Model for Generalist GUI Agents
2024/10/30 by Zhiyong Wu, Wu, Zhiyong, Zhenyu Wu +19 · 150 citations
Computer Science · Psychology · #Reinforcement Learning in Robotics #Multi-Agent Systems and Negotiation #Social Robot Interaction and HRI
- Can We Edit Factual Knowledge by In-Context Learning?
2023/05/22 by Ce Zheng, Lei Li, Zheng, Ce +11 · 68 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning in Healthcare #Natural Language Processing Techniques #Topic Modeling
- OS-Copilot: Towards Generalist Computer Agents with Self-Improvement
2024/02/12 by Zhiyong Wu, Chengcheng Han, Wu, Zhiyong +13 · 55 citations
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Multi-Agent Systems and Negotiation #Semantic Web and Ontologies
- Self-Adaptive In-Context Learning: An Information Compression Perspective for In-Context Example Selection and Ordering
2022/12/20 by Zhiyong Wu, Wu, Zhiyong, Yaoxiang Wang +5 · 37 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Music and Audio Processing #Text and Document Classification Technologies
- Compositional Exemplars for In-context Learning
2023/02/11 by Jiacheng Ye, Zhiyong Wu, Ye, Jiacheng +7 · 32 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Multimodal Machine Learning Applications
- OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis
2024/12/27 by Qiushi Sun, Sun, Qiushi, Kanzhi Cheng +27 · 58 citations
Computer Science · Engineering · Psychology · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Context-Aware Activity Recognition Systems #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Robotics and Automated Systems #Social Robot Interaction and HRI
- ZeroGen: Efficient Zero-shot Learning via Dataset Generation
2022/02/16 by Jiacheng Ye, Jiahui Gao, Ye, Jiacheng +13 · 20 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
- DiffuseStyleGesture: Stylized Audio-Driven Co-Speech Gesture Generation with Diffusion Models
2023/05/08 by Sicheng Yang, Zhiyong Wu, Yang, Sicheng +13 · 21 citations
Engineering · Computer Science · #Human Motion and Animation #Hand Gesture Recognition Systems #Human Pose and Action Recognition
- Corex: Pushing the Boundaries of Complex Reasoning through Multi-Model Collaboration
2023/09/30 by Qiushi Sun, Zhangyue Yin, Sun, Qiushi +9 · 19 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling
- TDAG: A Multi-Agent Framework based on Dynamic Task Decomposition and Agent Generation
2024/02/15 by Yaoxiang Wang, Zhiyong Wu, Wang, Yaoxiang +5 · 18 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multi-Agent Systems and Negotiation
- MFA-Conformer: Multi-scale Feature Aggregation Conformer for Automatic Speaker Verification
2022/03/29 by Yang Zhang, Zhiqiang Lv, Zhang, Yang +13 · 11 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- AV-SepFormer: Cross-Attention SepFormer for Audio-Visual Target Speaker Extraction
2023/06/25 by Jiuxin Lin, Xinyu Cai, Lin, Jiuxin +17 · 14 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Perturbed Masking: Parameter-free Probing for Analyzing and Interpreting BERT
2020/04/30 by Zhiyong Wu, Wu, Zhiyong, Yun Chen +4 · 8 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
- SECap: Speech Emotion Captioning with Large Language Model
2023/12/16 by Yaoxun Xu, Xu, Yaoxun, Hangting Chen +15 · 14 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- DiffuSeq-v2: Bridging Discrete and Continuous Text Spaces for Accelerated Seq2Seq Diffusion Models
2023/10/09 by Shansan Gong, Gong, Shansan, Mukai Li +7 · 12 citations
Biochemistry, Genetics and Molecular Biology · Computer Science · #Cancer-related molecular mechanisms research #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Speech Recognition and Synthesis #Speech and Audio Processing
- SCNet: Sparse Compression Network for Music Source Separation
2024/01/24 by Weinan Tong, Jiaxu Zhu, Tong, Weinan +13 · 12 citations
Computer Science · Engineering · #Speech and Audio Processing #Advanced Adaptive Filtering Techniques #Blind Source Separation Techniques
- Symbol-LLM: Towards Foundational Symbol-centric Interface For Large Language Models
2023/11/15 by Fangzhi Xu, Xu, Fangzhi, Zhiyong Wu +15 · 11 citations
Computer Science · Materials Science · #Topic Modeling #Advanced Text Analysis Techniques #Machine Learning in Materials Science
- Self-Guided Noise-Free Data Generation for Efficient Zero-Shot Learning
2022/05/25 by Jiahui Gao, Gao, Jiahui, Renjie Pi +16 · 7 citations
Computer Science · Medicine · #COVID-19 diagnosis using AI #Computation and Language (cs.CL) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Topic Modeling
- Automated Peer Reviewing in Paper SEA: Standardization, Evaluation, and Analysis
2024/07/09 by Jianxiang Yu, Yu, Jianxiang, Ding, Zichen +20 · 12 citations
Computer Science · #Computation and Language (cs.CL) #Digital Libraries (cs.DL) #Expert finding and Q&A systems #FOS: Computer and information sciences #Information Retrieval (cs.IR)
- Foundation Models for Music: A Survey
2024/08/26 by Yinghao Ma, Ma, Yinghao, Anders Øland +81 · 14 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music Technology and Sound Studies #Sound (cs.SD) #electronic engineering #information engineering
- UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
2025/09/02 by Haoming Wang, Wang, Haoming, Haoyang Zou +201 · 50 citations
Computer Science · Psychology · Engineering · #Context-Aware Activity Recognition Systems #Social Robot Interaction and HRI #Robotics and Automated Systems
- Co-Speech Gesture Video Generation via Motion-Decoupled Diffusion Model
2024/04/02 by He Xu, He, Xu, Qiaochu Huang +17 · 9 citations
Computer Science · Engineering · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Hand Gesture Recognition Systems #Human Motion and Animation #Human-Computer Interaction (cs.HC) #Multimedia (cs.MM) #Simulation and Modeling Applications
- φ-Decoding: Adaptive Foresight Sampling for Balanced Inference-Time Exploration and Exploitation
2025/03/17 by Fangzhi Xu, Hang Yan, Xu, Fangzhi +11 · 1 voice · 6 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.AI #cs.CL #cs.LG
- In-Context Learning with Many Demonstration Examples
2023/02/09 by Mukai Li, Li, Mukai, Shansan Gong +11 · 6 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Multimodal Machine Learning Applications
- SongCreator: Lyrics-based Universal Song Generation
2024/09/09 by Shun Lei, Lei, Shun, Yixuan Zhou +15 · 11 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
- Good for Misconceived Reasons: An Empirical Revisiting on the Need for Visual Context in Multimodal Machine Translation
2021/05/30 by Zhiyong Wu, Lingpeng Kong, Wu, Zhiyong +7 · 4 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
- An Approach to Mispronunciation Detection and Diagnosis with Acoustic, Phonetic and Linguistic (APL) Embeddings
2021/10/14 by Wenxuan Ye, Ye, Wenxuan, Shaoguang Mao +11 · 4 citations
Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- OpenICL: An Open-Source Framework for In-context Learning
2023/03/06 by Zhenyu Wu, YaoXiang Wang, Wu, Zhenyu +11 · 5 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
- Non-Autoregressive Transformer ASR with CTC-Enhanced Decoder Input
2020/10/28 by Xingchen Song, Zhiyong Wu, Song, Xingchen +9 · 4 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
- MC-SpEx: Towards Effective Speaker Extraction with Multi-Scale Interfusion and Conditional Speaker Modulation
2023/06/28 by Jun Chen, Wei Rao, Chen, Jun +13 · 5 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Enhancing Speaking Styles in Conversational Text-to-Speech Synthesis with Graph-based Multi-modal Context Modeling
2021/06/11 by Jingbei Li, Li, Jingbei, Meng Yi +11 · 3 citations
Computer Science · #Speech and dialogue systems #Topic Modeling #Speech Recognition and Synthesis
- Improving the Adversarial Robustness for Speaker Verification by Self-Supervised Learning
2021/06/01 by Haibin Wu, Xu Li, Wu, Haibin +9 · 3 citations
Computer Science · #Adversarial Robustness in Machine Learning #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Hate Speech and Cyberbullying Detection #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Adversarial Sample Detection for Speaker Verification by Neural Vocoders
2021/07/01 by Haibin Wu, Wu, Haibin, Po‐Chun Hsu +15 · 3 citations
Computer Science · #Adversarial Robustness in Machine Learning #Audio and Speech Processing (eess.AS) #Digital Media Forensic Detection #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Content-Dependent Fine-Grained Speaker Embedding for Zero-Shot Speaker Adaptation in Text-to-Speech Synthesis
2022/04/03 by Yixuan Zhou, Changhe Song, Zhou, Yixuan +13 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Freetalker: Controllable Speech and Text-Driven Gesture Generation Based on Diffusion Models for Enhanced Speaker Naturalness
2024/01/07 by Sicheng Yang, Yang, Sicheng, Zunnan Xu +11 · 4 citations
Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Face recognition and analysis #Human Motion and Animation #Human-Computer Interaction (cs.HC) #Multimedia (cs.MM) #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- Inter-SubNet: Speech Enhancement with Subband Interaction
2023/05/09 by Jun Chen, Chen, Jun, Wei Rao +13 · 3 citations
Computer Science · Engineering · #Speech and Audio Processing #Speech Recognition and Synthesis #Acoustic Wave Phenomena Research
- Adversarially learning disentangled speech representations for robust multi-factor voice conversion
2021/01/30 by Jie Wang, Jingbei Li, Wang, Jie +7 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Implicit Search via Discrete Diffusion: A Study on Chess
2025/02/27 by Jiacheng Ye, Zhenyu Wu, Ye, Jiacheng +11 · 9 citations
Computer Science · Decision Sciences · #Artificial Intelligence in Games #Reinforcement Learning in Robotics #Advanced Bandit Algorithms Research
- Emotion controllable speech synthesis using emotion-unlabeled dataset with the assistance of cross-domain speech emotion recognition
2020/10/26 by Xiong Cai, Dongyang Dai, Cai, Xiong +9 · 2 citations
Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Emotion and Mood Recognition #FOS: Computer and information sciences #FOS: Electrical engineering #I.2 #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- AudioComposer: Towards Fine-grained Audio Generation with Natural Language Descriptions
2024/09/19 by Yuanyuan Wang, Wang, Yuanyuan, Hangting Chen +7 · 5 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- Focus on the Sound around You: Monaural Target Speaker Extraction via Distance and Speaker Information
2023/06/28 by Jiuxin Lin, Peng Wang, Lin, Jiuxin +15 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Genius: A Generalizable and Purely Unsupervised Self-Training Framework For Advanced Reasoning
2025/04/11 by Fangzhi Xu, Hang Yan, Xu, Fangzhi +15 · 12 citations
Neuroscience · Social Sciences · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Education and Critical Thinking Development #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neuroscience, Education and Cognitive Function
- A Controlled Study on Long Context Extension and Generalization in LLMs
2024/09/18 by Yi Lü, Jing Nathan Yan, Lu, Yi +15 · 5 citations
Computer Science · #Computation and Language (cs.CL) #Data Mining Algorithms and Applications #FOS: Computer and information sciences #Machine Learning (cs.LG) #Semantic Web and Ontologies
- MagicMan: Generative Novel View Synthesis of Humans with 3D-Aware Diffusion and Iterative Refinement
2024/08/26 by He Xu, He, Xu, Xiaoyu Li +17 · 4 citations
Computer Science · #Advanced Vision and Imaging #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Video Surveillance and Tracking Methods
- SimCalib: Graph Neural Network Calibration based on Similarity between Nodes
2023/12/19 by Boshi Tang, Tang, Boshi, Zhiyong Wu +11 · 3 citations
Computer Science · Neuroscience · #Advanced Graph Neural Networks #Brain Tumor Detection and Classification #FOS: Computer and information sciences #Graph Theory and Algorithms #Machine Learning (cs.LG) #Social and Information Networks (cs.SI)
- TrimTail: Low-Latency Streaming ASR with Simple but Effective Spectrogram-Level Length Penalty
2022/11/01 by Xingchen Song, Di Wu, Song, Xingchen +15 · 2 citations
Engineering · Social Sciences · #Advanced Chemical Sensor Technologies #Advanced Computing and Algorithms #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #I.2.7 #Sound (cs.SD) #Underwater Vehicles and Communication Systems #electronic engineering #information engineering
- Speech Representation Disentanglement with Adversarial Mutual Information Learning for One-shot Voice Conversion
2022/08/18 by SiCheng Yang, Yang, SiCheng, Methawee Tantrawenith +19 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
2025/09/10 by Liyang Chen, Chen, Liyang, Tianxiang Ma +17 · 9 citations
Computer Science · #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications #Face recognition and analysis
- Leveraging Chain of Thought towards Empathetic Spoken Dialogue without Corresponding Question-Answering Data
2025/01/19 by Jingran Xie, Xie, Jingran, Shun Lei +11 · 6 citations
Computer Science · Psychology · #Speech and dialogue systems #Innovative Teaching and Learning Methods #Intelligent Tutoring Systems and Adaptive Learning
- When Fuzzing Meets LLMs: Challenges and Opportunities
2024/04/25 by Yu Jiang, Jiang, Yu, Jie Liang +19 · 3 citations
Business, Management and Accounting · Engineering · #Artificial Intelligence (cs.AI) #Big Data and Business Intelligence #Collaboration in agile enterprises #Digital Transformation in Industry #FOS: Computer and information sciences #Software Engineering (cs.SE)
- StyleSpeech: Self-supervised Style Enhancing with VQ-VAE-based Pre-training for Expressive Audiobook Speech Synthesis
2023/12/19 by Xueyuan Chen, Xi Wang, Chen, Xueyuan +11 · 3 citations
Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch
2024/12/11 by Xingchen Song, Song, Xingchen, Mengtao Xing +20 · 5 citations
Computer Science · Decision Sciences · #Context-Aware Activity Recognition Systems #Personal Information Management and User Behavior
- Spontaneous Style Text-to-Speech Synthesis with Controllable Spontaneous Behaviors Based on Language Models
2024/07/18 by Weiqin Li, Peiji Yang, Li, Weiqin +13 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
- LeVo: High-Quality Song Generation with Multi-Preference Alignment
2025/06/09 by Shun Lei, Yanying Xu, Lei, Shun +22 · 9 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Generative Adversarial Networks and Image Synthesis #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
- A Discourse-level Multi-scale Prosodic Model for Fine-grained Emotion Analysis
2023/09/21 by Xianhao Wei, Jia Jia, Wei, Xianhao +7 · 2 citations
Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Emotion and Mood Recognition #FOS: Computer and information sciences #FOS: Electrical engineering #Sentiment Analysis and Opinion Mining #Sound (cs.SD) #electronic engineering #information engineering
- Explore 3D Dance Generation via Reward Model from Automatically-Ranked Demonstrations
2023/12/18 by Zilin Wang, Haolin Zhuang, Wang, Zilin +15 · 2 citations
Computer Science · Engineering · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Human Motion and Animation #Human Pose and Action Recognition #Human-Computer Interaction (cs.HC) #I.3.7 #Reinforcement Learning in Robotics
- VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
2025/09/29 by Yixuan Zhou, Zhou, Yixuan, Guoqing Zeng +21 · 4 citations
Computer Science · #Speech Recognition and Synthesis #Speech and dialogue systems #Speech and Audio Processing
- How Vocabulary Sharing Facilitates Multilingualism in LLaMA?
2023/11/15 by Fei Yuan, Yuan, Fei, Shuai Yuan +5 · 2 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Text Readability and Simplification
- Speaker Independent and Multilingual/Mixlingual Speech-Driven Talking Head Generation Using Phonetic Posteriorgrams
2020/06/20 by Huirong Huang, Zhiyong Wu, Huang, Huirong +21 · 1 citation
Computer Science · #Speech and Audio Processing #Speech Recognition and Synthesis #Face recognition and analysis
- The Multi-speaker Multi-style Voice Cloning Challenge 2021
2021/04/05 by Qicong Xie, Xie, Qicong, Xiaohai Tian +21 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- PTeacher: a Computer-Aided Personalized Pronunciation Training System with Exaggerated Audio-Visual Corrective Feedback
2021/05/11 by Yaohua Bu, Tianyi Ma, Bu, Yaohua +27 · 1 citation
Computer Science · Psychology · #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Phonetics and Phonology Research #Speech Recognition and Synthesis #Speech and dialogue systems
- AutoStyle-TTS: Retrieval-Augmented Generation based Automatic Style Matching Text-to-Speech Synthesis
2025/04/14 by Dan Luo, Chengyuan Ma, Luo, Dan +9 · 4 citations
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling
- The Mars Orbiter Magnetometer of Tianwen-1: In-flight Performance and First Science Results
2023/01/02 by Yuming Wang, Tielong Zhang, Wang, Yuming +37 · 1 citation
Biochemistry, Genetics and Molecular Biology · Physics and Astronomy · #Astro and Planetary Science #Earth and Planetary Astrophysics (astro-ph.EP) #FOS: Physical sciences #Geomagnetism and Paleomagnetism Studies #Planetary Science and Exploration #Space Physics (physics.space-ph)
- Exploiting Audio-Visual Features with Pretrained AV-HuBERT for Multi-Modal Dysarthric Speech Reconstruction
2024/01/31 by Xueyuan Chen, Chen, Xueyuan, Yuejiao Wang +11 · 2 citations
Computer Science · Medicine · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Voice and Speech Disorders #electronic engineering #information engineering
- Speech Enhancement with Fullband-Subband Cross-Attention Network
2022/11/10 by Jun Chen, Wei Rao, Chen, Jun +13 · 1 citation
Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Hand Gesture Recognition Systems #Indoor and Outdoor Localization Technologies #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- Disentangled Speech Representation Learning for One-Shot Cross-lingual Voice Conversion Using β-VAE
2022/10/25 by Hui Lü, Lu, Hui, Disong Wang +9 · 1 citation
Computer Science · Medicine · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Voice and Speech Disorders #electronic engineering #information engineering
- Adversarial Speaker Disentanglement Using Unannotated External Data for Self-supervised Representation Based Voice Conversion
2023/05/16 by Xintao Zhao, Shuai Wang, Zhao, Xintao +7 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Context-aware Coherent Speaking Style Prediction with Hierarchical Transformers for Audiobook Speech Synthesis
2023/04/13 by Shun Lei, Yixuan Zhou, Lei, Shun +9 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- SnakeGAN: A Universal Vocoder Leveraging DDSP Prior Knowledge and Periodic Inductive Bias
2023/09/14 by Sipan Li, Songxiang Liu, Li, Sipan +13 · 1 citation
Computer Science · #Speech and Audio Processing #Music and Audio Processing #Music Technology and Sound Studies
- EMO: Earth Mover Distance Optimization for Auto-Regressive Language Modeling
2023/10/07 by Siyu Ren, Ren, Siyu, Zhiyong Wu +3 · 1 citation
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
- Interactive Evolution: A Neural-Symbolic Self-Training Framework For Large Language Models
2024/06/17 by Fangzhi Xu, Qiushi Sun, Xu, Fangzhi +9 · 2 citations
Computer Science · #Topic Modeling
- Towards Improving the Expressiveness of Singing Voice Synthesis with BERT Derived Semantic Information
2023/08/31 by Shaohuan Zhou, Zhou, Shaohuan, Shun Lei +13 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Rhythmic Foley: A Framework For Seamless Audio-Visual Alignment In Video-to-Audio Synthesis
2024/09/13 by Zhiqi Huang, Dan Luo, Huang, Zhiqi +9 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
- MuCodec: Ultra Low-Bitrate Music Codec
2024/09/20 by Yaoxun Xu, Xu, Yaoxun, Hangting Chen +12 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- CALM: Contrastive Cross-modal Speaking Style Modeling for Expressive Text-to-Speech Synthesis
2023/08/30 by Yi Meng, Xiang Li, Meng, Yi +15 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- An End-to-End Approach for Chord-Conditioned Song Generation
2024/09/10 by Shuochen Gao, Shun Lei, Gao, Shuochen +15 · 2 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Cellular Automata and Applications #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
- RobustSVC: HuBERT-based Melody Extractor and Adversarial Learning for Robust Singing Voice Conversion
2024/09/10 by Wei Chen, Chen, Wei, Xintao Zhao +9 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- OS-MAP: How Far Can Computer-Using Agents Go in Breadth and Depth?
2025/07/25 by X. R. Chen, Chen, Xuetian, Yuquan Chen +27 · 4 citations
Decision Sciences · Psychology · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human-Automation Interaction and Safety #Human-Computer Interaction (cs.HC) #Personal Information Management and User Behavior #Social Robot Interaction and HRI
- AdaMesh: Personalized Facial Expressions and Head Poses for Adaptive Speech-Driven 3D Facial Animation
2023/10/11 by Liyang Chen, Weihong Bao, Chen, Liyang +12 · 1 citation
Computer Science · Engineering · Medicine · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Face recognition and analysis #Facial Nerve Paralysis Treatment and Research #Human Motion and Animation #Multimedia (cs.MM)
- Singing Voice Conversion with Accompaniment Using Self-Supervised Representation-Based Melody Features
2025/02/07 by Wei Chen, Binzhu Sha, Chen, Wei +9 · 1 citation
Computer Science · #Speech and Audio Processing #Music and Audio Processing #Speech Recognition and Synthesis
- OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
2026/07/30 by Qiushi Sun, Kanzhi Cheng, Yian Wang +20 · 1 voice
Computer Science · #cs.AI #cs.CL #cs.CV
- Qwen-Music Technical Report
2026/07/27 by Jin Xu, Kangdi Wang, Ruibin Yuan +24
#cs.SD