vix.ing · top · new · best · stats · spec
  1. Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering
    2026/07/22 by Junyu Dai, Xinyue Fan, Weiqin Li +14 · 1 citation
    Computer Science · Engineering · #cs.AI #cs.SD #eess.AS
  2. Time-Frequency Consistency Learning for Robust Speech Deepfake Detection
    2026/07/20 by Jun Xue, Zhuolin Yi, Yanzhen Ren +6 · 1 voice
    #cs.SD #cs.AI
  3. Audio-Native Speech Recognition with a Frozen Discrete-Diffusion Language Model
    2026/07/14 by Harsha Vardhan Khurdula, Abhinav Kumar Singh, Yoeven D Khemlani +1 · 1 voice
    #cs.AI #cs.SD
  4. MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation
    2026/07/15 by Xiaohan Zhang, Yuqing Wen, Junlin Chen +9 · 1 citation
    #cs.CV #cs.SD
  5. Decomposer: Learning to Decompile Symbolic Music to Programs
    2026/07/02 by Yewon Kim, Apurva Gandhi, David Chung +2 · 2 voices
    Computer Science · #cs.LG #cs.AI #cs.SD
  6. MambAdapter: Lightweight Mamba-Based Adapters for Parameter-Efficient Transfer Learning in Speech and Audio
    2026/06/14 by Salman Hussain Ali, Umberto Cappellazzo, Mirco Ravanelli · 1 voice
    Engineering · Computer Science · #eess.AS #cs.SD
  7. MelT: A Portable, Single-GEMM Mel Audio Frontend via Non-Uniform DFT with Measured Latency and Energy Gains on GPUs
    2026/05/31 by Augusto Camargo, Marcelo Finger · 1 voice
    Computer Science · #cs.SD
  8. Continual Speaker Identity Unlearning with Minimal Interference
    2026/05/25 by Jinju Kim, Yunsung Kang, Gyeong-Moon Park +1 · 1 voice
    Computer Science · #cs.SD #cs.AI
  9. Stable Audio 3
    2026/05/18 by Zach Evans, Julian D. Parker, Matthew Rice +4 · 8 voices
    #cs.SD #cs.AI
  10. Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection
    2026/04/16 by Meng Chen, Kun Wang, Li Lu +2 · 4 voices · 2 citations
    #cs.CR #cs.AI #cs.SD
  11. Benchmarking Language Modeling for Lossless Compression of Full-Fidelity Audio
    2026/03/09 by Phillip Long, Zachary Novack, Chris Donahue · 1 voice
    Computer Science · Engineering · #cs.SD #cs.AI #cs.LG #eess.AS
  12. ACE-Step 1.5: Pushing the Boundaries of Open-Source Music Generation
    2026/01/31 by Junmin Gong, Yulin Song, Wenxiao Zhao +4 · 1 voice · 1 citation
    #cs.SD
  13. Qwen3-ASR Technical Report
    2026/01/29 by Xian Shi, Xiong Wang, Zhifang Guo +10 · 1 voice · 2 citations
    #cs.CL #cs.SD #eess.AS
  14. Embryonic Exposure to VPA Influences Chick Vocalisations: A Computational Study
    2026/01/18 by Antonella M. C. Torrisi, Inês Nolasco, Paola Sgadò +2 · 1 voice · 1 citation
    #cs.SD
  15. Elastic overtones: an equal temperament 12 tone music system with "perfect" fifths
    2026/01/12 by X. Hernandez, Luis Nasser, Pablo Garcia-Valenzuela · 2 voices
    #physics.soc-ph #cs.SD #eess.AS #physics.pop-ph
  16. PromptReverb: Multimodal Room Impulse Response Generation Through Latent Rectified Flow Matching
    2025/10/25 by Ali Vosoughi, Vosoughi, Ali, Yongyi Zang +9 · 1 voice · 1 citation
    #cs.SD #cs.AI
  17. WildElder: A Chinese Elderly Speech Dataset from the Wild with Fine-Grained Manual Annotations
    2025/10/10 by Hui Wang, Wang, Hui, Jiaming Zhou +7 · 1 citation
    #cs.SD #eess.AS
  18. Invisible Ears at Your Fingertips: Acoustic Eavesdropping via Mouse Sensors
    2025/09/16 by Mohamad Fakih, Fakih, Mohamad, Rahul Dharmaji +7 · 10 voices
    Computer Science · #Advanced Malware Detection Techniques #Interactive and Immersive Displays #User Authentication and Security Systems #cs.CR #cs.SD
  19. Flavors of Moonshine: Tiny Specialized ASR Models for Edge Devices
    2025/09/02 by Evan King, Adam Sabra, King, Evan +8 · 1 voice
    Computer Science · #Face recognition and analysis #Speech Recognition and Synthesis #Speech and Audio Processing #cs.CL #cs.LG #cs.SD
  20. Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning
    2025/08/27 by Abdullah Abdelfattah, Mahmoud I. Khalil, Abdelfattah, Abdullah +4 · 1 voice
    Arts and Humanities · Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Language, Linguistics, Cultural Analysis #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #cs.AI #cs.CL #cs.LG #cs.SD #eess.AS #electronic engineering #information engineering
  21. Foundation Models for Bioacoustics -- a Comparative Review
    2025/08/02 by Raphael Schwinger, Paria Vali Zadeh, Schwinger, Raphael +11 · 1 voice · 2 citations
    #cs.SD #cs.LG #eess.AS #q-bio.QM
  22. Combolutional Neural Networks
    2025/07/28 by Cameron Churchwell, Churchwell, Cameron, Minje Kim +3 · 1 voice
    Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #cs.LG #cs.SD #eess.AS #electronic engineering #information engineering
  23. StreamFlow: Streaming Flow Matching with Block-wise Guided Attention Mask for Speech Token Decoding
    2025/06/30 by Dake Guo, Jixun Yao, Guo, Dake +7 · 1 voice
    Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #cs.SD #eess.AS #electronic engineering #information engineering
  24. A Fourier Explanation of AI-music Artifacts
    2025/06/23 by Darius Afchar, Gabriel Meseguer-Brocal, Afchar, Darius +5 · 1 voice · 5 citations
    Computer Science · #FOS: Computer and information sciences #Music Technology and Sound Studies #Sound (cs.SD) #cs.SD
  25. Refining music sample identification with a self-supervised graph neural network
    2025/06/17 by Aditya Bhattacharjee, Bhattacharjee, Aditya, Ivan Meresman Higgs +6 · 1 voice · 1 citation
    Computer Science · #Image Processing and 3D Reconstruction #Music and Audio Processing #Neural Networks and Applications #cs.AI #cs.IR #cs.SD
  26. Addressing Pitfalls in Auditing Practices of Automatic Speech Recognition Technologies: A Case Study of People with Aphasia
    2025/06/10 by Katelyn Xiaoying Mei, Mei, Katelyn Xiaoying, Anna Seo Gyeong Choi +7 · 1 voice · 1 citation
    #cs.CY #cs.CL #cs.SD #eess.AS
  27. Audio synthesizer inversion in symmetric parameter spaces with approximately equivariant flow matching
    2025/06/08 by Ben Hayes, Ben J. Hayes, Charalampos Saitis +4 · 1 voice · 4 citations
    Computer Science · #Music Technology and Sound Studies #Music and Audio Processing #Speech and Audio Processing #cs.LG #cs.SD #eess.AS #eess.SP
  28. Can we reconstruct a dysarthric voice with the large speech model Parler TTS?
    2025/06/04 by Ariadna Sanchez, Sanchez, Ariadna, Simon King +1 · 1 voice · 1 citation
    Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #cs.CL #cs.SD #eess.AS #electronic engineering #information engineering
  29. Brain-tuned Speech Models Better Reflect Speech Processing Stages in the Brain
    2025/06/04 by Omer Moussa, Mariya Toneva, Moussa, Omer +1 · 1 voice · 3 citations
    Biochemistry, Genetics and Molecular Biology · Computer Science · Engineering · Neuroscience · Psychology · #Neural dynamics and brain function #Neurobiology of Language and Bilingualism #Phonetics and Phonology Research #cs.CL #cs.SD #eess.AS #q-bio.NC
  30. SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline
    2025/05/25 by Helin Wang, Wang, Helin, Jiarui Hai +17 · 1 voice · 4 citations
    Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #cs.AI #cs.SD #eess.AS #electronic engineering #information engineering

more