vix.ing · top · new · best · stats · spec

Jonathan Le Roux

  1. SDR - half-baked or well done?
    2018/11/06 by Jonathan Le Roux, Scott Wisdom, Roux, Jonathan Le +4 · 102 citations
    Computer Science · Neuroscience · Engineering · #Speech and Audio Processing #Hearing Loss and Rehabilitation #Advanced Adaptive Filtering Techniques
  2. Deep clustering: Discriminative embeddings for segmentation and separation
    2015/08/18 by John R. Hershey, Hershey, John R., Zhuo Chen +5 · 26 citations
    Computer Science · #Speech and Audio Processing #Music and Audio Processing #Speech Recognition and Synthesis
  3. Full-Capacity Unitary Recurrent Neural Networks
    2016/10/31 by Scott Wisdom, Thomas A. Powers, Wisdom, Scott +7 · 25 citations
    Computer Science · #Blind Source Separation Techniques #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Neural Networks and Applications #Neural and Evolutionary Computing (cs.NE)
  4. Deep Unfolding: Model-Based Inspiration of Novel Deep Architectures
    2014/09/09 by John R. Hershey, Hershey, John R., Jonathan Le Roux +3 · 15 citations
    Computer Science · #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Neural and Evolutionary Computing (cs.NE) #Speech Recognition and Synthesis #Speech and Audio Processing
  5. TF-Locoformer: Transformer with Local Modeling by Convolution for Speech Separation and Enhancement
    2024/08/06 by Kohei Saijo, Gordon Wichern, Saijo, Kohei +7 · 11 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  6. End-to-End Multi-speaker Speech Recognition with Transformer
    2020/02/10 by Xuankai Chang, Wangyou Zhang, Chang, Xuankai +7 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  7. Cold Diffusion for Speech Enhancement
    2022/11/04 by Hao Yen, François G. Germain, Yen, Hao +5 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  8. Improving Audio Captioning Models with Fine-grained Audio Features, Text Embedding Supervision, and LLM Mix-up Augmentation
    2023/09/29 by Shih-Lun Wu, Xuankai Chang, Wu, Shih-Lun +11 · 3 citations
    Computer Science · Arts and Humanities · #Music and Audio Processing #Subtitles and Audiovisual Media #Speech Recognition and Synthesis
  9. Why does music source separation benefit from cacophony?
    2024/02/28 by Chang-Bin Jeon, Gordon Wichern, Jeon, Chang-Bin +5 · 3 citations
    Neuroscience · Computer Science · #Hearing Loss and Rehabilitation #Music Technology and Sound Studies #Neuroscience and Music Perception
  10. (2.5+1)D Spatio-Temporal Scene Graphs for Video Question Answering
    2022/02/18 by Anoop Cherian, Cherian, Anoop, Chiori Hori +5 · 2 citations
    Computer Science · #Multimodal Machine Learning Applications #Human Pose and Action Recognition #Advanced Image and Video Retrieval Techniques
  11. The Cocktail Fork Problem: Three-Stem Audio Separation for Real-World Soundtracks
    2021/10/19 by Darius Petermann, Gordon Wichern, Petermann, Darius +5 · 2 citations
    Computer Science · #Speech and Audio Processing #Music and Audio Processing #Blind Source Separation Techniques
  12. NeuroHeed+: Improving Neuro-steered Speaker Extraction with Joint Auditory Attention Detection
    2023/12/12 by Zexu Pan, Gordon Wichern, Pan, Zexu +7 · 2 citations
    Computer Science · Neuroscience · #Audio and Speech Processing (eess.AS) #Blind Source Separation Techniques #EEG and Brain-Computer Interfaces #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  13. Unsupervised Speaker Adaptation using Attention-based Speaker Memory for\n End-to-End ASR
    2020/02/14 by Leda Sarı, Niko Moritz, Sarı, Leda +5 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  14. Unsupervised Domain Adaptation for Speech Recognition via Uncertainty\n Driven Self-Training
    2020/11/26 by Sameer Khurana, Khurana, Sameer, Niko Moritz +5 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  15. Style-transfer based Speech and Audio-visual Scene Understanding for Robot Action Sequence Acquisition from Videos
    2023/06/27 by Chiori Hori, Puyuan Peng, Hori, Chiori +17 · 2 citations
    Computer Science · Arts and Humanities · #Multimodal Machine Learning Applications #Human Pose and Action Recognition #Subtitles and Audiovisual Media
  16. Task-Aware Unified Source Separation
    2024/10/31 by Kohei Saijo, Janek Ebbers, Saijo, Kohei +7 · 1 voice · 2 citations
    #eess.AS #cs.SD
  17. Retrieval-Augmented Neural Field for HRTF Upsampling and Personalization
    2025/01/22 by Yoshiki Masuyama, Gordon Wichern, Masuyama, Yoshiki +7 · 2 citations
    Business, Management and Accounting · #AI and HR Technologies #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  18. Leveraging Audio-Only Data for Text-Queried Target Sound Extraction
    2024/09/20 by Kohei Saijo, Janek Ebbers, Saijo, Kohei +9 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  19. NABEATs: Noise-Aware Audio Representation Learning
    2026/07/18 by Takuya Fujimura, Yoshiki Masuyama, Gordon Wichern +3
    #eess.AS