Jonathan Le Roux
- SDR - half-baked or well done?
2018/11/06 by Jonathan Le Roux, Scott Wisdom, Roux, Jonathan Le +4 · 102 citations
Computer Science · Neuroscience · Engineering · #Speech and Audio Processing #Hearing Loss and Rehabilitation #Advanced Adaptive Filtering Techniques
- Deep clustering: Discriminative embeddings for segmentation and separation
2015/08/18 by John R. Hershey, Hershey, John R., Zhuo Chen +5 · 26 citations
Computer Science · #Speech and Audio Processing #Music and Audio Processing #Speech Recognition and Synthesis
- Full-Capacity Unitary Recurrent Neural Networks
2016/10/31 by Scott Wisdom, Thomas A. Powers, Wisdom, Scott +7 · 25 citations
Computer Science · #Blind Source Separation Techniques #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Neural Networks and Applications #Neural and Evolutionary Computing (cs.NE)
- Deep Unfolding: Model-Based Inspiration of Novel Deep Architectures
2014/09/09 by John R. Hershey, Hershey, John R., Jonathan Le Roux +3 · 15 citations
Computer Science · #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Neural and Evolutionary Computing (cs.NE) #Speech Recognition and Synthesis #Speech and Audio Processing
- TF-Locoformer: Transformer with Local Modeling by Convolution for Speech Separation and Enhancement
2024/08/06 by Kohei Saijo, Gordon Wichern, Saijo, Kohei +7 · 11 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- End-to-End Multi-speaker Speech Recognition with Transformer
2020/02/10 by Xuankai Chang, Wangyou Zhang, Chang, Xuankai +7 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Cold Diffusion for Speech Enhancement
2022/11/04 by Hao Yen, François G. Germain, Yen, Hao +5 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Improving Audio Captioning Models with Fine-grained Audio Features, Text Embedding Supervision, and LLM Mix-up Augmentation
2023/09/29 by Shih-Lun Wu, Xuankai Chang, Wu, Shih-Lun +11 · 3 citations
Computer Science · Arts and Humanities · #Music and Audio Processing #Subtitles and Audiovisual Media #Speech Recognition and Synthesis
- Why does music source separation benefit from cacophony?
2024/02/28 by Chang-Bin Jeon, Gordon Wichern, Jeon, Chang-Bin +5 · 3 citations
Neuroscience · Computer Science · #Hearing Loss and Rehabilitation #Music Technology and Sound Studies #Neuroscience and Music Perception
- (2.5+1)D Spatio-Temporal Scene Graphs for Video Question Answering
2022/02/18 by Anoop Cherian, Cherian, Anoop, Chiori Hori +5 · 2 citations
Computer Science · #Multimodal Machine Learning Applications #Human Pose and Action Recognition #Advanced Image and Video Retrieval Techniques
- The Cocktail Fork Problem: Three-Stem Audio Separation for Real-World Soundtracks
2021/10/19 by Darius Petermann, Gordon Wichern, Petermann, Darius +5 · 2 citations
Computer Science · #Speech and Audio Processing #Music and Audio Processing #Blind Source Separation Techniques
- NeuroHeed+: Improving Neuro-steered Speaker Extraction with Joint Auditory Attention Detection
2023/12/12 by Zexu Pan, Gordon Wichern, Pan, Zexu +7 · 2 citations
Computer Science · Neuroscience · #Audio and Speech Processing (eess.AS) #Blind Source Separation Techniques #EEG and Brain-Computer Interfaces #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- Unsupervised Speaker Adaptation using Attention-based Speaker Memory for\n End-to-End ASR
2020/02/14 by Leda Sarı, Niko Moritz, Sarı, Leda +5 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Unsupervised Domain Adaptation for Speech Recognition via Uncertainty\n Driven Self-Training
2020/11/26 by Sameer Khurana, Khurana, Sameer, Niko Moritz +5 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Style-transfer based Speech and Audio-visual Scene Understanding for Robot Action Sequence Acquisition from Videos
2023/06/27 by Chiori Hori, Puyuan Peng, Hori, Chiori +17 · 2 citations
Computer Science · Arts and Humanities · #Multimodal Machine Learning Applications #Human Pose and Action Recognition #Subtitles and Audiovisual Media
- Task-Aware Unified Source Separation
2024/10/31 by Kohei Saijo, Janek Ebbers, Saijo, Kohei +7 · 1 voice · 2 citations
#eess.AS #cs.SD
- Retrieval-Augmented Neural Field for HRTF Upsampling and Personalization
2025/01/22 by Yoshiki Masuyama, Gordon Wichern, Masuyama, Yoshiki +7 · 2 citations
Business, Management and Accounting · #AI and HR Technologies #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Leveraging Audio-Only Data for Text-Queried Target Sound Extraction
2024/09/20 by Kohei Saijo, Janek Ebbers, Saijo, Kohei +9 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- NABEATs: Noise-Aware Audio Representation Learning
2026/07/18 by Takuya Fujimura, Yoshiki Masuyama, Gordon Wichern +3
#eess.AS