vix.ing · top · new · best · stats · spec

Hershey, John R.

  1. SDR - half-baked or well done?
    2018/11/06 by Jonathan Le Roux, Scott Wisdom, Roux, Jonathan Le +4 · 112 citations
    Computer Science · Neuroscience · Engineering · #Speech and Audio Processing #Hearing Loss and Rehabilitation #Advanced Adaptive Filtering Techniques
  2. Deep clustering: Discriminative embeddings for segmentation and separation
    2015/08/18 by John R. Hershey, Hershey, John R., Zhuo Chen +5 · 28 citations
    Computer Science · #Speech and Audio Processing #Music and Audio Processing #Speech Recognition and Synthesis
  3. Full-Capacity Unitary Recurrent Neural Networks
    2016/10/31 by Scott Wisdom, Wisdom, Scott, Thomas A. Powers +7 · 26 citations
    Computer Science · #Blind Source Separation Techniques #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Neural Networks and Applications #Neural and Evolutionary Computing (cs.NE)
  4. Deep Unfolding: Model-Based Inspiration of Novel Deep Architectures
    2014/09/09 by John R. Hershey, Hershey, John R., Jonathan Le Roux +3 · 21 citations
    Computer Science · #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Neural and Evolutionary Computing (cs.NE) #Speech Recognition and Synthesis #Speech and Audio Processing
  5. Integration of speech separation, diarization, and recognition for multi-speaker meetings: System description, comparison, and analysis
    2020/11/03 by Raj, Desh, Denisov, Pavel, Chen, Zhuo +11 · 11 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  6. Attention-Based Multimodal Fusion for Video Description
    2017/01/11 by Hori, Chiori, Hori, Takaaki, Lee, Teng-Yok +3 · 7 citations
    #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimedia (cs.MM)
  7. Universal Sound Separation
    2019/05/08 by Kavalerov, Ilya, Wisdom, Scott, Erdogan, Hakan +4 · 5 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Sound (cs.SD) #electronic engineering #information engineering
  8. Improving Bird Classification with Unsupervised Sound Separation
    2021/10/07 by Tom Denton, Scott Wisdom, Denton, Tom +3 · 6 citations
    Biochemistry, Genetics and Molecular Biology · Computer Science · Environmental Science · #Animal Vocal Communication and Behavior #Speech and Audio Processing #Marine animal studies overview
  9. Into the Wild with AudioScope: Unsupervised Audio-Visual Separation of On-Screen Sounds
    2020/11/02 by Efthymios Tzinis, Scott Wisdom, Tzinis, Efthymios +11 · 4 citations
    Computer Science · Neuroscience · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Hearing Loss and Rehabilitation #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  10. Single-Channel Multi-Speaker Separation using Deep Clustering
    2016/07/07 by Isik, Yusuf, Roux, Jonathan Le, Chen, Zhuo +2 · 3 citations
    #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Sound (cs.SD)
  11. Separating the "Chirp" from the "Chat": Self-supervised Visual Grounding of Sound and Language
    2024/06/09 by Mark Hamilton, Hamilton, Mark, Andrew Zisserman +5 · 6 citations
    Computer Science · #Video Analysis and Summarization
  12. End-to-End Speech Separation with Unfolded Iterative Phase Reconstruction
    2018/04/26 by Wang, Zhong-Qiu, Roux, Jonathan Le, Wang, DeLiang +1 · 2 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Sound (cs.SD) #electronic engineering #information engineering
  13. DF-Conformer: Integrated architecture of Conv-TasNet and Conformer using linear complexity self-attention for speech enhancement
    2021/06/30 by Yuma Koizumi, Koizumi, Yuma, Shigeki Karita +11 · 2 citations
    Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Indoor and Outdoor Localization Technologies #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  14. AudioScopeV2: Audio-Visual Attention Architectures for Calibrated Open-Domain On-Screen Sound Separation
    2022/07/20 by Efthymios Tzinis, Scott Wisdom, Tzinis, Efthymios +5 · 2 citations
    Computer Science · Engineering · Neuroscience · #Advanced Adaptive Filtering Techniques #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Hearing Loss and Rehabilitation #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  15. Distance-Based Sound Separation
    2022/07/01 by Katharine Patterson, Kevin R. Wilson, Patterson, Katharine +5 · 2 citations
    Computer Science · Engineering · #Speech and Audio Processing #Advanced Adaptive Filtering Techniques #Blind Source Separation Techniques
  16. TokenSplit: Using Discrete Speech Representations for Direct, Refined, and Transcript-Conditioned Speech Separation and Recognition
    2023/08/21 by Erdogan, Hakan, Wisdom, Scott, Chang, Xuankai +4 · 2 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  17. Differentiable Consistency Constraints for Improved Deep Speech Enhancement
    2018/11/20 by Wisdom, Scott, Hershey, John R., Wilson, Kevin +4 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  18. A Purely End-to-end System for Multi-speaker Speech Recognition
    2018/05/15 by Seki, Hiroshi, Hori, Takaaki, Watanabe, Shinji +2 · 1 citation
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (stat.ML) #Sound (cs.SD) #electronic engineering #information engineering
  19. Adapting Speech Separation to Real-World Meetings Using Mixture Invariant Training
    2021/10/20 by Sivaraman, Aswin, Wisdom, Scott, Erdogan, Hakan +1 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  20. CycleGAN-Based Unpaired Speech Dereverberation
    2022/03/29 by Hannah Muckenhirn, Muckenhirn, Hannah, Aleksandr Safin +11 · 1 citation
    Computer Science · Psychology · #Speech and Audio Processing #Speech Recognition and Synthesis #Phonetics and Phonology Research
  21. Towards Sub-millisecond Latency Real-Time Speech Enhancement Models on Hearables
    2024/09/26 by Dementyev, Artem, Reddy, Chandan K. A., Wisdom, Scott +3 · 2 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  22. Recomposer: Event-roll-guided generative audio editing
    2025/09/05 by Ellis, Daniel P. W., Fonseca, Eduardo, Weiss, Ron J. +7 · 2 voices · 1 citation
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  23. Continuous Speech Separation Using Speaker Inventory for Long Multi-talker Recording
    2020/12/17 by Cong Han, Yi Luo, Han, Cong +19 · 1 citation
    Computer Science · #Speech and Audio Processing #Speech Recognition and Synthesis #Music and Audio Processing