vix.ing · top · new · best · stats · spec

Pan, Zexu

  1. Muse: Multi-modal target speaker extraction with visual cues
    2020/10/15 by Pan, Zexu, Tao, Ruijie, Xu, Chenglin +1 · 5 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Image and Video Processing (eess.IV) #Multimedia (cs.MM) #Sound (cs.SD) #electronic engineering #information engineering
  2. TF-Locoformer: Transformer with Local Modeling by Convolution for Speech Separation and Enhancement
    2024/08/06 by Kohei Saijo, Saijo, Kohei, Gordon Wichern +7 · 11 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  3. USEV: Universal Speaker Extraction with Visual Cue
    2021/09/30 by Zexu Pan, Meng Ge, Pan, Zexu +3 · 4 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  4. Multi-modal Attention for Speech Emotion Recognition
    2020/09/09 by Zexu Pan, Zhaojie Luo, Pan, Zexu +5 · 4 citations
    Psychology · Computer Science · #Emotion and Mood Recognition #Speech and Audio Processing #Human Pose and Action Recognition
  5. NIIRF: Neural IIR Filter Field for HRTF Upsampling and Personalization
    2024/02/27 by Masuyama, Yoshiki, Wichern, Gordon, Germain, François G. +4 · 5 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  6. InspireMusic: Integrating Super Resolution and Large Language Model for High-Fidelity Long-Form Music Generation
    2025/02/28 by Zhang, Chong, Ma, Yukun, Chen, Qian +12 · 6 citations
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  7. Selective Listening by Synchronizing Speech with Lips
    2021/06/14 by Pan, Zexu, Tao, Ruijie, Xu, Chenglin +1 · 2 citations
    #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #electronic engineering #information engineering
  8. Target Active Speaker Detection with Audio-visual Cues
    2023/05/22 by Yidi Jiang, Ruijie Tao, Jiang, Yidi +5 · 2 citations
    Computer Science · Engineering · #Speech and Audio Processing #Speech Recognition and Synthesis #Advanced Adaptive Filtering Techniques
  9. ImagineNET: Target Speaker Extraction with Intermittent Visual Cue through Embedding Inpainting
    2022/10/31 by Pan, Zexu, Wang, Wupeng, Borsdorf, Marvin +1 · 2 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  10. Scenario-Aware Audio-Visual TF-GridNet for Target Speech Extraction
    2023/10/30 by Pan, Zexu, Wichern, Gordon, Masuyama, Yoshiki +4 · 2 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #electronic engineering #information engineering
  11. NeuroHeed+: Improving Neuro-steered Speaker Extraction with Joint Auditory Attention Detection
    2023/12/12 by Zexu Pan, Pan, Zexu, Gordon Wichern +7 · 2 citations
    Computer Science · Neuroscience · #Audio and Speech Processing (eess.AS) #Blind Source Separation Techniques #EEG and Brain-Computer Interfaces #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  12. Generation or Replication: Auscultating Audio Latent Diffusion Models
    2023/10/16 by Bralios, Dimitrios, Wichern, Gordon, Germain, François G. +4 · 2 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  13. SSM2Mel: State Space Model to Reconstruct Mel Spectrogram from the EEG
    2025/01/03 by Fan, Cunhang, Zhang, Sheng, Zhang, Jingjing +2 · 2 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Signal Processing (eess.SP) #Sound (cs.SD) #electronic engineering #information engineering
  14. A Hybrid Continuity Loss to Reduce Over-Suppression for Time-domain Target Speaker Extraction
    2022/03/31 by Zexu Pan, Meng Ge, Pan, Zexu +3 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  15. ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment
    2025/06/24 by Shengkui Zhao, Zhao, Shengkui, Bin Ma +2 · 4 citations
    Computer Science · #Speech and dialogue systems
  16. Late Audio-Visual Fusion for In-The-Wild Speaker Diarization
    2022/11/02 by Pan, Zexu, Wichern, Gordon, Germain, François G. +2 · 1 citation
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  17. VCSE: Time-Domain Visual-Contextual Speaker Extraction Network
    2022/10/09 by Junjie Li, Meng Ge, Li, Junjie +7 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  18. HiFi-SR: A Unified Generative Transformer-Convolutional Adversarial Network for High-Fidelity Speech Super-Resolution
    2025/01/17 by Zhao, Shengkui, Zhou, Kun, Pan, Zexu +3 · 2 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  19. NeuroHeed: Neuro-Steered Speaker Extraction using EEG Signals
    2023/07/26 by Pan, Zexu, Borsdorf, Marvin, Cai, Siqi +2 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #electronic engineering #information engineering
  20. LocSelect: Target Speaker Localization with an Auditory Selective Hearing Mechanism
    2023/10/16 by Chen, Yu, Qian, Xinyuan, Pan, Zexu +2 · 1 citation
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  21. Improved Feature Extraction Network for Neuro-Oriented Target Speaker Extraction
    2025/01/03 by Fan, Cunhang, Gao, Youdian, Pan, Zexu +4 · 2 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  22. M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction
    2025/05/31 by Cunhang Fan, Fan, Cunhang, Ying Chen +14 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  23. Conditional Latent Diffusion-Based Speech Enhancement Via Dual Context Learning
    2025/01/17 by Zhao, Shengkui, Pan, Zexu, Zhou, Kun +3 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering