vix.ing · top · new · best · stats · spec

Zexu Pan

  1. TF-Locoformer: Transformer with Local Modeling by Convolution for Speech Separation and Enhancement
    2024/08/06 by Kohei Saijo, Gordon Wichern, Saijo, Kohei +7 · 11 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  2. USEV: Universal Speaker Extraction with Visual Cue
    2021/09/30 by Zexu Pan, Pan, Zexu, Meng Ge +3 · 4 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  3. Multi-modal Attention for Speech Emotion Recognition
    2020/09/09 by Zexu Pan, Pan, Zexu, Zhaojie Luo +5 · 3 citations
    Psychology · Computer Science · #Emotion and Mood Recognition #Speech and Audio Processing #Human Pose and Action Recognition
  4. Target Active Speaker Detection with Audio-visual Cues
    2023/05/22 by Yidi Jiang, Ruijie Tao, Jiang, Yidi +5 · 2 citations
    Computer Science · Engineering · #Speech and Audio Processing #Speech Recognition and Synthesis #Advanced Adaptive Filtering Techniques
  5. NeuroHeed+: Improving Neuro-steered Speaker Extraction with Joint Auditory Attention Detection
    2023/12/12 by Zexu Pan, Pan, Zexu, Gordon Wichern +7 · 2 citations
    Computer Science · Neuroscience · #Audio and Speech Processing (eess.AS) #Blind Source Separation Techniques #EEG and Brain-Computer Interfaces #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  6. A Hybrid Continuity Loss to Reduce Over-Suppression for Time-domain Target Speaker Extraction
    2022/03/31 by Zexu Pan, Meng Ge, Pan, Zexu +3 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  7. VCSE: Time-Domain Visual-Contextual Speaker Extraction Network
    2022/10/09 by Junjie Li, Meng Ge, Li, Junjie +7 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  8. M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction
    2025/05/31 by Cunhang Fan, Fan, Cunhang, Ying Chen +14 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  9. TF-MossFormer: Integrating Convolution Gated Local-Global Attentions for Enhanced Time-Frequency Domain Monaural Speech Separation
    2026/07/23 by Shengkui Zhao, Zexu Pan, Haoxu Wang +3
    #cs.SD