Zexu Pan
- TF-Locoformer: Transformer with Local Modeling by Convolution for Speech Separation and Enhancement
2024/08/06 by Kohei Saijo, Gordon Wichern, Saijo, Kohei +7 · 11 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- USEV: Universal Speaker Extraction with Visual Cue
2021/09/30 by Zexu Pan, Pan, Zexu, Meng Ge +3 · 4 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Multi-modal Attention for Speech Emotion Recognition
2020/09/09 by Zexu Pan, Pan, Zexu, Zhaojie Luo +5 · 3 citations
Psychology · Computer Science · #Emotion and Mood Recognition #Speech and Audio Processing #Human Pose and Action Recognition
- Target Active Speaker Detection with Audio-visual Cues
2023/05/22 by Yidi Jiang, Ruijie Tao, Jiang, Yidi +5 · 2 citations
Computer Science · Engineering · #Speech and Audio Processing #Speech Recognition and Synthesis #Advanced Adaptive Filtering Techniques
- NeuroHeed+: Improving Neuro-steered Speaker Extraction with Joint Auditory Attention Detection
2023/12/12 by Zexu Pan, Pan, Zexu, Gordon Wichern +7 · 2 citations
Computer Science · Neuroscience · #Audio and Speech Processing (eess.AS) #Blind Source Separation Techniques #EEG and Brain-Computer Interfaces #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- A Hybrid Continuity Loss to Reduce Over-Suppression for Time-domain Target Speaker Extraction
2022/03/31 by Zexu Pan, Meng Ge, Pan, Zexu +3 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- VCSE: Time-Domain Visual-Contextual Speaker Extraction Network
2022/10/09 by Junjie Li, Meng Ge, Li, Junjie +7 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction
2025/05/31 by Cunhang Fan, Fan, Cunhang, Ying Chen +14 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- TF-MossFormer: Integrating Convolution Gated Local-Global Attentions for Enhanced Time-Frequency Domain Monaural Speech Separation
2026/07/23 by Shengkui Zhao, Zexu Pan, Haoxu Wang +3
#cs.SD