vix.ing · top · new · best · stats · spec

Zou, Heqing

  1. Speech Emotion Recognition with Co-Attention based Multi-level Acoustic Information
    2022/03/29 by Heqing Zou, Yuke Si, Zou, Heqing +7 · 5 citations
    Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Emotion and Mood Recognition #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  2. Leveraging Modality-specific Representations for Audio-visual Speech Recognition via Reinforcement Learning
    2022/12/10 by Chen, Chen, Hu, Yuchen, Zhang, Qiang +3 · 4 citations
    #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Sound (cs.SD) #electronic engineering #information engineering
  3. Cross-Modality and Within-Modality Regularization for Audio-Visual DeepFake Detection
    2024/01/11 by Zou, Heqing, Shen, Meng, Hu, Yuchen +3 · 5 citations
    #FOS: Computer and information sciences #Multimedia (cs.MM)
  4. From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding
    2024/09/27 by Zou, Heqing, Tianze Luo, Luo, Tianze +15 · 6 citations
    Computer Science · #Multimodal Machine Learning Applications
  5. Cross-Modal Global Interaction and Local Alignment for Audio-Visual Speech Recognition
    2023/05/16 by Yu‐Chen Hu, Ruizhe Li, Hu, Yuchen +9 · 2 citations
    Computer Science · Neuroscience · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Hearing Loss and Rehabilitation #Multimedia (cs.MM) #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  6. Datasets and Recipes for Video Temporal Grounding via Reinforcement Learning
    2025/07/24 by Ruizhe Chen, Zhiting Fan, Chen, Ruizhe +16 · 8 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Video Analysis and Summarization
  7. Interactive Audio-text Representation for Automated Audio Captioning with Contrastive Learning
    2022/03/29 by Chen, Chen, Hou, Nana, Hu, Yuchen +3 · 1 citation
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  8. HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding
    2025/01/03 by Zou, Heqing, Luo, Tianze, Xie, Guiyang +7 · 3 citations
    #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  9. Text-based Talking Video Editing with Cascaded Conditional Diffusion
    2024/07/20 by Bo Han, Heqing Zou, Han, Bo +6 · 1 citation
    Computer Science · #Video Analysis and Summarization #Advanced Data Compression Techniques #Music and Audio Processing