vix.ing · top · new · best · stats · spec

Haizhou Li

  1. Ego4D: Around the World in 3,000 Hours of Egocentric Video
    2021/10/13 by Kristen Grauman, Andrew Westbury, Grauman, Kristen +175 · 1 voice · 378 citations
    Computer Science · Social Sciences · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Participatory Visual Research Methods #Video Surveillance and Tracking Methods #cs.AI #cs.CV
  2. ADD 2022: the First Audio Deep Synthesis Detection Challenge
    2022/02/17 by Jiangyan Yi, Ruibo Fu, Yi, Jiangyan +36 · 39 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Digital Media Forensic Detection #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  3. Seen and Unseen emotional style transfer for voice conversion with a new emotional speech dataset
    2020/10/28 by Kun Zhou, Berrak Sisman, Zhou, Kun +6 · 34 citations
    Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #cs.CL #cs.SD #eess.AS #electronic engineering #information engineering
  4. SpEx+: A Complete Time Domain Speaker Extraction Network
    2020/05/10 by Meng Ge, Chenglin Xu, Ge, Meng +9 · 23 citations
    Computer Science · Engineering · Mathematics · #Algorithm #Artificial intelligence #Audio and Speech Processing (eess.AS) #Computer science #Computer vision #Decoding methods #Domain (mathematical analysis) #Embedding #Encoder #FOS: Computer and information sciences #FOS: Electrical engineering #Frequency domain #Mathematics #Music and Audio Processing #Pipeline (software) #SIGNAL (programming language) #Sound (cs.SD) #Speaker recognition #Speaker verification #Speech Recognition and Synthesis #Speech and Audio Processing #Speech recognition #Time domain #cs.SD #eess.AS #electronic engineering #information engineering
  5. Minimizing the Accumulated Trajectory Error to Improve Dataset Distillation
    2022/11/20 by Jiawei Du, Yidi Jiang, Du, Jiawei +7 · 31 citations
    Computer Science · #Advanced Neural Network Applications #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications
  6. AlphaFlowTSE: One-Step Generative Target Speaker Extraction via Conditional AlphaFlow
    2026/03/11 by Duojia Li, Shuhan Zhang, Zihan Qian +5 · 1 voice
    Computer Science · Psychology · #Emotion and Mood Recognition #Generalization #Generative model #Matching (statistics) #Pattern recognition (psychology) #Sampling (signal processing) #Similarity (geometry) #Speech Recognition and Synthesis #Speech and Audio Processing #Trajectory #Utterance #cs.AI #cs.SD