vix.ing · top · new · best · stats · spec

Xiaoxu Zhu

  1. InteractiveOmni: A Unified Omni-modal Model for Audio-Visual Multi-turn Dialogue
    2025/10/15 by Wenwen Tong, Tong, Wenwen, Dongchuan Ran +46 · 4 citations
    Arts and Humanities · Computer Science · Psychology · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Language, Metaphor, and Cognition #Speech and dialogue systems #Subtitles and Audiovisual Media