vix.ing · top · new · best · stats · spec

Yin Xie

  1. LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
    2025/09/28 by Xiang An, An, Xiang, Yin Xie +43 · 1 voice · 46 citations
    #cs.CV
  2. RealSyn: An Effective and Scalable Multimodal Interleaved Document Transformation Paradigm
    2025/02/18 by Tiancheng Gu, Gu, Tiancheng, Kai-Cheng Yang +15 · 4 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Natural Language Processing Techniques #Semantic Web and Ontologies
  3. Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
    2026/07/27 by Senqiao Yang, Kaichen Zhang, Zhaoyang Jia +20
    #cs.CV #cs.CL