Yin Xie
- LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
2025/09/28 by Xiang An, An, Xiang, Yin Xie +43 · 1 voice · 46 citations
#cs.CV
- RealSyn: An Effective and Scalable Multimodal Interleaved Document Transformation Paradigm
2025/02/18 by Tiancheng Gu, Gu, Tiancheng, Kai-Cheng Yang +15 · 4 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Natural Language Processing Techniques #Semantic Web and Ontologies
- Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
2026/07/27 by Senqiao Yang, Kaichen Zhang, Zhaoyang Jia +20
#cs.CV #cs.CL