vix.ing · top · new · best · stats · spec

Kevin Qinghong Lin

  1. Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
    2024/08/22 by Jinheng Xie, Weijia Mao, Xie, Jinheng +17 · 167 citations
    Computer Science · #Speech and dialogue systems
  2. Paper2Video: Automatic Video Generation from Scientific Papers
    2025/10/06 by Zeyu Zhu, Kevin Qinghong Lin, Zhu, Zeyu +3 · 8 voices · 6 citations
    Biochemistry, Genetics and Molecular Biology · Computer Science · #Biomedical Text Mining and Ontologies #Mathematics, Computing, and Information Processing #Video Analysis and Summarization #cs.AI #cs.CL #cs.CV #cs.MA #cs.MM
  3. VideoLLM-online: Online Video Large Language Model for Streaming Video
    2024/06/17 by Joya Chen, Chen, Joya, Zhaoyang Lv +17 · 36 citations
    Computer Science · Social Sciences · #Video Analysis and Summarization #Multimedia Communication and Technology
  4. ShowUI: One Vision-Language-Action Model for GUI Visual Agent
    2024/11/26 by Kevin Qinghong Lin, Lin, Kevin Qinghong, Linjie Li +15 · 51 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Multimodal Machine Learning Applications
  5. Egocentric Video-Language Pretraining
    2022/06/03 by Kevin Qinghong Lin, Alex Jinpeng Wang, Lin, Kevin Qinghong +29 · 19 citations
    Biochemistry, Genetics and Molecular Biology · Computer Science · #Artificial Intelligence (cs.AI) #Cancer-related molecular mechanisms research #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Multimodal Machine Learning Applications
  6. EgoVLPv2: Egocentric Video-Language Pre-training with Fusion in the Backbone
    2023/07/11 by Shraman Pramanick, Pramanick, Shraman, Yale Song +13 · 16 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Human Pose and Action Recognition #Multimodal Machine Learning Applications
  7. AssistGPT: A General Multi-modal Assistant that can Plan, Execute, Inspect, and Learn
    2023/06/14 by Difei Gao, Gao, Difei, Lei Ji +10 · 13 citations
    Computer Science · #Natural Language Processing Techniques #Multimodal Machine Learning Applications #Topic Modeling
  8. UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction
    2025/03/19 by Shravan Nayak, Nayak, Shravan, Xiangru Jian +28 · 1 voice · 22 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Gaze Tracking and Assistive Technology #Video Surveillance and Tracking Methods #Virtual Reality Applications and Impacts #cs.AI #cs.CL #cs.CV
  9. Paper2Poster: Towards Multimodal Poster Automation from Scientific Papers
    2025/05/27 by Wei Pang, Kevin Qinghong Lin, Pang, Wei +7 · 1 voice · 15 citations
    #cs.CV #cs.AI #cs.CL #cs.MA
  10. COSMO: COntrastive Streamlined MultimOdal Model with Interleaved Pre-Training
    2024/01/01 by Alex Jinpeng Wang, Linjie Li, Wang, Alex Jinpeng +13 · 3 citations
    Computer Science · #Multimodal Machine Learning Applications #Topic Modeling #Natural Language Processing Techniques
  11. VideoGUI: A Benchmark for GUI Automation from Instructional Videos
    2024/06/14 by Kevin Qinghong Lin, Lin, Kevin Qinghong, Linjie Li +13 · 5 citations
    Computer Science · #Video Analysis and Summarization
  12. Agents' Last Exam
    2026/06/03 by Yiyou Sun, Xinyang Han, Weichen Zhang +306 · 4 voices
    #cs.AI #cs.CL #cs.LG
  13. Egocentric Video-Language Pretraining @ Ego4D Challenge 2022
    2022/07/04 by Kevin Qinghong Lin, Alex Jinpeng Wang, Lin, Kevin Qinghong +29 · 1 citation
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
  14. DiffusionVMR: Diffusion Model for Joint Video Moment Retrieval and Highlight Detection
    2023/08/29 by Henghao Zhao, Zhao, Henghao, Kevin Qinghong Lin +5 · 1 citation
    Computer Science · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Video Analysis and Summarization
  15. VLog: Video-Language Models by Generative Retrieval of Narration Vocabulary
    2025/03/12 by Kevin Qinghong Lin, Lin, Kevin Qinghong, Mike Zheng Shou +1 · 2 citations
    Computer Science · #Video Analysis and Summarization #Topic Modeling #Multimodal Machine Learning Applications
  16. GUI Action Narrator: Where and When Did That Action Take Place?
    2024/06/19 by Qinchen Wu, Difei Gao, Wu, Qinchen +15 · 2 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Persona Design and Applications