vix.ing · top · new · best · stats · spec

Luo, Siwen

  1. PDFVQA: A New Dataset for Real-World VQA on PDF Documents
    2023/04/13 by Ding, Yihao, Luo, Siwen, Chung, Hyunsuk +1 · 5 citations
    #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  2. MAGIC-VQA: Multimodal And Grounded Inference with Commonsense Knowledge for Visual Question Answering
    2025/03/24 by Yang, Shuo, Luo, Siwen, Han, Soyeon Caren +1 · 7 citations
    #Computation and Language (cs.CL) #FOS: Computer and information sciences
  3. VICTR: Visual Information Captured Text Representation for Text-to-Image Multimodal Tasks
    2020/10/07 by Han, Soyeon Caren, Long, Siqu, Luo, Siwen +2 · 1 citation
    #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  4. PDF-MVQA: A Dataset for Multimodal Information Retrieval in PDF-based Visual Question Answering
    2024/04/19 by Ding, Yihao, Ren, Kaixuan, Huang, Jiabin +2 · 1 citation
    #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  5. Multimodal Commonsense Knowledge Distillation for Visual Question Answering
    2024/11/05 by Shuo Yang, Siwen Luo, Yang, Shuo +3 · 1 citation
    Computer Science · #Advanced Image and Video Retrieval Techniques #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Speech and dialogue systems