vix.ing · top · new · best · stats · spec

Shi, Yaya

  1. mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
    2023/04/27 by Qinghao Ye, Ye, Qinghao, Haiyang Xu +32 · 69 citations
    Computer Science · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Topic Modeling
  2. EMScore: Evaluating Video Captioning via Coarse-Grained and Fine-Grained Embedding Matching
    2021/11/17 by Shi, Yaya, Yang, Xu, Xu, Haiyang +4 · 6 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  3. MIBench: Evaluating Multimodal Large Language Models over Multiple Images
    2024/07/21 by Liu, Haowei, Zhang, Xi, Xu, Haiyang +8 · 10 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  4. Youku-mPLUG: A 10 Million Large-scale Chinese Video-Language Dataset for Pre-training and Benchmarks
    2023/06/07 by Haiyang Xu, Qinghao Ye, Xu, Haiyang +29 · 4 citations
    Biochemistry, Genetics and Molecular Biology · Computer Science · #Cancer-related molecular mechanisms research #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications
  5. mPLUG-PaperOwl: Scientific Diagram Analysis with the Multimodal Large Language Model
    2023/11/30 by Anwen Hu, Yaya Shi, Hu, Anwen +17 · 3 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Handwritten Text Recognition Techniques #Multimedia (cs.MM) #Natural Language Processing Techniques #Topic Modeling
  6. iMOVE: Instance-Motion-Aware Video Understanding
    2025/02/17 by Li, Jiaze, Shi, Yaya, Ma, Zongyang +7 · 5 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  7. Object Relational Graph with Teacher-Recommended Learning for Video Captioning
    2020/02/26 by Zhang, Ziqi, Shi, Yaya, Yuan, Chunfeng +4 · 1 citation
    #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  8. TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types
    2025/02/14 by Chen, Jiankang, Zhang, Tianke, Liu, Changyi +8 · 2 citations
    #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences