vix.ing · top · new · best · stats · spec

Shehan Munasinghe

  1. VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos
    2024/11/07 by Shehan Munasinghe, Munasinghe, Shehan, Hanan Gani +11 · 15 citations
    Computer Science · #Advanced Vision and Imaging #Video Surveillance and Tracking Methods #Video Analysis and Summarization
  2. PG-Video-LLaVA: Pixel Grounding Large Video-Language Models
    2023/11/22 by Shehan Munasinghe, R. Thushara, Munasinghe, Shehan +11 · 5 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling