Shehan Munasinghe
- VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos
2024/11/07 by Shehan Munasinghe, Munasinghe, Shehan, Hanan Gani +11 · 15 citations
Computer Science · #Advanced Vision and Imaging #Video Surveillance and Tracking Methods #Video Analysis and Summarization
- PG-Video-LLaVA: Pixel Grounding Large Video-Language Models
2023/11/22 by Shehan Munasinghe, R. Thushara, Munasinghe, Shehan +11 · 5 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling