MM-SeR: Multimodal Self-Refinement for Lightweight Image Captioning
2025/08/29 by Song, Junha, Jo, Yongsik, Min, So Yeon +4
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2508.21451
Abstract
Systems such as video chatbots and navigation robots often depend on streaming image captioning to interpret visual inputs. Existing approaches typically employ large multimodal language models (MLLMs) for this purpose, but their substantial computational cost hinders practical application. This limitation motivates our development of a lightweight captioning model. Our investigation begins by replacing the large-scale language component in MLLMs with a compact 125M-parameter model. Surprisingly, this compact model, despite a 93x reduction in size, achieves comparable performance to MLLMs, suggesting that factual image captioning does not significantly require the complex reasoning abilities of LLMs. Despite this promising result, our lightweight model still lacks reliability. To address this, we draw inspiration from the human visual process: perceiving a global and coarse understanding of the scene before attending to finer details. Accordingly, we propose a multimodal self-refinement framework that guides the model to utilize features from salient regions, identified by referencing the previous coarse caption, and to produce a refined description. Experimental results demonstrate the superiority of our model in both single-sentence and detailed captioning, extending even to long-range video QA tasks.
Citations
- T*: Re-thinking Temporal Search for Long-Form Video Understanding
- Painting with Words: Elevating Detailed Image Captioning with Benchmark and Alignment Learning
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- Qwen2.5-VL Technical Report
- Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models
- Qwen2.5 Technical Report
- Embodied-RAG: General Non-parametric Embodied Memory for Retrieval and Generation
- Attention Prompting on Image for Large Vision-Language Models
- Training Language Models to Self-Correct via Reinforcement Learning
- LLaVA-OneVision: Easy Visual Task Transfer
- MoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language Models
- LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
- VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks
- VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos
- Benchmarking and Improving Detail Image Caption
- Large Language Models Can Self-Correct with Key Condition Verification
- Efficient Multimodal Large Language Models: A Survey
- Hallucination of Multimodal Large Language Models: A Survey
- LocCa: Visual Pretraining with Location-aware Captioners
- InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
- VideoAgent: Long-form Video Understanding with Large Language Model as Agent
- RoboEXP: Action-Conditioned Scene Graph via Interactive Exploration for Robotic Manipulation
- MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark
- StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback
- Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
- A Simple LLM Framework for Long-Range Video Question-Answering
- A Picture is Worth More Than 77 Text Tokens: Evaluating CLIP-Style Models on Dense Captions
- Hallucination Augmented Contrastive Learning for Multimodal Large Language Model
- Multimodal Large Language Models: A Survey
- ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
- mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
- GLaMM: Pixel Grounding Large Multimodal Model
- CLAIR: Evaluating Image Captions with Large Language Models
- From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
- PaLI-3 Vision Language Models: Smaller, Faster, Stronger
- Improved Baselines with Visual Instruction Tuning
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- Automatically Correcting Large Language Models: Surveying the landscape of diverse self-correction strategies
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
- Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
- Image Captioners Are Scalable Vision Learners Too
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- FACTUAL: A Benchmark for Faithful and Consistent Textual Scene Graph Parsing
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
- Visual Instruction Tuning
- DINOv2: Learning Robust Visual Features without Supervision
- Self-Refine: Iterative Refinement with Self-Feedback
- Sigmoid Loss for Language Image Pre-Training
- Positive-Augmented Contrastive Learning for Image and Video Captioning Evaluation
- GPT-4 Technical Report
- Tag2Text: Guiding Vision-Language Model via Image Tagging
- EcoTTA: Memory-Efficient Continual Test-time Adaptation via Self-distilled Regularization
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- Learning Video Representations from Large Language Models
- Semantic-Conditional Diffusion Networks for Image Captioning
- SmallCap: Lightweight Image Captioning Prompted with Retrieval Augmentation
- Analog Bits: Generating Discrete Data using Diffusion Models with Self-Conditioning
- A Survey on Masked Autoencoder for Self-supervised Learning in Vision and Beyond
- CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning
- Emergent Abilities of Large Language Models
- GIT: A Generative Image-to-text Transformer for Vision and Language
- OPT: Open Pre-trained Transformer Language Models
- CaMEL: Mean Teacher Learning for Image Captioning
- I-Tuning: Tuning Frozen Language Models with Image for Lightweight Image Captioning
- Scaling Up Vision-Language Pre-training for Image Captioning
- ClipCap: CLIP Prefix for Image Captioning
- Masked Autoencoders Are Scalable Vision Learners
- An Image is Worth More Than a Thousand Words: Towards Disentanglement in the Wild
- Learning Transferable Visual Models From Natural Language Supervision
- BERTScore: Evaluating Text Generation with BERT
- Neural Baby Talk
- Feature Pyramid Networks for Object Detection
- Microsoft COCO Captions: Data Collection and Evaluation Server
- CIDEr: Consensus-based Image Description Evaluation
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Related