Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
2023/06/27 by Keqin Chen, Chen, Keqin, Zhao Zhang +9 · 89 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Speech and dialogue systems
paper · pdf · doi:10.48550/arxiv.2306.15195
openalex publication_date 2023/06/27 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
In human conversations, individuals can indicate relevant regions within a scene while addressing others. In turn, the other person can then respond by referring to specific regions if necessary. This natural referential ability in dialogue remains absent in current Multimodal Large Language Models (MLLMs). To fill this gap, this paper proposes an MLLM called Shikra, which can handle spatial coordinate inputs and outputs in natural language. Its architecture consists of a vision encoder, an alignment layer, and a LLM. It is designed to be straightforward and simple, without the need for extra vocabularies, position encoder, pre-/post-detection modules, or external plug-in models. All inputs and outputs are in natural language form. Referential dialogue is a superset of various vision-language (VL) tasks. Shikra can naturally handle location-related tasks like REC and PointQA, as well as conventional VL tasks such as Image Captioning and VQA. Experimental results showcase Shikra's promising performance. Furthermore, it enables numerous exciting applications, like providing mentioned objects' coordinates in chains of thoughts and comparing user-pointed regions similarities. Our code, model and dataset are accessed at https://github.com/shikras/shikra.
Cited by
- Mixture-of-Thought-Tokens: Unifying Perception and Reasoning for Free-form Multimodal Grounding
- Generative Digital Twins: Vision-Language Simulation Models for Executable Industrial Systems
- Unified Multimodal Brain Decoding via Cross-Subject Soft-ROI Fusion
- 4D-RGPT: Toward Region-level 4D Understanding via Perceptual Distillation
- VEGAS: Mitigating Hallucinations in Large Vision-Language Models via Vision-Encoder Attention Guided Adaptive Steering
- UFVideo: Towards Unified Fine-Grained Video Cooperative Understanding with Large Language Models
- VGent: Visual Grounding via Modular Design for Disentangling Reasoning and Prediction
- Grounding Everything in Tokens for Multimodal Large Language Models
- SATGround: A Spatially-Aware Approach for Visual Grounding in Remote Sensing
- CoT4Det: A Chain-of-Thought Framework for Perception-Oriented Vision-Language Tasks
- Visual Reasoning Tracer: Object-Level Grounded Reasoning Benchmark
- GeoViS: Geospatially Rewarded Visual Search for Remote Sensing Visual Grounding
- Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous Driving
- Intervene-All-Paths: Unified Mitigation of LVLM Hallucinations across Alignment Formats
- Octopus: Agentic Multimodal Reasoning with Six-Capability Orchestration
- PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model
- PAS : Prelim Attention Score for Detecting Object Hallucinations in Large Vision--Language Models
- Unveiling Deep Semantic Uncertainty Perception for Language-Anchored Multi-modal Vision-Brain Alignment
- PETAR: Localized Findings Generation with Mask-Aware Vision-Language Modeling for PET Automated Reporting
- PixelRefer: A Unified Framework for Spatio-Temporal Object Referring with Arbitrary Granularity
- Top-Down Semantic Refinement for Image Captioning
- Moving Beyond Diffusion: Hierarchy-to-Hierarchy Autoregression for fMRI-to-Image Reconstruction
- HyperET: Efficient Training in Hyperbolic Space for Multi-modal Large Language Models
- Why LVLMs Are More Prone to Hallucinations in Longer Responses: The Role of Context
- Fake-in-Facext: Towards Fine-Grained Explainable DeepFake Analysis
- Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs
- SceneCOT: Eliciting Grounded Chain-of-Thought Reasoning in 3D Scenes
- CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
- Talking Points: Describing and Localizing Pixels
- Spatial Preference Rewarding for MLLMs Spatial Understanding
- Vision-Centric Activation and Coordination for Multimodal Large Language Models
- The Mechanistic Emergence of Symbol Grounding in Language Models
- Detect Anything via Next Point Prediction
- MIMO: A medical vision language model with visual referring multimodal input and pixel grounding multimodal output
- Reallocating Attention Across Layers to Reduce Multimodal Hallucination
- Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
- Task-Aware Resolution Optimization for Visual Large Language Models
- Constructive Distortion: Improving MLLMs with Attention-Guided Image Warping
- Person-Centric Annotations of LAION-400M: Auditing Bias and Its Transfer to Models
- From Seeing to Predicting: A Vision-Language Framework for Trajectory Forecasting and Controlled Video Generation
- VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
- From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models
- Mixture-of-Visual-Thoughts: Exploring Context-Adaptive Reasoning Mode Selection for General Visual Reasoning
- MIRG-RL: Multi-Image Reasoning and Grounding with Reinforcement Learning
- Beyond Frame Selection: Generative Latent Evidence Aggregation for Long-Video Understanding
- The Photographer Eye: Teaching Multimodal Large Language Models to Understand Image Aesthetics like Photographers
- UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning
- Unleashing the Potential of Multimodal LLMs for Zero-Shot Spatio-Temporal Video Grounding
- Improving Generalized Visual Grounding with Instance-aware Joint Learning
- Re-purposing SAM into Efficient Visual Projectors for MLLM-Based Referring Image Segmentation
- MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook
- 3D Aware Region Prompted Vision Language Model
- MindVL: Towards Efficient and Effective Training of Multimodal Large Language Models on Ascend NPUs
- Towards Understanding Visual Grounding in Visual Language Models
- Tracing and Mitigating Hallucinations in Multimodal LLMs via Dynamic Attention Localization
- Text4Seg++: Advancing Image Segmentation via Generative Language Modeling
- MM-SeR: Multimodal Self-Refinement for Lightweight Image Captioning
- GLSim: Detecting Object Hallucinations in LVLMs via Global-Local Similarity
- Mitigating Hallucinations in Multimodal LLMs via Object-aware Preference Optimization
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Directed-Tokens: A Robust Multi-Modality Alignment Approach to Large Language-Vision Models
- Simple o3: Towards Interleaved Vision-Language Reasoning
- Agentic Design Review System
- SHALE: A Scalable Benchmark for Fine-grained Hallucination Evaluation in LVLMs
- ExpVG: Investigating the Design Space of Visual Grounding in Multimodal Large Language Model
- Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model
- Text-guided Visual Prompt DINO for Generic Segmentation
- SAVER: Mitigating Hallucinations in Large Vision-Language Models via Style-Aware Visual Early Revision
- Modality Bias in LVLMs: Analyzing and Mitigating Object Hallucination via Attention Lens
- Benchmarking and Bridging Emotion Conflicts for Multimodal Emotion Reasoning
- MIHBench: Benchmarking and Mitigating Multi-Image Hallucinations in Multimodal Large Language Models
- Phi-Ground Tech Report: Advancing Perception in GUI Grounding
- Describe, Adapt and Combine: Empowering CLIP Encoders for Open-set 3D Object Retrieval
- Enhancing Large Multimodal Models with Adaptive Sparsity and KV Cache Compression
- RingMo-Agent: A Unified Remote Sensing Foundation Model for Multi-Platform and Multi-Modal Reasoning
- Object-centric Video Question Answering with Visual Grounding and Referring
- IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning
- LMM-Det: Make Large Multimodal Models Excel in Object Detection
- ReMeREC: Relation-aware and Multi-entity Referring Expression Comprehension
- Advancing Visual Large Language Model for Multi-granular Versatile Perception
- Argus: Leveraging Multiview Images for Improved 3-D Scene Understanding With Large Language Models
- Describe Anything Model for Visual Question Answering on Text-rich Images
- KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model
- FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text
- MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines
- LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance
- Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning
- CoT-Segmenter: Enhancing OOD Detection in Dense Road Scenes via Chain-of-Thought Reasoning
- BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset
Related