LISA: Reasoning Segmentation via Large Language Model
2023/08/01 by Lai, Xin, Tian, Zhuotao, Chen, Yukang +4 · 151 citations
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2308.00692
Abstract
Although perception systems have made remarkable advancements in recent years, they still rely on explicit human instruction or pre-defined categories to identify the target objects before executing visual recognition tasks. Such systems cannot actively reason and comprehend implicit user intention. In this work, we propose a new segmentation task -- reasoning segmentation. The task is designed to output a segmentation mask given a complex and implicit query text. Furthermore, we establish a benchmark comprising over one thousand image-instruction-mask data samples, incorporating intricate reasoning and world knowledge for evaluation purposes. Finally, we present LISA: large Language Instructed Segmentation Assistant, which inherits the language generation capabilities of multimodal Large Language Models (LLMs) while also possessing the ability to produce segmentation masks. We expand the original vocabulary with a token and propose the embedding-as-mask paradigm to unlock the segmentation capability. Remarkably, LISA can handle cases involving complex reasoning and world knowledge. Also, it demonstrates robust zero-shot capability when trained exclusively on reasoning-free datasets. In addition, fine-tuning the model with merely 239 reasoning segmentation data samples results in further performance enhancement. Both quantitative and qualitative experiments show our method effectively unlocks new reasoning segmentation capabilities for multimodal LLMs. Code, models, and data are available at https://github.com/dvlab-research/LISA.
Cited by
- Mixture-of-Thought-Tokens: Unifying Perception and Reasoning for Free-form Multimodal Grounding
- UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling
- StAR: Segment Anything Reasoner
- LogicLens: Visual-Logical Co-Reasoning for Text-Centric Forgery Analysis
- RLLaVA: An RL-central Framework for Language and Vision Assistants
- SegEarth-R2: Towards Comprehensive Language-guided Segmentation for Remote Sensing Images
- ReasonCD: A Multimodal Reasoning Large Model for Implicit Change-of-Interest Semantic Mining
- AMUSE: Audio-Visual Benchmark and Alignment Framework for Agentic Multi-Speaker Understanding
- Vibe Spaces for Creatively Connecting and Expressing Visual Concepts
- A4-Agent: An Agentic Framework for Zero-Shot Affordance Reasoning
- DrivePI: Spatial-aware 4D MLLM for Unified Autonomous Driving Understanding, Perception, Prediction and Planning
- WeDetect: Fast Open-Vocabulary Object Detection as Retrieval
- Moment and Highlight Detection via MLLM Frame Segmentation
- Cross-modal Context-aware Learning for Visual Prompt Guided Multimodal Image Understanding in Remote Sensing
- UFVideo: Towards Unified Fine-Grained Video Cooperative Understanding with Large Language Models
- VGent: Visual Grounding via Modular Design for Disentangling Reasoning and Prediction
- MeViS: A Multi-Modal Dataset for Referring Motion Expression Video Segmentation
- Grounding Everything in Tokens for Multimodal Large Language Models
- GLACIA: Instance-Aware Positional Reasoning for Glacial Lake Segmentation via Multimodal Large Language Model
- SATGround: A Spatially-Aware Approach for Visual Grounding in Remote Sensing
- Towards Unified Semantic and Controllable Image Fusion: A Diffusion Transformer Approach
- Less Is More, but Where? Dynamic Token Compression via LLM-Guided Keyframe Prior
- Omni-Referring Image Segmentation
- CoT4Det: A Chain-of-Thought Framework for Perception-Oriented Vision-Language Tasks
- VG-Refiner: Towards Tool-Refined Referring Grounded Reasoning via Agentic Reinforcement Learning
- Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language Models
- Visual Reasoning Tracer: Object-Level Grounded Reasoning Benchmark
- Malicious Image Analysis via Vision-Language Segmentation Fusion: Detection, Element, and Location in One-shot
- SAM3-I: Segment Anything with Instructions
- OneThinker: All-in-one Reasoning Model for Image and Video
- ReVSeg: Incentivizing the Reasoning Chain for Video Segmentation with Reinforcement Learning
- Learning Visual Affordance from Audio
- Artemis: Structured Visual Reasoning for Perception Policy Learning
- Better, Stronger, Faster: Tackling the Trilemma in MLLM-based Segmentation with Simultaneous Textual Mask Prediction
- DenseScan: Advancing 3D Scene Understanding with 2D Dense Annotation
- RobotSeg: A Model and Dataset for Segmenting Robots in Image and Video
- UniGeoSeg: Towards Unified Open-World Segmentation for Geospatial Scenes
- Ar2Can: An Architect and an Artist Leveraging a Canvas for Multi-Human Generation
- Beyond Real versus Fake Towards Intent-Aware Video Analysis
- UMind-VL: A Generalist Ultrasound Vision-Language Model for Unified Grounded Perception and Comprehensive Interpretation
- MoE3D: Mixture of Experts meets Multi-Modal 3D Understanding
- Syn-GRPO: Self-Evolving Data Synthesis for MLLM Perception Reasoning
- Vision-Language Enhanced Foundation Model for Semi-supervised Medical Image Segmentation
- Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous Driving
- MedSAM3: Delving into Segment Anything with Medical Concepts
- Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
- AVERY: Adaptive VLM Split Computing through Embodied Self-Awareness for Efficient Disaster Response Systems
- CORA: Consistency-Guided Semi-Supervised Framework for Reasoning Segmentation
- UniSER: A Foundation Model for Unified Soft Effects Removal
- Instruction-Guided Lesion Segmentation for Chest X-rays with Automatically Generated Large-Scale Dataset
- Direct Visual Grounding by Directing Attention of Visual Tokens
- Reasoning Text-to-Video Retrieval via Digital Twin Video Representations and Large Language Models
- Fast Reasoning Segmentation for Images and Videos
- Constructing and Interpreting Digital Twin Representations for Visual Reasoning via Reinforcement Learning
- MediRound: Multi-Round Entity-Level Reasoning Segmentation in Medical Images
- PointCubeNet: 3D Part-level Reasoning with 3x3x3 Point Cloud Blocks
- NOVO: Bridging LLaVA and SAM with Visual-only Prompts for Reasoning Segmentation
- S2LM: Towards Semantic Steganography via Large Language Models
- Medical Referring Image Segmentation via Next-Token Mask Prediction
- UniChange: Unifying Change Detection with Multimodal Large Language Model
- URDF-Anything: Constructing Articulated Objects with 3D Multimodal Language Model
- Saliency-R1: Incentivizing Unified Saliency Reasoning Capability in MLLM with Confidence-Guided Reinforcement Learning
- Understanding the Implicit User Intention via Reasoning with Large Language Model for Image Editing
- LangHOPS: Language Grounded Hierarchical Open-Vocabulary Part Segmentation
- REALM: An MLLM-Agent Framework for Open World 3D Reasoning Segmentation and Editing on Gaussian Splatting
- Beyond Reasoning Gains: Mitigating General Capabilities Forgetting in Large Reasoning Models
- FineRS: Fine-grained Reasoning and Segmentation of Small Objects with Reinforcement Learning
- ARGenSeg: Image Segmentation with Autoregressive Image Generation Model
- HyperET: Efficient Training in Hyperbolic Space for Multi-modal Large Language Models
- Fake-in-Facext: Towards Fine-Grained Explainable DeepFake Analysis
- Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation
- Seg the HAB: Language-Guided Geospatial Algae Bloom Reasoning and Segmentation
- Beyond Single Models: Mitigating Multimodal Hallucinations via Adaptive Token Ensemble Decoding
- Enrich and Detect: Video Temporal Grounding with Multimodal LLMs
- Segmentation as A Plug-and-Play Capability for Frozen Multimodal LLMs
- MultiVerse: A Multi-Turn Conversation Benchmark for Evaluating Large Vision and Language Models
- CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
- VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models
- Detect Anything via Next Point Prediction
- GenCellAgent: Generalizable, Training-Free Cellular Image Segmentation via Large Language Model Agents
- CoPRS: Learning Positional Prior from Chain-of-Thought for Reasoning Segmentation
- GIR-Bench: Versatile Benchmark for Generating Images with Reasoning
- ViSurf: Visual Supervised-and-Reinforcement Fine-Tuning for Large Vision-and-Language Models
- Unified Open-World Segmentation with Multi-Modal Prompts
- SaFiRe: Saccade-Fixation Reiteration with Mamba for Referring Image Segmentation
- Complementary and Contrastive Learning for Audio-Visual Segmentation
- MIMO: A medical vision language model with visual referring multimodal input and pixel grounding multimodal output
- Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
- Holistic Order Prediction in Natural Scenes
- Constructive Distortion: Improving MLLMs with Attention-Guided Image Warping
- LTCA: Long-range Temporal Context Attention for Referring Video Object Segmentation
- XYZCylinder: Towards Compatible Feed-Forward 3D Gaussian Splatting for Driving Scenes via Unified Cylinder Lifting Method
- CIR-CoT: Towards Interpretable Composed Image Retrieval via End-to-End Chain-of-Thought Reasoning
- Deforming Videos to Masks: Flow Matching for Referring Video Segmentation
- ChainMPQ: Interleaved Text-Image Reasoning Chains for Mitigating Relation Hallucinations
- UGround: Towards Unified Visual Grounding with Unrolled Transformers
- CoT Referring: Improving Referring Expression Tasks with Grounded Reasoning
- Retrv-R1: A Reasoning-Driven MLLM Framework for Universal and Efficient Multimodal Retrieval
- VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
- From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models
- SVAC: Scaling Is All You Need For Referring Video Object Segmentation
- RAU: Reference-based Anatomical Understanding with Vision Language Models
- Geo-R1: Improving Few-Shot Geospatial Referring Expression Understanding with Reinforcement Fine-Tuning
- Spatial Reasoning in Foundation Models: Benchmarking Object-Centric Spatial Understanding
- MIRG-RL: Multi-Image Reasoning and Grounding with Reinforcement Learning
- Guiding Audio Editing with Audio Language Model
- CAMILA: Context-Aware Masking for Image Editing with Language Alignment
- Beyond Frame Selection: Generative Latent Evidence Aggregation for Long-Video Understanding
- Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer
- Propose and Rectify: A Forensics-Driven MLLM Framework for Image Manipulation Localization
- PoRe: Position-Reweighted Visual Token Pruning for Vision Language Models
- Citrus-V: Advancing Medical Foundation Models with Unified Medical Image Grounding for Clinical Reasoning
- UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning
- SimToken: A Simple Baseline for Referring Audio-Visual Segmentation
- The 1st Solution for 7th LSVOS RVOS Track: SaSaSa2VA
- MDF-MLLM: Deep Fusion Through Cross-Modal Feature Alignment for Contextually Aware Fundoscopic Image Classification
- SmolRGPT: Efficient Spatial Reasoning for Warehouse Environments with 600M Parameters
- Unleashing the Potential of Multimodal LLMs for Zero-Shot Spatio-Temporal Video Grounding
- Improving Generalized Visual Grounding with Instance-aware Joint Learning
- Explain Before You Answer: A Survey on Compositional Visual Reasoning
- MINGLE: VLMs for Semantically Complex Region Detection in Urban Scenes
- CLAIRE: A Dual Encoder Network with RIFT Loss and Phi-3 Small Language Model Based Interpretability for Cross-Modality Synthetic Aperture Radar and Optical Land Cover Segmentation
- Detecting Text Manipulation in Images using Vision Language Models
- Towards Understanding Visual Grounding in Visual Language Models
- Point Linguist Model: Segment Any Object via Bridged Large 3D-Language Model
- Text4Seg++: Advancing Image Segmentation via Generative Language Modeling
- PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination
- PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding?
- GENNAV: Polygon Mask Generation for Generalized Referring Navigable Regions
- PathMR: Multimodal Visual Reasoning for Interpretable Pathology Diagnosis
- Temporal Grounding as a Learning Signal for Referring Video Object Segmentation
- EVTP-IVS: Effective Visual Token Pruning For Unifying Instruction Visual Segmentation In Multi-Modal Large Language Models
- Agentic Design Review System
- ExpVG: Investigating the Design Space of Visual Grounding in Multimodal Large Language Model
- Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model
- Pose-RFT: Enhancing MLLMs for 3D Pose Generation via Hybrid Action Reinforcement Fine-Tuning
- SafePLUG: Empowering Multimodal LLMs with Pixel-Level Insight and Temporal Grounding for Traffic Accident Understanding
- NEP: Autoregressive Image Editing via Next Editing Token Prediction
- User-Intent-Driven Semantic Communication via Adaptive Deep Understanding
- SPEX: A Vision-Language Model for Land Cover Extraction on Spectral Remote Sensing Images
- Latent Expression Generation for Referring Image Segmentation and Grounding
- SAMPO-Path: Segmentation Intent-Aligned Preference Optimization for Pathology Foundation Model Segmentation
- Set Pivot Learning: Redefining Generalized Segmentation with Vision Foundation Models
- Referring Remote Sensing Image Segmentation with Cross-view Semantics Interaction Network
- Fine-grained Spatiotemporal Grounding on Egocentric Videos
- Multimodal Referring Segmentation: A Survey
- RAGNet: Large-scale Reasoning-based Affordance Segmentation Benchmark towards General Grasping
- ART: Adaptive Relation Tuning for Generalized Relation Prediction
- Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation
- RingMo-Agent: A Unified Remote Sensing Foundation Model for Multi-Platform and Multi-Modal Reasoning
- Region-based Cluster Discrimination for Visual Representation Learning
Related