Seg-R1: Segmentation Can Be Surprisingly Simple with Reinforcement Learning
2025/06/27 by You, Zuyao, Wu, Zuxuan · 9 citations
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2506.22624
Abstract
We present Seg-R1, a preliminary exploration of using reinforcement learning (RL) to enhance the pixel-level understanding and reasoning capabilities of large multimodal models (LMMs). Starting with foreground segmentation tasks, specifically camouflaged object detection (COD) and salient object detection (SOD), our approach enables the LMM to generate point and bounding box prompts in the next-token fashion, which are then used to guide SAM2 in producing segmentation masks. We introduce Group Relative Policy Optimization (GRPO) into the segmentation domain, equipping the LMM with pixel-level comprehension through a carefully designed training strategy. Notably, Seg-R1 achieves remarkable performance with purely RL-based training, achieving .873 S-measure on COD10K without complex model modification. Moreover, we found that pure RL training demonstrates strong open-world generalization. Despite being trained solely on foreground segmentation image-mask pairs without text supervision, Seg-R1 achieves impressive zero-shot performance on referring segmentation and reasoning segmentation tasks, with 71.4 cIoU on RefCOCOg test and 56.7 gIoU on ReasonSeg test, outperforming models fully supervised on these datasets.
Citations
- SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL
- Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding
- VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
- Qwen2.5-VL Technical Report
- Pix2Cap-COCO: Advancing Visual Comprehension via Pixel-Level Captioning
- FOCUS: Towards Universal Foreground Segmentation
- Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- Leveraging Hallucinations to Reduce Manual Prompt Dependency in Promptable Segmentation
- SAM 2: Segment Anything in Images and Videos
- OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding
- BioXP-0.5B: Explainable Medical-AI via RL-GRPO
- Bilateral Reference for High-Resolution Dichotomous Image Segmentation
- Relax Image-Specific Prompt Requirement in SAM: A Single Generic Prompt for Segmenting Camouflaged Objects
- PixelLM: Pixel Reasoning with Large Multimodal Model
- GLaMM: Pixel Grounding Large Multimodal Model
- Efficient Memory Management for Large Language Model Serving with PagedAttention
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- LISA: Reasoning Segmentation via Large Language Model
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
- MMBench: Is Your Multi-modal Model an All-around Player?
- GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest
- Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
- MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
- GRES: Generalized Referring Expression Segmentation
- Explicit Visual Prompting for Universal Foreground Segmentations
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Evaluating Object Hallucination in Large Vision-Language Models
- mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
- Visual Instruction Tuning
- Segment Everything Everywhere All at Once
- Micrograph segmentations for DDEVD
- Segment Anything
- Feature Shrinkage Pyramid for Camouflaged Object Detection with Transformers
- Explicit Visual Prompting for Low-Level Structure Segmentations
- Annotated Point Clouds and Images from a Spatial-Semantic Perception Pipeline: Robotized Deconstruction at the Reference Construction Site, Aachen
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- Generalized Decoding for Pixel, Image, and Language
- Open-Vocabulary Semantic Segmentation with Mask-adapted CLIP
- A Weakly Supervised Learning Framework for Salient Object Detection via Hybrid Labels
- Texture-guided Saliency Distilling for Unsupervised Salient Object Detection
- Highly Accurate Dichotomous Image Segmentation
- Zoom In and Out: A Mixed-scale Triplet Network for Camouflaged Object Detection
- Training language models to follow instructions with human feedback
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
- Masked-attention Mask Transformer for Universal Image Segmentation
- Receptive Field Broadening and Boosting for Salient Object Detection
- Per-Pixel Classification is Not All You Need for Semantic Segmentation
- Camouflaged Object Segmentation with Distraction Mining
- Panoptic Segmentation
- Structure-measure: A New Way to Evaluate Foreground Maps
- Structure-Measure: A New Way to Evaluate Foreground Maps
- Proximal Policy Optimization Algorithms
- Modeling Context in Referring Expressions
- A Diagram Is Worth A Dozen Images
- Flickr30k Entities: Collecting Region-to-Phrase Correspondences for\n Richer Image-to-Sentence Models
- Visual Saliency Based on Multiscale Deep Features
- Microsoft COCO: Common Objects in Context
- Reinforcement Learning: An Introduction
Cited by
Related