Text4Seg++: Advancing Image Segmentation via Generative Language Modeling
2025/09/08 by Lan, Mengcheng, Chen, Chaofeng, Xu, Jiaxing +6 · 1 citation
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2509.06321
Abstract
Multimodal Large Language Models (MLLMs) have shown exceptional capabilities in vision-language tasks. However, effectively integrating image segmentation into these models remains a significant challenge. In this work, we propose a novel text-as-mask paradigm that casts image segmentation as a text generation problem, eliminating the need for additional decoders and significantly simplifying the segmentation process. Our key innovation is semantic descriptors, a new textual representation of segmentation masks where each image patch is mapped to its corresponding text label. We first introduce image-wise semantic descriptors, a patch-aligned textual representation of segmentation masks that integrates naturally into the language modeling pipeline. To enhance efficiency, we introduce the Row-wise Run-Length Encoding (R-RLE), which compresses redundant text sequences, reducing the length of semantic descriptors by 74% and accelerating inference by 3×, without compromising performance. Building upon this, our initial framework Text4Seg achieves strong segmentation performance across a wide range of vision tasks. To further improve granularity and compactness, we propose box-wise semantic descriptors, which localizes regions of interest using bounding boxes and represents region masks via structured mask tokens called semantic bricks. This leads to our refined model, Text4Seg++, which formulates segmentation as a next-brick prediction task, combining precision, scalability, and generative efficiency. Comprehensive experiments on natural and remote sensing datasets show that Text4Seg++ consistently outperforms state-of-the-art models across diverse benchmarks without any task-specific fine-tuning, while remaining compatible with existing MLLM backbones. Our work highlights the effectiveness, scalability, and generalizability of text-driven image segmentation within the MLLM framework.
Citations
- ALTo: Adaptive-Length Tokenizer for Autoregressive Mask Generation
- Qwen3 Technical Report
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding
- SegEarth-R1: Geospatial Pixel Reasoning via Large Language Model
- POPEN: Preference-Based Optimization and Ensemble for LVLM-Based Reasoning Segmentation
- MMR: A Large-scale Benchmark Dataset for Multi-target and Multi-granularity Reasoning Segmentation
- HiMTok: Learning Hierarchical Mask Tokens for Image Segmentation with Large Multimodal Model
- SegAgent: Exploring Pixel Understanding Capabilities in MLLMs by Imitating Human Annotator Trajectories
- Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement
- UFO: A Unified Approach to Fine-grained Visual Perception via Open-ended Language Interface
- Qwen2.5-VL Technical Report
- SAMRefiner: Taming Segment Anything Model for Universal Mask Refinement
- Pix2Cap-COCO: Advancing Visual Comprehension via Pixel-Level Captioning
- InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models
- DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- GeoGround: A Unified Large Vision-Language Model for Remote Sensing Visual Grounding
- Text4Seg: Reimagining Image Segmentation as Text Generation
- Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- SAM4MLLM: Enhance Multi-Modal Large Language Model for Referring Expression Segmentation
- SWIFT:A Scalable lightWeight Infrastructure for Fine-Tuning
- ProxyCLIP: Proxy Attention Improves CLIP for Open-Vocabulary Segmentation
- LLaVA-OneVision: Easy Visual Task Transfer
- ClearCLIP: Decomposing CLIP Representations for Dense Vision-Language Inference
- GenArtist: Multimodal LLM as an Agent for Unified Image Generation and Editing
- OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding
- VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks
- How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
- Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models
- LaSagnA: Language-based Segmentation Assistant for Complex Queries
- MoMA: Multimodal LLM Adapter for Fast Personalized Image Generation
- Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
- LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
- DeepSeek-VL: Towards Real-World Vision-Language Understanding
- Multi-modal Instruction Tuned LLMs with Fine-grained Visual Perception
- GROUNDHOG: Grounding Large Language Models to Holistic Segmentation
- Jack of All Tasks, Master of Many: Designing General-purpose Coarse-to-Fine Vision-Language Model
- Rotated Multi-Scale Interaction Network for Referring Remote Sensing Image Segmentation
- GSVA: Generalized Segmentation via Multimodal Large Language Models
- PixelLM: Pixel Reasoning with Large Multimodal Model
- Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models
- Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks
- NExT-Chat: An LMM for Chat, Detection and Segmentation
- mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
- GLaMM: Pixel Grounding Large Multimodal Model
- SmooSeg: Smoothness Prior for Unsupervised Semantic Segmentation
- Improved Baselines with Visual Instruction Tuning
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
- LISA: Reasoning Segmentation via Large Language Model
- Towards Open Vocabulary Learning: A Survey
- Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
- Kosmos-2: Grounding Multimodal Large Language Models to the World
- GRES: Generalized Referring Expression Segmentation
- VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
- Otter: A Multi-Modal Model with In-Context Instruction Tuning
- Transformer-Based Visual Segmentation: A Survey
- Visual Instruction Tuning
- Micrograph segmentations for DDEVD
- Segment Anything
- Sigmoid Loss for Language Image Pre-Training
- Universal Instance Perception as Object Discovery and Retrieval
- Annotated Point Clouds and Images from a Spatial-Semantic Perception Pipeline: Robotized Deconstruction at the Reference Construction Site, Aachen
- LLaMA: Open and Efficient Foundation Language Models
- Side Adapter Network for Open-Vocabulary Semantic Segmentation
- PolyFormer: Referring Image Segmentation as Sequential Polygon Generation
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- Open-Vocabulary Semantic Segmentation with Mask-adapted CLIP
- Open-Vocabulary Universal Image Segmentation with MaskCLIP
- Flamingo: a Visual Language Model for Few-Shot Learning
- GroupViT: Semantic Segmentation Emerges from Text Supervision
- Pix2seq: A Language Modeling Framework for Object Detection
- LoRA: Low-Rank Adaptation of Large Language Models
- Learning Transferable Visual Models From Natural Language Supervision
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- PhraseCut: Language-based Image Segmentation in the Wild
- ZeRO: Memory Optimizations Toward Training Trillion Parameter Models
- Panoptic Feature Pyramid Networks
- Semantic Understanding of Scenes through the ADE20K Dataset
- Semantic Understanding of Scenes Through the ADE20K Dataset
- Generation and Comprehension of Unambiguous Object Descriptions
- Microsoft COCO: Common Objects in Context
Cited by
Related