Masked-attention Mask Transformer for Universal Image Segmentation
2021/12/02 by Bowen Cheng, Ishan Misra, Cheng, Bowen +7 · 504 citations
Computer Science · #Advanced Neural Network Applications #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Medical Image Segmentation Techniques #cs.AI #cs.CV #cs.LG
paper · pdf · doi:10.48550/arxiv.2112.01527
CVPR 2022. Project page/code/models: https://bowenc0221.github.io/mask2former
openalex publication_date 2021/12/02 · arxiv created 2022/06/15 · arxiv updated 2022/06/17 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Image segmentation is about grouping pixels with different semantics, e.g., category or instance membership, where each choice of semantics defines a task. While only the semantics of each task differ, current research focuses on designing specialized architectures for each task. We present Masked-attention Mask Transformer (Mask2Former), a new architecture capable of addressing any image segmentation task (panoptic, instance or semantic). Its key components include masked attention, which extracts localized features by constraining cross-attention within predicted mask regions. In addition to reducing the research effort by at least three times, it outperforms the best specialized architectures by a significant margin on four popular datasets. Most notably, Mask2Former sets a new state-of-the-art for panoptic segmentation (57.8 PQ on COCO), instance segmentation (50.1 AP on COCO) and semantic segmentation (57.7 mIoU on ADE20K).
Citations
Cited by
- Retrieve and Segment: Are a Few Examples Enough to Bridge the Supervision Gap in Open-Vocabulary Segmentation?
- Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
- RoadVGGT: Road-Structure-Aware Feed-Forward Road Surface Reconstruction
- QueenVIS: Rethinking Image-Only Training for Video Instance Segmentation via Query Enrichment
- ESRVS: Extreme Semi-Supervised Retinal Vessel Segmentation with a Single Annotated Image
- DishSeg24k: A Large-Scale Benchmark for Food Segmentation with Stochastic Expert Decoding
- Room-Mediated Co-occurrence for Zero-Shot Object-Centric Semantic Navigation via Frontier Scoring
- Leak-Free Cross-Validated Stacking with Per-Architecture Calibration for Sand-Boil Segmentation in Earthen Levees
- Advancing All-Weather Building Damage Mapping to the Instance Level: Outcomes and Insights from the 2026 Bright Challenge
- UP-Fuse: Uncertainty-guided LiDAR-Camera Fusion for 3D Panoptic Segmentation
- A Lightweight Multi-Scale Attention Framework for Real-Time Spinal Endoscopic Instance Segmentation
- SegEarth-R2: Towards Comprehensive Language-guided Segmentation for Remote Sensing Images
- BiCoR-Seg: Bidirectional Co-Refinement Framework for High-Resolution Remote Sensing Image Segmentation
- Watch Closely: Mitigating Object Hallucinations in Large Vision-Language Models with Disentangled Decoding
- Multifaceted Exploration of Spatial Openness in Rental Housing: A Big Data Analysis in Tokyo's 23 Wards
- RadarGen: Automotive Radar Point Cloud Generation from Cameras
- GenEval 2: Addressing Benchmark Drift in Text-to-Image Evaluation
- Task-Oriented Data Synthesis and Control-Rectify Sampling for Remote Sensing Semantic Segmentation
- Causal-Tune: Mining Causal Factors from Vision Foundation Models for Domain Generalized Semantic Segmentation
- Using Gaussian Splats to Create High-Fidelity Facial Geometry and Texture
- PixelArena: A benchmark for Pixel-Precision Visual Intelligence
- SynthSeg-Agents: Multi-Agent Synthetic Data Generation for Zero-Shot Weakly Supervised Semantic Segmentation
- S2D: Sparse-To-Dense Keymask Distillation for Unsupervised Video Instance Segmentation
- ST-DETrack: Identity-Preserving Branch Tracking in Entangled Plant Canopies via Dual Spatiotemporal Evidence
- Tracking spatial temporal details in ultrasound long video via wavelet analysis and memory bank
- AMD-HookNet++: Evolution of AMD-HookNet with Hybrid CNN-Transformer Feature Enhancement for Glacier Calving Front Segmentation
- Unified Semantic Transformer for 3D Scene Understanding
- TorchTraceAP: A New Benchmark Dataset for Detecting Performance Anti-Patterns in Computer Vision Models
- Pancakes: Consistent Multi-Protocol Image Segmentation Across Biomedical Domains
- Learning to Generate Cross-Task Unexploitable Examples
- Test-Time Modification: Inverse Domain Transformation for Robust Perception
- MeViS: A Multi-Modal Dataset for Referring Motion Expression Video Segmentation
- Super4DR: 4D Radar-centric Self-supervised Odometry and Gaussian-based Map Optimization
- Hot Hém: Sài Gòn Giũa Cái Nóng Hông Còng Bàng -- Saigon in Unequal Heat
- From SAM to DINOv2: Towards Distilling Foundation Models to Lightweight Baselines for Generalized Polyp Segmentation
- SegEarth-OV3: Exploring SAM 3 for Open-Vocabulary Semantic Segmentation in Remote Sensing Images
- XR-DT: Extended Reality-Enhanced Digital Twin for Safe Motion Planning via Human-Aware Model Predictive Path Integral Control
- Generalization vs. Specialization: Evaluating Segment Anything Model (SAM3) Zero-Shot Segmentation Against Fine-Tuned YOLO Detectors
- LiDAS: Lighting-driven Dynamic Active Sensing for Nighttime Perception
- Structure-Aware Feature Rectification with Region Adjacency Graphs for Training-Free Open-Vocabulary Semantic Segmentation
- Enhancing Urban Sensing Utility with Sensor-enabled Vehicles and Easily Accessible Data
- Balanced Learning for Domain Adaptive Semantic Segmentation
- Towards Robust Pseudo-Label Learning in Semantic Segmentation: An Encoding Perspective
- Boosting Unsupervised Video Instance Segmentation with Automatic Quality-Guided Self-Training
- Pseudo-Label Refinement for Robust Wheat Head Segmentation via Two-Stage Hybrid Training
- Are AI-Generated Driving Videos Ready for Autonomous Driving? A Diagnostic Evaluation Framework
- See in Depth: Training-Free Surgical Scene Segmentation with Monocular Depth Priors
- Performance Evaluation of Deep Learning for Tree Branch Segmentation in Autonomous Forestry Systems
- The SAM2-to-SAM3 Gap in the Segment Anything Model Family: Why Prompt-Based Expertise Fails in Concept-Driven Image Segmentation
- Exploiting Domain Properties in Language-Driven Domain Generalization for Semantic Segmentation
- NAS-LoRA: Empowering Parameter-Efficient Fine-Tuning for Visual Foundation Models with Searchable Adaptation
- Flexible Gravitational-Wave Parameter Estimation with Transformers
- Rethinking Surgical Smoke: A Smoke-Type-Aware Laparoscopic Video Desmoking Method and Dataset
- ESACT: An End-to-End Sparse Accelerator for Compute-Intensive Transformers via Local Similarity
- AirSim360: A Panoramic Simulation Platform within Drone View
- SGDiff: Scene Graph Guided Diffusion Model for Image Collaborative SegCaptioning
- ELVIS: Enhance Low-Light for Video Instance Segmentation in the Dark
- Language-Guided Open-World Anomaly Segmentation
- Panda: Self-distillation of Reusable Sensor-level Representations for High Energy Physics
- FOM-Nav: Frontier-Object Maps for Object Goal Navigation
- VFM-ISRefiner: Towards Better Adapting Vision Foundation Models for Interactive Segmentation of Remote Sensing Images
- DEAL-300K: Diffusion-based Editing Area Localization with a 300K-Scale Dataset and Frequency-Prompted Baseline
- UniGeoSeg: Towards Unified Open-World Segmentation for Geospatial Scenes
- Semantic-Centric Alignment for Zero-shot Panoptic Segmentation with Limited Data
- UMind-VL: A Generalist Ultrasound Vision-Language Model for Unified Grounded Perception and Comprehensive Interpretation
- Autonomous labeling of surgical resection margins using a foundation model
- Progress by Pieces: Test-Time Scaling for Autoregressive Image Generation
- Open Vocabulary Compositional Explanations for Neuron Alignment
- V2-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object Correspondence
- CrossEarth-Gate: Fisher-Guided Adaptive Tuning Engine for Efficient Adaptation of Cross-Domain Remote Sensing Semantic Segmentation
- SAM3-Adapter: Efficient Adaptation of Segment Anything 3 for Camouflage Object Segmentation, Shadow Detection, and Medical Image Segmentation
- Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous Driving
- PhysDNet: Physics-Guided Decomposition Network of Side-Scan Sonar Imagery
- Semantic Prioritization in Visual Counterfactual Explanations with Weighted Segmentation and Auto-Adaptive Region Selection
- Lightweight Transformer Framework for Weakly Supervised Semantic Segmentation
- Illustrator's Depth: Monocular Layer Index Prediction for Image Decomposition
- MobileOcc: A Human-Aware Semantic Occupancy Dataset for Mobile Robots
- VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning
- SVG360: Editable Multiview Vector Graphics from a Single SVG
- PairHuman: A High-Fidelity Photographic Dataset for Customized Dual-Person Generation
- InfoCLIP: Bridging Vision-Language Pretraining and Open-Vocabulary Semantic Segmentation via Information-Theoretic Alignment Transfer
- Click2Graph: Interactive Panoptic Video Scene Graphs from a Single Click
- Unsupervised Image Classification with Adaptive Nearest Neighbor Selection and Cluster Ensembles
- MaskMed: Decoupled Mask and Class Prediction for Medical Image Segmentation
- WarNav: An Autonomous Driving Benchmark for Segmentation of Navigable Zones in War Scenes
- The changing surface of the world's roads
- FGNet: Leveraging Feature-Guided Attention to Refine SAM2 for 3D EM Neuron Segmentation
- Fine-Grained Representation for Lane Topology Reasoning
- Seg-VAR: Image Segmentation with Visual Autoregressive Modeling
- Empowering DINO Representations for Underwater Instance Segmentation via Aligner and Prompter
- SkelSplat: Robust Multi-view 3D Human Pose Estimation with Differentiable Gaussian Rendering
- Navigating the Wild: Pareto-Optimal Visual Decision-Making in Image Space
- Visual Bridge: Universal Visual Perception Representations Generating
- Relative Energy Learning for LiDAR Out-of-Distribution Detection
- Leveraging Text-Driven Semantic Variation for Robust OOD Segmentation
- EIDSeg: A Pixel-Level Semantic Segmentation Dataset for Post-Earthquake Damage Assessment from Social Media Images
- Polymap: generating high definition map based on rasterized polygons
- From Words to Safety: Language-Conditioned Safety Filtering for Robot Navigation
- Another BRIXEL in the Wall: Towards Cheaper Dense Features
- No Pose Estimation? No Problem: Pose-Agnostic and Instance-Aware Test-Time Adaptation for Monocular Depth Estimation
- OneOcc: Semantic Occupancy Prediction for Legged Robots with a Single Panoramic Camera
- MIQ-SAM3D: From Single-Point Prompt to Multi-Instance Segmentation via Competitive Query Refinement
- Differentiable Hierarchical Visual Tokenization
- Grounding Surgical Action Triplets with Instrument Instance Segmentation: A Dataset and Target-Aware Fusion Approach
- EPARA: Parallelizing Categorized AI Inference in Edge Clouds
- BeetleFlow: An Integrative Deep Learning Pipeline for Beetle Image Processing
- Generative Semantic Coding for Ultra-Low Bitrate Visual Communication and Analysis
- NaviTrace: Evaluating Embodied Navigation of Vision-Language Models
- Revisiting Generative Infrared and Visible Image Fusion Based on Human Cognitive Laws
- DVPSFormer: Efficient Online Depth-aware Video Panoptic Segmentation for Autonomous Driving
- Cataract-LMM: Large-Scale, Multi-Source, Multi-Task Benchmark for Deep Learning in Surgical Video Analysis
- LangHOPS: Language Grounded Hierarchical Open-Vocabulary Part Segmentation
- Region-CAM: Towards Accurate Object Regions in Class Activation Maps for Weakly Supervised Learning Tasks
- AtlasGS: Atlanta-world Guided Surface Reconstruction with Implicit Structured Gaussians
- IGGT: Instance-Grounded Geometry Transformer for Semantic 3D Reconstruction
- SRSR: Enhancing Semantic Accuracy in Real-World Image Super-Resolution with Spatially Re-Focused Text-Conditioning
- Simplifying Knowledge Transfer in Pretrained Models
- AURASeg: Attention Guided Upsampling with Residual Boundary-Assistive Refinement for Drivable-Area Segmentation
- Dynamic Semantic-Aware Correlation Modeling for UAV Tracking
- Controllable-LPMoE: Adapting to Challenging Object Segmentation via Dynamic Local Priors from Mixture-of-Experts
- ARGenSeg: Image Segmentation with Autoregressive Image Generation Model
- SFGFusion: Surface Fitting Guided 3D Object Detection with 4D Radar and Camera Fusion
- Symmetric Entropy-Constrained Video Coding for Machines
- Self-Supervised Learning to Fly using Efficient Semantic Segmentation and Metric Depth Estimation for Low-Cost Autonomous UAVs
- Aria Gen 2 Pilot Dataset
- MOBIUS: Big-to-Mobile Universal Instance Segmentation via Multi-modal Bottleneck Fusion and Calibrated Decoder Pruning
- CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
- Multi-modal video data-pipelines for machine learning with minimal human supervision
- EuroMineNet: A Multitemporal Sentinel-2 Benchmark for Spatiotemporal Mining Footprint Analysis in the European Union (2015-2024)
- Towards Generalist Intelligence in Dentistry: Vision Foundation Models for Oral and Maxillofacial Radiology
- UrbanVerse: Scaling Urban Simulation by Watching City-Tour Videos
- UniVector: Unified Vector Extraction via Instance-Geometry Interaction
- A Semi-amortized Lifted Learning-to-Optimize Masked (SALLO-M) Transformer Model for Scalable and Generalizable Beamforming
- UniFusion: Vision-Language Model as Unified Encoder in Image Generation
- Generalisation of automatic tumour segmentation in histopathological whole-slide images across multiple cancer types
- MSCloudCAM: Multi-Scale Context Adaptation with Convolutional Cross-Attention for Multispectral Cloud Segmentation
- A Machine Learning Perspective on Automated Driving Corner Cases
- Unified Open-World Segmentation with Multi-Modal Prompts
- FRIEREN: Federated Learning with Vision-Language Regularization for Segmentation
- Probabilistic Hyper-Graphs using Multiple Randomly Masked Autoencoders for Semi-supervised Multi-modal Multi-task Learning
- Complementary and Contrastive Learning for Audio-Visual Segmentation
- Explainable Human-in-the-Loop Segmentation via Critic Feedback Signals
- ClustViT: Clustering-based Token Merging for Semantic Segmentation
- What Matters in RL-Based Methods for Object-Goal Navigation? An Empirical Study and A Unified Framework
- Holistic Order Prediction in Natural Scenes
- Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and Grounding
- SilvaScenes: Tree Detection and Species Classification from Under-Canopy Images in Natural Forests
- X2Video: Adapting Diffusion Models for Multimodal Controllable Neural Video Rendering
- LTCA: Long-range Temporal Context Attention for Referring Video Object Segmentation
- UniMMVSR: A Unified Multi-Modal Framework for Cascaded Video Super-Resolution
- AlignGS: Aligning Geometry and Semantics for Robust Indoor Reconstruction from Sparse Views
- Locality-Sensitive Hashing-Based Efficient Point Transformer for Charged Particle Reconstruction
- Geometry-Aware Cross Modal Alignment for Light Field-LiDAR Semantic Segmentation
- Data Factory with Minimal Human Effort Using VLMs
- Human Action Recognition from Point Clouds over Time
- From Filters to VLMs: Benchmarking Defogging Methods through Object Detection and Segmentation Performance
- UGround: Towards Unified Visual Grounding with Unrolled Transformers
- IMAGEdit: Let Any Subject Transform
- KeySG: Hierarchical Keyframe-Based 3D Scene Graphs
- Semantic Visual Simultaneous Localization and Mapping: A Survey on State of the Art, Challenges, and Future Directions
- Robust Context-Aware Object Recognition
- SAGE-LD: Towards Scalable and Generalizable End-to-End Language Diarization via Simulated Data Augmentation
- Stitch: Training-Free Position Control in Multimodal Diffusion Transformers
- IRIS: Intrinsic Reward Image Synthesis
- CLASP: Adaptive Spectral Clustering for Unsupervised Per-Image Segmentation
- CORE-3D: Context-aware Open-vocabulary Retrieval by Embeddings in 3D
- K-Prism: A Knowledge-Guided and Prompt Integrated Universal Medical Image Segmentation Model
- HieraTok: Multi-Scale Visual Tokenizer Improves Image Reconstruction and Generation
- Token Merging via Spatiotemporal Information Mining for Surgical Video Understanding
- Foundation Model-Based Adaptive Semantic Image Transmission for Dynamic Wireless Environments
- UniMapGen: A Generative Framework for Large-Scale Map Construction from Multi-modal Data
- Learning What To Hear: Boosting Sound-Source Association For Robust Audiovisual Instance Segmentation
- CubistMerge: Spatial-Preserving Token Merging For Diverse ViT Backbones
- Queryable 3D Scene Representation: A Multi-Modal Framework for Semantic Reasoning and Robotic Task Planning
- Boosting LiDAR-Based Localization with Semantic Insight: Camera Projection versus Direct LiDAR Segmentation
- ZMIS-SAM: Segment Anything Model Enhanced with Wavelet Transform for Zooplankton Microscopy Image Instance Segmentation
- IGME: Efficient Chained Method Ensemble for Transferable Semantic Segmentation Attacks
- MLF-4DRCNet: Multi-Level Fusion with 4D Radar and Camera for 3D Object Detection in Autonomous Driving
- Frequency-Domain Decomposition and Recomposition for Robust Audio-Visual Segmentation
- Surgical Video Understanding with Label Interpolation
- Visual Instruction Pretraining for Domain-Specific Foundation Models
- Region-Aware Deformable Convolutions
- White Aggregation and Restoration for Few-shot 3D Point Cloud Semantic Segmentation
- Masked Feature Modeling Enhances Adaptive Segmentation
- Improving Generalized Visual Grounding with Instance-aware Joint Learning
- Re-purposing SAM into Efficient Visual Projectors for MLLM-Based Referring Image Segmentation
- Mitigating Query Selection Bias in Referring Video Object Segmentation
- First Place Solution to the MLCAS 2025 GWFSS Challenge: The Devil is in the Detail and Minority
- NavMoE: Hybrid Model- and Learning-based Traversability Estimation for Local Navigation via Mixture of Experts
- Road Obstacle Video Segmentation
- RailSafeNet: Visual Scene Understanding for Tram Safety
- Exploring Efficient Open-Vocabulary Segmentation in the Remote Sensing
- CLAIRE: A Dual Encoder Network with RIFT Loss and Phi-3 Small Language Model Based Interpretability for Cross-Modality Synthetic Aperture Radar and Optical Land Cover Segmentation
- Microsurgical Instrument Segmentation for Robot-Assisted Surgery
- Multimodal SAM-adapter for Semantic Segmentation
- I-Segmenter: Integer-Only Vision Transformer for Efficient Semantic Segmentation
- Leveraging Multi-View Weak Supervision for Occlusion-Aware Multi-Human Parsing
- DGFusion: Depth-Guided Sensor Fusion for Robust Semantic Perception
- Prompt-Driven Image Analysis with Multimodal Generative AI: Detection, Segmentation, Inpainting, and Interpretation
- UNO: Unifying One-stage Video Scene Graph Generation via Object-Centric Visual Representation Learning
- A biologically inspired separable learning vision model for real-time traffic object perception in Dark
- Decoding Tourist Perception in Historic Urban Quarters with Multimodal Social Media Data: An AI-Based Framework and Evidence from Shanghai
- Skywork UniPic 2.0: Building Kontext Model with Online RL for Unified Multimodal Model
- HAMSt3R: Human-Aware Multi-view Stereo 3D Reconstruction
- Unsupervised Instance Segmentation with Superpixels
- InstaDA: Augmenting Instance Segmentation Data with Dual-Agent System
- SOPSeg: Prompt-based Small Object Instance Segmentation in Remote Sensing Imagery
- EM3M: An Electron Micrograph Dataset for Microstructural Segmentation and Generation
- An Investigation of Visual Foundation Models Robustness
- MedDINOv3: How to adapt vision foundation models for medical image segmentation?
- RAGSR: Regional Attention Guided Diffusion for Image Super-Resolution
- No More Sibling Rivalry: Debiasing Human-Object Interaction Detection
- Can General-Purpose Omnimodels Compete with Specialists? A Case Study in Medical Image Segmentation
- DGL-RSIS: Decoupling Global Spatial Context and Local Class Semantics for Training-Free Remote Sensing Image Segmentation
- Learning Yourself: Class-Incremental Semantic Segmentation with Language-Inspired Bootstrapped Disentanglement
- Representation Learning with Adaptive Superpixel Coding
- Inference-Time Alignment Control for Diffusion Models with Reinforcement Learning Guidance
- Webly-Supervised Image Manipulation Localization via Category-Aware Auto-Annotation
- GLOW: A Unified Particle Flow Transformer
- Image Quality Assessment for Machines: Paradigm, Large-scale Database, and Models
- AutoQ-VIS: Improving Unsupervised Video Instance Segmentation via Automatic Quality Assessment
- FreeVPS: Repurposing Training-Free SAM2 for Generalizable Video Polyp Segmentation
- IELDG: Suppressing Domain-Specific Noise with Inverse Evolution Layers for Domain Generalized Semantic Segmentation
- Autoregressive Universal Video Segmentation Model
- RoofSeg: An edge-aware transformer-based network for end-to-end roof plane segmentation
- PseudoMapTrainer: Learning Online Mapping without HD Maps
- CM2LoD3: Reconstructing LoD3 Building Models Using Semantic Conflict Maps
- EventSSEG: Event-driven Self-Supervised Segmentation with Probabilistic Attention
- Multiscale Video Transformers for Class Agnostic Segmentation in Autonomous Driving
- A Comprehensive Review of Agricultural Parcel and Boundary Delineation from Remote Sensing Images: Recent Progress and Future Perspectives
- MoVieDrive: Urban Scene Synthesis with Multi-Modal Multi-View Video Diffusion Transformer
- LENS: Learning to Segment Anything with Unified Reinforced Reasoning
- SIS-Challenge: Event-based Spatio-temporal Instance Segmentation Challenge at the CVPR 2025 Event-based Vision Workshop
- Temporal Grounding as a Learning Signal for Referring Video Object Segmentation
- EVTP-IVS: Effective Visual Token Pruning For Unifying Instruction Visual Segmentation In Multi-Modal Large Language Models
- OVSegDT: Segmenting Transformer for Open-Vocabulary Object Goal Navigation
- Generalized Decoupled Learning for Enhancing Open-Vocabulary Dense Perception
- CRISP: Contrastive Residual Injection and Semantic Prompting for Continual Video Instance Segmentation
- Unlocking Robust Semantic Segmentation Performance via Label-only Elastic Deformations against Implicit Label Noise
- From Pixel to Mask: A Survey of Out-of-Distribution Segmentation
- TRACE: Learning 3D Gaussian Physical Dynamics from Multi-view Videos
- FM4NPP: A Scaling Foundation Model for Nuclear and Particle Physics
- RelayFormer: A Unified Local-Global Attention Framework for Scalable Image and Video Manipulation Localization
- Revisiting Efficient Semantic Segmentation: Learning Offsets for Better Spatial and Class Feature Alignment
- SHREC 2025: Retrieval of Optimal Objects for Multi-modal Enhanced Language and Spatial Assistance (ROOMELSA)
- Hierarchical Visual Prompt Learning for Continual Video Instance Segmentation
- Vision Generalist Model: A Survey
- Planner-Refiner: Dynamic Space-Time Refinement for Vision-Language Alignment in Videos
- EventRR: Event Referential Reasoning for Referring Video Object Segmentation
- Text-guided Visual Prompt DINO for Generic Segmentation
- UGD-IML: A Unified Generative Diffusion-based Framework for Constrained and Unconstrained Image Manipulation Localization
- Learning 3D Texture-Aware Representations for Parsing Diverse Human Clothing and Body Parts
- SGDFuse: SAM-Guided Diffusion for High-Fidelity Infrared and Visible Image Fusion
- SPEX: A Vision-Language Model for Land Cover Extraction on Spectral Remote Sensing Images
- Decoupling Continual Semantic Segmentation
- Temporal Cluster Assignment for Efficient Real-Time Video Segmentation
- TSMS-SAM2: Multi-scale Temporal Sampling Augmentation and Memory-Splitting Pruning for Promptable Video Object Segmentation and Tracking in Surgical Scenarios
- X-SAM: From Segment Anything to Any Segmentation
- Two-Way Garment Transfer: Unified Diffusion Framework for Dressing and Undressing Synthesis
- A2Mamba: Attention-augmented State Space Models for Visual Recognition
- What Holds Back Open-Vocabulary Segmentation?
- TNet: Terrace Convolutional Decoder Network for Remote Sensing Image Semantic Segmentation
- DOMR: Establishing Cross-View Segmentation via Dense Object Matching
- Prototype-Driven Structure Synergy Network for Remote Sensing Images Segmentation
- MetaScope: Optics-Driven Neural Network for Ultra-Micro Metalens Endoscopy
- ParticleSAM: Small Particle Segmentation for Material Quality Monitoring in Recycling Processes
- AlignCAT: Visual-Linguistic Alignment of Category and Attribute for Weakly Supervised Visual Grounding
- Multi-Granularity Feature Calibration via VFM for Domain Generalized Semantic Segmentation
- LT-Gaussian: Long-Term Map Update Using 3D Gaussian Splatting for Autonomous Driving
- Rein++: Efficient Generalization and Adaptation for Semantic Segmentation with Vision Foundation Models
- Set Pivot Learning: Redefining Generalized Segmentation with Vision Foundation Models
- IAUNet: Instance-Aware U-Net
- RMT-PPAD: Real-time Multi-task Learning for Panoptic Perception in Autonomous Driving
- Training-Free Class Purification for Open-Vocabulary Semantic Segmentation
- SDMatte: Grafting Diffusion Models for Interactive Matting
- Multimodal Referring Segmentation: A Survey
- UIS-Mamba: Exploring Mamba for Underwater Instance Segmentation via Dynamic Tree Scan and Hidden State Weaken
- MagicRoad: Semantic-Aware 3D Road Surface Reconstruction via Obstacle Inpainting
- Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation
- Segment Anything for Video: A Comprehensive Review of Video Object Segmentation and Tracking from Past to Future
- AlphaDent: A dataset for automated tooth pathology detection
- Exploring Probabilistic Modeling Beyond Domain Generalization for Semantic Segmentation
- ObjectGS: Object-aware Scene Reconstruction and Scene Understanding via Gaussian Splatting
- Lightweight Transformer-Driven Segmentation of Hotspots and Snail Trails in Solar PV Thermal Imagery
- Conditional Video Generation for High-Efficiency Video Compression
- SeeDiff: Off-the-Shelf Seeded Mask Generation from Diffusion Models
- Latest Object Memory Management for Temporally Consistent Video Instance Segmentation
- SurgPIS: Surgical-instrument-level Instances and Part-level Semantics for Weakly-supervised Part-aware Instance Segmentation
- MixA-Q: Revisiting Activation Sparsity for Vision Transformers from a Mixed-Precision Quantization Perspective
- Synthetic Data Augmentation for Enhanced Chicken Carcass Instance Segmentation
- Privacy-Preserving Semantic Segmentation from Ultra-Low-Resolution RGB Inputs
- Advancing Complex Video Object Segmentation via Progressive Concept Construction
- Object segmentation in the wild with foundation models: application to vision assisted neuro-prostheses for upper limbs
- GVCCS : Ground Visible Camera Contrail Sequences
- GTPBD: A Fine-Grained Global Terraced Parcel and Boundary Dataset
- Moving Object Detection from Moving Camera Using Focus of Expansion Likelihood and Segmentation
- IndoorBEV: Joint Detection and Footprint Completion of Objects via Mask-based Prediction in Indoor Scenarios for Bird's-Eye View Perception
- Leveraging Pathology Foundation Models for Panoptic Segmentation of Melanoma in H&E Images
- Test-time Prompt Refinement for Text-to-Image Models
- Part Segmentation of Human Meshes via Multi-View Human Parsing
- Comparative validation of surgical phase recognition, instrument keypoint estimation, and instrument instance segmentation in endoscopy: Results of the PhaKIR 2024 challenge
- Advancing Visual Large Language Model for Multi-granular Versatile Perception
- SCORE: Scene Context Matters in Open-Vocabulary Remote Sensing Instance Segmentation
- NLI4VolVis: Natural Language Interaction for Volume Visualization via LLM Multi-Agents and Editable 3D Gaussian Splatting
- AD-GS: Object-Aware B-Spline Gaussian Splatting for Self-Supervised Autonomous Driving
- Frequency-Dynamic Attention Modulation for Dense Prediction
- From Binary to Semantic: Utilizing Large-Scale Binary Occupancy Data for 3D Semantic Occupancy Prediction
- Spatial Frequency Modulation for Semantic Segmentation
- Data-Efficient Challenges in Visual Inductive Priors: A Retrospective
- SGLoc: Semantic Localization System for Camera Pose Estimation from 3D Gaussian Splatting Representation
- DCD: A Semantic Segmentation Model for Fetal Ultrasound Four-Chamber View
- DEARLi: Decoupled Enhancement of Recognition and Localization for Semi-supervised Panoptic Segmentation
- Stereo-based 3D Anomaly Object Detection for Autonomous Driving: A New Dataset and Baseline
- Advancing Medical Image Segmentation via Self-supervised Instance-adaptive Prototype Learning
- Objectomaly: Objectness-Aware Refinement for OoD Segmentation with Structural Consistency and Boundary Precision
- Rethinking Query-based Transformer for Continual Image Segmentation
- A multi-modal dataset for insect biodiversity with imagery and DNA at the trap and individual level
- SemRaFiner: Panoptic Segmentation in Sparse and Noisy Radar Point Clouds
- First Investigation of Deep Learning for Intraoperative Gauze Segmentation in Minimally Invasive Abdominal Surgery
- What Demands Attention in Urban Street Scenes? From Scene Understanding towards Road Safety: A Survey of Vision-driven Datasets and Studies
- AnthroTAP: Learning Point Tracking with Real-World Motion
- SPADE: Spatial-Aware Denoising Network for Open-vocabulary Panoptic Scene Graph Generation with Long- and Local-range Context Reasoning
- Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation
- Learning from Adversity: Semantic-Aware Mask Refinement through Adversarial Perturbation
- All in One: Visual-Description-Guided Unified Point Cloud Segmentation
- From Vision To Language through Graph of Events in Space and Time: An Explainable Self-supervised Approach
- MOSU: Autonomous Long-range Robot Navigation with Multi-modal Scene Understanding
- VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs
- OpenWorldSAM: Extending SAM2 for Universal Image Segmentation with Language Prompts
- Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation
- NRSeg: Noise-Resilient Learning for BEV Semantic Segmentation via Driving World Models
- PLUS: Plug-and-Play Enhanced Liver Lesion Diagnosis Model on Non-Contrast CT Scans
- PhenoBench: A Comprehensive Benchmark for Cell Phenotyping
- Leveraging Image Generators to Address Data Scarcity: The Gen4Regen Dataset for Forest Regeneration Mapping
- No time to train! Training-Free Reference-Based Instance Segmentation
- SIU3R: Simultaneous Scene Understanding and 3D Reconstruction Beyond Feature Alignment
- Diagnosing and Correcting Concept Omission in Multimodal Diffusion Transformers
- ViRefSAM: Visual Reference-Guided Segment Anything Model for Remote Sensing Segmentation
- DeRIS: Decoupling Perception and Cognition for Enhanced Referring Image Segmentation through Loopback Synergy
- A Gift from the Integration of Discriminative and Diffusion-based Generative Learning: Boundary Refinement Remote Sensing Semantic Segmentation
- Learning an Ensemble Token from Task-driven Priors in Facial Analysis
- Learning from Random Subspace Exploration: Generalized Test-Time Augmentation with Self-supervised Distillation
- SelvaBox: A high-resolution dataset for tropical tree crown detection
- Perception Characteristics Distance: Measuring Stability and Robustness of Perception System in Dynamic Conditions under a Certain Decision Rule
- MedSAM-CA: A CNN-Augmented ViT with Attention-Enhanced Multi-Scale Fusion for Medical Image Segmentation
- Revisiting Audio-Visual Segmentation with Vision-Centric Transformer
- Probabilistic Prototype Calibration of Vision-Language Models for Generalized Few-shot Semantic Segmentation
- FreeGave: 3D Physics Learning from Dynamic Videos by Gaussian Velocity
- Seg-R1: Segmentation Can Be Surprisingly Simple with Reinforcement Learning
- Partial CLIP is Enough: Chimera-Seg for Zero-shot Semantic Segmentation
- Improving the quality of respiratory signals extracted from the segmented mask area
- MADrive: Memory-Augmented Driving Scene Modeling
- CAT-SG: A Large Dynamic Scene Graph Dataset for Fine-Grained Understanding of Cataract Surgery
- SAM2-SGP: Enhancing SAM2 for Medical Image Segmentation via Support-Set Guided Prompting
- Flow-Anything: Learning Real-World Optical Flow Estimation from Large-Scale Single-view Images
- Pre-Trained LLM is a Semantic-Aware and Generalizable Segmentation Booster
- Relation3D: Enhancing Relation Modeling for Point Cloud Instance Segmentation
- FOCUS: Unified Vision-Language Modeling for Interactive Editing Driven by Referential Segmentation
- Loupe: A Generalizable and Adaptive Framework for Image Forgery Detection
- Stepping Out of Similar Semantic Space for Open-Vocabulary Segmentation
- Image Segmentation with Large Language Models: A Survey with Perspectives for Intelligent Transportation Systems
- Leader360V: The Large-scale, Real-world 360 Video Dataset for Multi-task Learning in Diverse Environment
- Segmenting Visuals With Querying Words: Language Anchors For Semi-Supervised Image Segmentation
- Open-Set LiDAR Panoptic Segmentation Guided by Uncertainty-Aware Learning
- FOAM: A General Frequency-Optimized Anti-Overlapping Framework for Overlapping Object Perception
- Unleashing Diffusion and State Space Models for Medical Image Segmentation
- Affogato: Open-Vocabulary Affordance Grounding with Automated Data Generation at Scale
- O2Former:Direction-Aware and Multi-Scale Query Enhancement for SAR Ship Instance Segmentation
- Symmetrical Flow Matching: Unified Image Generation, Segmentation, and Classification with Score-Based Generative Models
- MSTAR: Box-free Multi-query Scene Text Retrieval with Attention Recycling
- GLD-Road:A global-local decoding road network extraction model for remote sensing images
- O-MaMa: Learning Object Mask Matching between Egocentric and Exocentric Views
- Query Nearby: Offset-Adjusted Mask2Former enhances small-organ segmentation
- GS4: Generalizable Sparse Splatting Semantic SLAM
- WoundAIssist: A Patient-Centered Mobile App for AI-Assisted Wound Care With Physicians in the Loop
- Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025
- Bringing SAM to new heights: Leveraging elevation data for tree crown segmentation from drone imagery
- Predicting Road Surface Anomalies by Visual Tracking of a Preceding Vehicle
- PhenoStitch: Training-Free Panoptic Crop Mapping from Satellite Image Time Series
- Are Synthetic Corruptions A Reliable Proxy For Real-World Corruptions?
- DeCLIP: Decoupled Learning for Open-Vocabulary Dense Perception
- BiXFormer: A Robust Framework for Maximizing Modality Effectiveness in Multi-Modal Semantic Segmentation
- ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads
- Unified Attention Modeling for Efficient Free-Viewing and Visual Search via Shared Representations
- Towards In-the-wild 3D Plane Reconstruction from a Single Image
- Pan-Arctic Permafrost Landform and Human-built Infrastructure Feature Detection with Vision Transformers and Location Embeddings
- InterRVOS: Interaction-aware Referring Video Object Segmentation
- unMORE: Unsupervised Multi-Object Segmentation via Center-Boundary Reasoning
- G4Seg: Generation for Inexact Segmentation Refinement with Diffusion Models
- ADEPT: Adaptive Diffusion Environment for Policy Transfer Sim-to-Real
- Seg2Any: Open-set Segmentation-Mask-to-Image Generation with Precise Shape and Semantic Control
- Understanding while Exploring: Semantics-driven Active Mapping
- Panoramic Out-of-Distribution Segmentation
- Point or Line? Using Line-based Representation for Panoptic Symbol Spotting in CAD Drawings
- PixelThink: Towards Efficient Chain-of-Pixel Reasoning
- Advancing Generalizable Tumor Segmentation with Anomaly-Aware Open-Vocabulary Attention Maps and Frozen Foundation Diffusion Models
- UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning
- CAST: Contrastive Adaptation and Distillation for Semi-Supervised Instance Segmentation
- Precise Object and Effect Removal with Adaptive Target-Aware Attention
- S2AFormer: Strip Self-Attention for Efficient Vision Transformer
- On Geometry-Enhanced Parameter-Efficient Fine-Tuning for 3D Scene Segmentation
- SANSA: Unleashing the Hidden Semantics in SAM2 for Few-Shot Segmentation
- The Missing Point in Vision Transformers for Universal Image Segmentation
- What You Perceive Is What You Conceive: A Cognition-Inspired Framework for Open Vocabulary Image Segmentation
- Location-guided lesions representation learning via image generation for assessing plant leaf diseases severity
- EOVSAM: Efficient Open-Vocabulary Segmentation with SAM 3 in One Pass
- Reasoning Segmentation for Images and Videos: A Survey
- OmniGenBench: A Benchmark for Omnipotent Multimodal Generation across 50+ Tasks
- REN: Fast and Efficient Region Encodings from Patch-Based Image Encoders
- Semantic segmentation with reward
- Sketchy Bounding-box Supervision for 3D Instance Segmentation
- Grounding Agentic VLMs with Dedicated Segmentation for Fine-Grained Vehicle Damage Assessment
- Native Segmentation Vision Transformers
- DF3: World Modeling via Decoder-Free Feature Forecasting in Autonomous Navigation
- A Methodology to Evaluate Strategies Predicting Rankings on Unseen Domains
- Advancing Marine Research: UWSAM Framework and UIIS10K Dataset for Precise Underwater Instance Segmentation
- Multi-View Projection for Unsupervised Domain Adaptation in 3D Semantic Segmentation
- RAZER: Robust Accelerated Zero-Shot 3D Open-Vocabulary Panoptic Reconstruction with Spatio-Temporal Aggregation
- gen2seg: Generative Models Enable Generalizable Instance Segmentation
- Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels
- Generalizable Multispectral Land Cover Classification via Frequency-Aware Mixture of Low-Rank Token Experts
- InstanceBEV: Unifying Instance and BEV Representation for 3D Panoptic Segmentation
- Industrial Synthetic Segment Pre-training
- Is Semantic SLAM Ready for Embedded Systems ? A Comparative Survey
- Mask-Based Priors Are More Persistent than Query-Key Initializations
- Pseudo-Label Quality Decoupling and Correction for Semi-Supervised Instance Segmentation
- Seeing Sound, Hearing Sight: Uncovering Modality Bias and Conflict of AI models in Sound Localization
- StoryReasoning Dataset: Using Chain-of-Thought for Scene Understanding and Grounded Story Generation
- Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving
- MESSI: A Multi-Elevation Semantic Segmentation Image Dataset of an Urban Environment
- Technical Report for ICRA 2025 GOOSE 2D Semantic Segmentation Challenge: Leveraging Color Shift Correction, RoPE-Swin Backbone, and Quantile-based Label Denoising Strategy for Robust Outdoor Scene Understanding
- UnfoldIR: Rethinking Deep Unfolding Network in Illumination Degradation Image Restoration
- Visual Affordance Prediction: Survey and Reproducibility
- Split Matching for Inductive Zero-shot Semantic Segmentation
- Joint Super-Resolution and Segmentation for 1-m Impervious Surface Area Mapping in China's Yangtze River Economic Belt
- Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation
- Spotting the Unexpected (STU): A 3D LiDAR Dataset for Anomaly Segmentation in Autonomous Driving
- Adversarial Robustness of Deep Learning Models for Inland Water Body Segmentation from SAR Images
- VSC: Visual Search Compositional Text-to-Image Diffusion Model
- Global Collinearity-aware Polygonizer for Polygonal Building Mapping in Remote Sensing
- Vision Transformers and Graph Neural Networks for Charged Particle Tracking in the ATLAS Muon Spectrometer
- URBAN-SPIN: A street-level bikeability index to inform design implementations in historical city centres
- Vision as Unified Multimodal Generation
- Assessing and mapping public visual perception across urban public space typologies based on geotagged social media images
- Vision Transformers Need More Than Registers
- Mcity Data Engine: Iterative Model Improvement Through Open-Vocabulary Data Selection
- OG-HFYOLO :Orientation gradient guidance and heterogeneous feature fusion for deformation table cell instance segmentation
- Large-scale visual SLAM for in-the-wild videos
- Learning Streaming Video Representation via Multitask Training
- Foundation Model-Driven Framework for Human-Object Interaction Prediction with Segmentation Mask Integration
- Open-set Anomaly Segmentation in Complex Scenarios
- BARIS: Boundary-Aware Refinement with Environmental Degradation Priors for Robust Underwater Instance Segmentation
- PhenoAssistant: A Conversational Multi-Agent AI System for Automated Plant Phenotyping
- Segmenting Objectiveness and Task-awareness Unknown Region for Autonomous Driving
- PAD: Phase-Amplitude Decoupling Fusion for Multi-Modal Land Cover Classification
- CARL: Camera-Agnostic Representation Learning for Spectral Image Analysis
- Toward Fully Autonomous Driving: AI, Challenges, Opportunities, and Needs
- Network-Adaptive Cloud Processing for Visual Neuroprostheses
- The Semantic Lifecycle in Embodied AI: Acquisition, Representation and Storage via Foundation Models
- iFAN: Inference-Aware Learning for Plain Mask Transformers
- Hear to See: Discerning Stateful Listening for Audio-Visual Instance Segmentation
- Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding
- Enhancing rice breeding efficiency through semi-supervised detection and segmentation of panicles and leaves
- What is the Added Value of UDA in the VFM Era?
- LaRI: Layered Ray Intersections for Single-view 3D Geometric Reasoning
- StaticSegFormer: An Efficient High-Performance Semantic Segmentation Based on Static Structured Pruning
- From Transparent Labware Segmentation to Collision Avoidance: A Real-Time Edge-Aware Perception Pipeline
- CoMBO: Conflict Mitigation via Branched Optimization for Class Incremental Segmentation
- S4M: Boosting Semi-Supervised Instance Segmentation with SAM
- Texture2LoD3: Enabling LoD3 Building Reconstruction With Panoramic Images
- Beyond Anonymization: Object Scrubbing for Privacy-Preserving 2D and 3D Vision Tasks
- DreamO: A Unified Framework for Image Customization
- UINO-FSS: Unifying Representation Learning and Few-shot Segmentation via Hierarchical Distillation and Mamba-HyperCorrelation
- EmoSEM: Segment and Explain Emotion Stimuli in Visual Art
- LOOPE: Learnable Optimal Patch Order in Positional Embeddings for Vision Transformers
- Occlusion-Ordered Semantic Instance Segmentation
- LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models
- Fighting Fires from Space: Leveraging Vision Transformers for Enhanced Wildfire Detection and Characterization
- Multiscale Tensor Summation Factorization as a New Neural Network Layer (MTS Layer) for Multidimensional Data Processing
- View2CAD: Reconstructing View-Centric CAD Models from Single RGB-D Scans
- Towards Learning to Complete Anything in Lidar
- A Complex-valued SAR Foundation Model Based on Physically Inspired Representation Learning
- EgoExo-Gen: Ego-centric Video Prediction by Watching Exo-centric Videos
- SCI-CLIP: Segment-Centric Inference with Reference Memory for Training-Free Open-Vocabulary Segmentation
- DocSAM: Unified Document Image Segmentation via Query Decomposition and Heterogeneous Mixed Learning
- SEED: Simple ViT and Evolving Harness for Explainable Text Forgery Detection
- FLOSS: Free Lunch in Open-vocabulary Semantic Segmentation
- M2S-RoAD: Multi-Modal Semantic Segmentation for Road Damage Using Camera and LiDAR Data
- Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding
- SegEarth-R1: Geospatial Pixel Reasoning via Large Language Model
- TextSplat: Text-Guided Semantic Fusion for Generalizable Gaussian Splatting
- AerOSeg: Harnessing SAM for Open-Vocabulary Segmentation in Remote Sensing Images
- Embodied Image Captioning: Self-supervised Learning Agents for Spatially Coherent Image Descriptions
- LMM4LMM: Benchmarking and Evaluating Large-multimodal Image Generation with LMMs
- FMLGS: Fast Multilevel Language Embedded Gaussians for Part-level Interactive Agents
- Hypergraph Vision Transformers: Images are More than Nodes, More than Edges
- DGOcc: Depth-aware Global Query-based Network for Monocular 3D Occupancy Prediction
- Domain Generalization through Attenuation of Domain-Specific Information
- GraspClutter6D: A Large-scale Real-world Dataset for Robust Perception and Grasping in Cluttered Scenes
- TMT: Cross-domain Semantic Segmentation with Region-adaptive Transferability Estimation
- Earth-Adapter: Bridge the Geospatial Domain Gaps with Mixture of Frequency Adaptation
- Prior2Former -- Evidential Modeling of Mask Transformers for Assumption-Free Open-World Panoptic Segmentation
- DFormerv2: Geometry Self-Attention for RGBD Semantic Segmentation
- BoxSeg: Quality-Aware and Peer-Assisted Learning for Box-supervised Instance Segmentation
Related