Masked-attention Mask Transformer for Universal Image Segmentation
2021/12/02 by Bowen Cheng, Ishan Misra, Cheng, Bowen +7 · 323 citations
Computer Science · #Advanced Neural Network Applications #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Medical Image Segmentation Techniques
paper · pdf · doi:10.48550/arxiv.2112.01527
openalex publication_date 2021/12/02 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Image segmentation is about grouping pixels with different semantics, e.g., category or instance membership, where each choice of semantics defines a task. While only the semantics of each task differ, current research focuses on designing specialized architectures for each task. We present Masked-attention Mask Transformer (Mask2Former), a new architecture capable of addressing any image segmentation task (panoptic, instance or semantic). Its key components include masked attention, which extracts localized features by constraining cross-attention within predicted mask regions. In addition to reducing the research effort by at least three times, it outperforms the best specialized architectures by a significant margin on four popular datasets. Most notably, Mask2Former sets a new state-of-the-art for panoptic segmentation (57.8 PQ on COCO), instance segmentation (50.1 AP on COCO) and semantic segmentation (57.7 mIoU on ADE20K).
Citations
Cited by
- Retrieve and Segment: Are a Few Examples Enough to Bridge the Supervision Gap in Open-Vocabulary Segmentation?
- Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
- RoadVGGT: Road-Structure-Aware Feed-Forward Road Surface Reconstruction
- QueenVIS: Rethinking Image-Only Training for Video Instance Segmentation via Query Enrichment
- ESRVS: Extreme Semi-Supervised Retinal Vessel Segmentation with a Single Annotated Image
- DishSeg24k: A Large-Scale Benchmark for Food Segmentation with Stochastic Expert Decoding
- Room-Mediated Co-occurrence for Zero-Shot Object-Centric Semantic Navigation via Frontier Scoring
- Leak-Free Cross-Validated Stacking with Per-Architecture Calibration for Sand-Boil Segmentation in Earthen Levees
- Advancing All-Weather Building Damage Mapping to the Instance Level: Outcomes and Insights from the 2026 Bright Challenge
- UP-Fuse: Uncertainty-guided LiDAR-Camera Fusion for 3D Panoptic Segmentation
- A Lightweight Multi-Scale Attention Framework for Real-Time Spinal Endoscopic Instance Segmentation
- SegEarth-R2: Towards Comprehensive Language-guided Segmentation for Remote Sensing Images
- BiCoR-Seg: Bidirectional Co-Refinement Framework for High-Resolution Remote Sensing Image Segmentation
- Watch Closely: Mitigating Object Hallucinations in Large Vision-Language Models with Disentangled Decoding
- Multifaceted Exploration of Spatial Openness in Rental Housing: A Big Data Analysis in Tokyo's 23 Wards
- RadarGen: Automotive Radar Point Cloud Generation from Cameras
- GenEval 2: Addressing Benchmark Drift in Text-to-Image Evaluation
- Task-Oriented Data Synthesis and Control-Rectify Sampling for Remote Sensing Semantic Segmentation
- Causal-Tune: Mining Causal Factors from Vision Foundation Models for Domain Generalized Semantic Segmentation
- Using Gaussian Splats to Create High-Fidelity Facial Geometry and Texture
- PixelArena: A benchmark for Pixel-Precision Visual Intelligence
- SynthSeg-Agents: Multi-Agent Synthetic Data Generation for Zero-Shot Weakly Supervised Semantic Segmentation
- S2D: Sparse-To-Dense Keymask Distillation for Unsupervised Video Instance Segmentation
- ST-DETrack: Identity-Preserving Branch Tracking in Entangled Plant Canopies via Dual Spatiotemporal Evidence
- Tracking spatial temporal details in ultrasound long video via wavelet analysis and memory bank
- AMD-HookNet++: Evolution of AMD-HookNet with Hybrid CNN-Transformer Feature Enhancement for Glacier Calving Front Segmentation
- Unified Semantic Transformer for 3D Scene Understanding
- TorchTraceAP: A New Benchmark Dataset for Detecting Performance Anti-Patterns in Computer Vision Models
- Pancakes: Consistent Multi-Protocol Image Segmentation Across Biomedical Domains
- Learning to Generate Cross-Task Unexploitable Examples
- Test-Time Modification: Inverse Domain Transformation for Robust Perception
- MeViS: A Multi-Modal Dataset for Referring Motion Expression Video Segmentation
- Super4DR: 4D Radar-centric Self-supervised Odometry and Gaussian-based Map Optimization
- Hot Hém: Sài Gòn Giũa Cái Nóng Hông Còng Bàng -- Saigon in Unequal Heat
- From SAM to DINOv2: Towards Distilling Foundation Models to Lightweight Baselines for Generalized Polyp Segmentation
- SegEarth-OV3: Exploring SAM 3 for Open-Vocabulary Semantic Segmentation in Remote Sensing Images
- XR-DT: Extended Reality-Enhanced Digital Twin for Safe Motion Planning via Human-Aware Model Predictive Path Integral Control
- Generalization vs. Specialization: Evaluating Segment Anything Model (SAM3) Zero-Shot Segmentation Against Fine-Tuned YOLO Detectors
- LiDAS: Lighting-driven Dynamic Active Sensing for Nighttime Perception
- Structure-Aware Feature Rectification with Region Adjacency Graphs for Training-Free Open-Vocabulary Semantic Segmentation
- Enhancing Urban Sensing Utility with Sensor-enabled Vehicles and Easily Accessible Data
- Balanced Learning for Domain Adaptive Semantic Segmentation
- Towards Robust Pseudo-Label Learning in Semantic Segmentation: An Encoding Perspective
- Boosting Unsupervised Video Instance Segmentation with Automatic Quality-Guided Self-Training
- Pseudo-Label Refinement for Robust Wheat Head Segmentation via Two-Stage Hybrid Training
- Are AI-Generated Driving Videos Ready for Autonomous Driving? A Diagnostic Evaluation Framework
- See in Depth: Training-Free Surgical Scene Segmentation with Monocular Depth Priors
- Performance Evaluation of Deep Learning for Tree Branch Segmentation in Autonomous Forestry Systems
- The SAM2-to-SAM3 Gap in the Segment Anything Model Family: Why Prompt-Based Expertise Fails in Concept-Driven Image Segmentation
- Exploiting Domain Properties in Language-Driven Domain Generalization for Semantic Segmentation
- NAS-LoRA: Empowering Parameter-Efficient Fine-Tuning for Visual Foundation Models with Searchable Adaptation
- Flexible Gravitational-Wave Parameter Estimation with Transformers
- Rethinking Surgical Smoke: A Smoke-Type-Aware Laparoscopic Video Desmoking Method and Dataset
- ESACT: An End-to-End Sparse Accelerator for Compute-Intensive Transformers via Local Similarity
- AirSim360: A Panoramic Simulation Platform within Drone View
- SGDiff: Scene Graph Guided Diffusion Model for Image Collaborative SegCaptioning
- ELVIS: Enhance Low-Light for Video Instance Segmentation in the Dark
- Language-Guided Open-World Anomaly Segmentation
- Panda: Self-distillation of Reusable Sensor-level Representations for High Energy Physics
- FOM-Nav: Frontier-Object Maps for Object Goal Navigation
- VFM-ISRefiner: Towards Better Adapting Vision Foundation Models for Interactive Segmentation of Remote Sensing Images
- DEAL-300K: Diffusion-based Editing Area Localization with a 300K-Scale Dataset and Frequency-Prompted Baseline
- UniGeoSeg: Towards Unified Open-World Segmentation for Geospatial Scenes
- Semantic-Centric Alignment for Zero-shot Panoptic Segmentation with Limited Data
- UMind-VL: A Generalist Ultrasound Vision-Language Model for Unified Grounded Perception and Comprehensive Interpretation
- Autonomous labeling of surgical resection margins using a foundation model
- Progress by Pieces: Test-Time Scaling for Autoregressive Image Generation
- Open Vocabulary Compositional Explanations for Neuron Alignment
- V2-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object Correspondence
- CrossEarth-Gate: Fisher-Guided Adaptive Tuning Engine for Efficient Adaptation of Cross-Domain Remote Sensing Semantic Segmentation
- SAM3-Adapter: Efficient Adaptation of Segment Anything 3 for Camouflage Object Segmentation, Shadow Detection, and Medical Image Segmentation
- Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous Driving
- PhysDNet: Physics-Guided Decomposition Network of Side-Scan Sonar Imagery
- Semantic Prioritization in Visual Counterfactual Explanations with Weighted Segmentation and Auto-Adaptive Region Selection
- Lightweight Transformer Framework for Weakly Supervised Semantic Segmentation
- Illustrator's Depth: Monocular Layer Index Prediction for Image Decomposition
- MobileOcc: A Human-Aware Semantic Occupancy Dataset for Mobile Robots
- VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning
- SVG360: Editable Multiview Vector Graphics from a Single SVG
- PairHuman: A High-Fidelity Photographic Dataset for Customized Dual-Person Generation
- InfoCLIP: Bridging Vision-Language Pretraining and Open-Vocabulary Semantic Segmentation via Information-Theoretic Alignment Transfer
- Click2Graph: Interactive Panoptic Video Scene Graphs from a Single Click
- Unsupervised Image Classification with Adaptive Nearest Neighbor Selection and Cluster Ensembles
- MaskMed: Decoupled Mask and Class Prediction for Medical Image Segmentation
- WarNav: An Autonomous Driving Benchmark for Segmentation of Navigable Zones in War Scenes
- The changing surface of the world's roads
- FGNet: Leveraging Feature-Guided Attention to Refine SAM2 for 3D EM Neuron Segmentation
- Fine-Grained Representation for Lane Topology Reasoning
- Seg-VAR: Image Segmentation with Visual Autoregressive Modeling
- Empowering DINO Representations for Underwater Instance Segmentation via Aligner and Prompter
- SkelSplat: Robust Multi-view 3D Human Pose Estimation with Differentiable Gaussian Rendering
- Navigating the Wild: Pareto-Optimal Visual Decision-Making in Image Space
- Visual Bridge: Universal Visual Perception Representations Generating
- Relative Energy Learning for LiDAR Out-of-Distribution Detection
- Leveraging Text-Driven Semantic Variation for Robust OOD Segmentation
- EIDSeg: A Pixel-Level Semantic Segmentation Dataset for Post-Earthquake Damage Assessment from Social Media Images
- Polymap: generating high definition map based on rasterized polygons
- From Words to Safety: Language-Conditioned Safety Filtering for Robot Navigation
- Another BRIXEL in the Wall: Towards Cheaper Dense Features
- No Pose Estimation? No Problem: Pose-Agnostic and Instance-Aware Test-Time Adaptation for Monocular Depth Estimation
- OneOcc: Semantic Occupancy Prediction for Legged Robots with a Single Panoramic Camera
- MIQ-SAM3D: From Single-Point Prompt to Multi-Instance Segmentation via Competitive Query Refinement
- Differentiable Hierarchical Visual Tokenization
- Grounding Surgical Action Triplets with Instrument Instance Segmentation: A Dataset and Target-Aware Fusion Approach
- EPARA: Parallelizing Categorized AI Inference in Edge Clouds
- BeetleFlow: An Integrative Deep Learning Pipeline for Beetle Image Processing
- Generative Semantic Coding for Ultra-Low Bitrate Visual Communication and Analysis
- NaviTrace: Evaluating Embodied Navigation of Vision-Language Models
- Revisiting Generative Infrared and Visible Image Fusion Based on Human Cognitive Laws
- DVPSFormer: Efficient Online Depth-aware Video Panoptic Segmentation for Autonomous Driving
- Cataract-LMM: Large-Scale, Multi-Source, Multi-Task Benchmark for Deep Learning in Surgical Video Analysis
- LangHOPS: Language Grounded Hierarchical Open-Vocabulary Part Segmentation
- Region-CAM: Towards Accurate Object Regions in Class Activation Maps for Weakly Supervised Learning Tasks
- AtlasGS: Atlanta-world Guided Surface Reconstruction with Implicit Structured Gaussians
- IGGT: Instance-Grounded Geometry Transformer for Semantic 3D Reconstruction
- SRSR: Enhancing Semantic Accuracy in Real-World Image Super-Resolution with Spatially Re-Focused Text-Conditioning
- Simplifying Knowledge Transfer in Pretrained Models
- AURASeg: Attention Guided Upsampling with Residual Boundary-Assistive Refinement for Drivable-Area Segmentation
- Dynamic Semantic-Aware Correlation Modeling for UAV Tracking
- Controllable-LPMoE: Adapting to Challenging Object Segmentation via Dynamic Local Priors from Mixture-of-Experts
- ARGenSeg: Image Segmentation with Autoregressive Image Generation Model
- SFGFusion: Surface Fitting Guided 3D Object Detection with 4D Radar and Camera Fusion
- Symmetric Entropy-Constrained Video Coding for Machines
- Self-Supervised Learning to Fly using Efficient Semantic Segmentation and Metric Depth Estimation for Low-Cost Autonomous UAVs
- Aria Gen 2 Pilot Dataset
- MOBIUS: Big-to-Mobile Universal Instance Segmentation via Multi-modal Bottleneck Fusion and Calibrated Decoder Pruning
- CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
- Multi-modal video data-pipelines for machine learning with minimal human supervision
- EuroMineNet: A Multitemporal Sentinel-2 Benchmark for Spatiotemporal Mining Footprint Analysis in the European Union (2015-2024)
- Towards Generalist Intelligence in Dentistry: Vision Foundation Models for Oral and Maxillofacial Radiology
- UrbanVerse: Scaling Urban Simulation by Watching City-Tour Videos
- UniVector: Unified Vector Extraction via Instance-Geometry Interaction
- Transformer-based Scalable Beamforming Optimization via Deep Residual Learning
- UniFusion: Vision-Language Model as Unified Encoder in Image Generation
- Generalisation of automatic tumour segmentation in histopathological whole-slide images across multiple cancer types
- MSCloudCAM: Multi-Scale Context Adaptation with Convolutional Cross-Attention for Multispectral Cloud Segmentation
- A Machine Learning Perspective on Automated Driving Corner Cases
- Unified Open-World Segmentation with Multi-Modal Prompts
- FRIEREN: Federated Learning with Vision-Language Regularization for Segmentation
- Probabilistic Hyper-Graphs using Multiple Randomly Masked Autoencoders for Semi-supervised Multi-modal Multi-task Learning
- Complementary and Contrastive Learning for Audio-Visual Segmentation
- Explainable Human-in-the-Loop Segmentation via Critic Feedback Signals
- ClustViT: Clustering-based Token Merging for Semantic Segmentation
- What Matters in RL-Based Methods for Object-Goal Navigation? An Empirical Study and A Unified Framework
- Holistic Order Prediction in Natural Scenes
- Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and Grounding
- SilvaScenes: Tree Detection and Species Classification from Under-Canopy Images in Natural Forests
- X2Video: Adapting Diffusion Models for Multimodal Controllable Neural Video Rendering
- LTCA: Long-range Temporal Context Attention for Referring Video Object Segmentation
- UniMMVSR: A Unified Multi-Modal Framework for Cascaded Video Super-Resolution
- AlignGS: Aligning Geometry and Semantics for Robust Indoor Reconstruction from Sparse Views
- Locality-Sensitive Hashing-Based Efficient Point Transformer for Charged Particle Reconstruction
- Geometry-Aware Cross Modal Alignment for Light Field-LiDAR Semantic Segmentation
- Data Factory with Minimal Human Effort Using VLMs
- Human Action Recognition from Point Clouds over Time
- From Filters to VLMs: Benchmarking Defogging Methods through Object Detection and Segmentation Performance
- UGround: Towards Unified Visual Grounding with Unrolled Transformers
- IMAGEdit: Let Any Subject Transform
- KeySG: Hierarchical Keyframe-Based 3D Scene Graphs
- Semantic Visual Simultaneous Localization and Mapping: A Survey on State of the Art, Challenges, and Future Directions
- Robust Context-Aware Object Recognition
- SAGE-LD: Towards Scalable and Generalizable End-to-End Language Diarization via Simulated Data Augmentation
- Stitch: Training-Free Position Control in Multimodal Diffusion Transformers
- IRIS: Intrinsic Reward Image Synthesis
- CLASP: Adaptive Spectral Clustering for Unsupervised Per-Image Segmentation
- CORE-3D: Context-aware Open-vocabulary Retrieval by Embeddings in 3D
- K-Prism: A Knowledge-Guided and Prompt Integrated Universal Medical Image Segmentation Model
- HieraTok: Multi-Scale Visual Tokenizer Improves Image Reconstruction and Generation
- Token Merging via Spatiotemporal Information Mining for Surgical Video Understanding
- Foundation Model-Based Adaptive Semantic Image Transmission for Dynamic Wireless Environments
- UniMapGen: A Generative Framework for Large-Scale Map Construction from Multi-modal Data
- Learning What To Hear: Boosting Sound-Source Association For Robust Audiovisual Instance Segmentation
- CubistMerge: Spatial-Preserving Token Merging For Diverse ViT Backbones
- Queryable 3D Scene Representation: A Multi-Modal Framework for Semantic Reasoning and Robotic Task Planning
- Boosting LiDAR-Based Localization with Semantic Insight: Camera Projection versus Direct LiDAR Segmentation
- ZMIS-SAM: Segment Anything Model Enhanced with Wavelet Transform for Zooplankton Microscopy Image Instance Segmentation
- IGME: Efficient Chained Method Ensemble for Transferable Semantic Segmentation Attacks
- MLF-4DRCNet: Multi-Level Fusion with 4D Radar and Camera for 3D Object Detection in Autonomous Driving
- Frequency-Domain Decomposition and Recomposition for Robust Audio-Visual Segmentation
- Surgical Video Understanding with Label Interpolation
- Visual Instruction Pretraining for Domain-Specific Foundation Models
- Region-Aware Deformable Convolutions
- White Aggregation and Restoration for Few-shot 3D Point Cloud Semantic Segmentation
- Masked Feature Modeling Enhances Adaptive Segmentation
- Improving Generalized Visual Grounding with Instance-aware Joint Learning
- Re-purposing SAM into Efficient Visual Projectors for MLLM-Based Referring Image Segmentation
- Mitigating Query Selection Bias in Referring Video Object Segmentation
- First Place Solution to the MLCAS 2025 GWFSS Challenge: The Devil is in the Detail and Minority
- NavMoE: Hybrid Model- and Learning-based Traversability Estimation for Local Navigation via Mixture of Experts
- Road Obstacle Video Segmentation
- RailSafeNet: Visual Scene Understanding for Tram Safety
- Exploring Efficient Open-Vocabulary Segmentation in the Remote Sensing
- CLAIRE: A Dual Encoder Network with RIFT Loss and Phi-3 Small Language Model Based Interpretability for Cross-Modality Synthetic Aperture Radar and Optical Land Cover Segmentation
- Microsurgical Instrument Segmentation for Robot-Assisted Surgery
- Multimodal SAM-adapter for Semantic Segmentation
- I-Segmenter: Integer-Only Vision Transformer for Efficient Semantic Segmentation
- Leveraging Multi-View Weak Supervision for Occlusion-Aware Multi-Human Parsing
- DGFusion: Depth-Guided Sensor Fusion for Robust Semantic Perception
- Prompt-Driven Image Analysis with Multimodal Generative AI: Detection, Segmentation, Inpainting, and Interpretation
- UNO: Unifying One-stage Video Scene Graph Generation via Object-Centric Visual Representation Learning
- A biologically inspired separable learning vision model for real-time traffic object perception in Dark
- A Multidimensional AI-powered Framework for Analyzing Tourist Perception in Historic Urban Quarters: A Case Study in Shanghai
- Skywork UniPic 2.0: Building Kontext Model with Online RL for Unified Multimodal Model
- HAMSt3R: Human-Aware Multi-view Stereo 3D Reconstruction
- Unsupervised Instance Segmentation with Superpixels
- InstaDA: Augmenting Instance Segmentation Data with Dual-Agent System
- SOPSeg: Prompt-based Small Object Instance Segmentation in Remote Sensing Imagery
- EM3M: An Electron Micrograph Dataset for Microstructural Segmentation and Generation
- An Investigation of Visual Foundation Models Robustness
- MedDINOv3: How to adapt vision foundation models for medical image segmentation?
- RAGSR: Regional Attention Guided Diffusion for Image Super-Resolution
- No More Sibling Rivalry: Debiasing Human-Object Interaction Detection
- Can General-Purpose Omnimodels Compete with Specialists? A Case Study in Medical Image Segmentation
- DGL-RSIS: Decoupling Global Spatial Context and Local Class Semantics for Training-Free Remote Sensing Image Segmentation
- Learning Yourself: Class-Incremental Semantic Segmentation with Language-Inspired Bootstrapped Disentanglement
- Representation Learning with Adaptive Superpixel Coding
- Inference-Time Alignment Control for Diffusion Models with Reinforcement Learning Guidance
- Webly-Supervised Image Manipulation Localization via Category-Aware Auto-Annotation
- GLOW: A Unified Particle Flow Transformer
- Image Quality Assessment for Machines: Paradigm, Large-scale Database, and Models
- AutoQ-VIS: Improving Unsupervised Video Instance Segmentation via Automatic Quality Assessment
- FreeVPS: Repurposing Training-Free SAM2 for Generalizable Video Polyp Segmentation
- IELDG: Suppressing Domain-Specific Noise with Inverse Evolution Layers for Domain Generalized Semantic Segmentation
- Autoregressive Universal Video Segmentation Model
- RoofSeg: An edge-aware transformer-based network for end-to-end roof plane segmentation
- PseudoMapTrainer: Learning Online Mapping without HD Maps
- CM2LoD3: Reconstructing LoD3 Building Models Using Semantic Conflict Maps
- EventSSEG: Event-driven Self-Supervised Segmentation with Probabilistic Attention
- Multiscale Video Transformers for Class Agnostic Segmentation in Autonomous Driving
- A Comprehensive Review of Agricultural Parcel and Boundary Delineation from Remote Sensing Images: Recent Progress and Future Perspectives
- MoVieDrive: Multi-Modal Multi-View Urban Scene Video Generation
- LENS: Learning to Segment Anything with Unified Reinforced Reasoning
- SIS-Challenge: Event-based Spatio-temporal Instance Segmentation Challenge at the CVPR 2025 Event-based Vision Workshop
- Temporal Grounding as a Learning Signal for Referring Video Object Segmentation
- EVTP-IVS: Effective Visual Token Pruning For Unifying Instruction Visual Segmentation In Multi-Modal Large Language Models
- OVSegDT: Segmenting Transformer for Open-Vocabulary Object Goal Navigation
- Generalized Decoupled Learning for Enhancing Open-Vocabulary Dense Perception
- CRISP: Contrastive Residual Injection and Semantic Prompting for Continual Video Instance Segmentation
- Unlocking Robust Semantic Segmentation Performance via Label-only Elastic Deformations against Implicit Label Noise
- From Pixel to Mask: A Survey of Out-of-Distribution Segmentation
- TRACE: Learning 3D Gaussian Physical Dynamics from Multi-view Videos
- FM4NPP: A Scaling Foundation Model for Nuclear and Particle Physics
- RelayFormer: A Unified Local-Global Attention Framework for Scalable Image and Video Manipulation Localization
- Revisiting Efficient Semantic Segmentation: Learning Offsets for Better Spatial and Class Feature Alignment
- SHREC 2025: Retrieval of Optimal Objects for Multi-modal Enhanced Language and Spatial Assistance (ROOMELSA)
- Hierarchical Visual Prompt Learning for Continual Video Instance Segmentation
- Planner-Refiner: Dynamic Space-Time Refinement for Vision-Language Alignment in Videos
- EventRR: Event Referential Reasoning for Referring Video Object Segmentation
- Text-guided Visual Prompt DINO for Generic Segmentation
- UGD-IML: A Unified Generative Diffusion-based Framework for Constrained and Unconstrained Image Manipulation Localization
- Learning 3D Texture-Aware Representations for Parsing Diverse Human Clothing and Body Parts
- SGDFuse: SAM-Guided Diffusion for High-Fidelity Infrared and Visible Image Fusion
- SPEX: A Vision-Language Model for Land Cover Extraction on Spectral Remote Sensing Images
- Decoupling Continual Semantic Segmentation
- Temporal Cluster Assignment for Efficient Real-Time Video Segmentation
- TSMS-SAM2: Multi-scale Temporal Sampling Augmentation and Memory-Splitting Pruning for Promptable Video Object Segmentation and Tracking in Surgical Scenarios
- X-SAM: From Segment Anything to Any Segmentation
- Two-Way Garment Transfer: Unified Diffusion Framework for Dressing and Undressing Synthesis
- A2Mamba: Attention-augmented State Space Models for Visual Recognition
- What Holds Back Open-Vocabulary Segmentation?
- TNet: Terrace Convolutional Decoder Network for Remote Sensing Image Semantic Segmentation
- DOMR: Establishing Cross-View Segmentation via Dense Object Matching
- Prototype-Driven Structure Synergy Network for Remote Sensing Images Segmentation
- MetaScope: Optics-Driven Neural Network for Ultra-Micro Metalens Endoscopy
- ParticleSAM: Small Particle Segmentation for Material Quality Monitoring in Recycling Processes
- AlignCAT: Visual-Linguistic Alignment of Category and Attribute for Weakly Supervised Visual Grounding
- Multi-Granularity Feature Calibration via VFM for Domain Generalized Semantic Segmentation
- LT-Gaussian: Long-Term Map Update Using 3D Gaussian Splatting for Autonomous Driving
- Rein++: Efficient Generalization and Adaptation for Semantic Segmentation with Vision Foundation Models
- Set Pivot Learning: Redefining Generalized Segmentation with Vision Foundation Models
- IAUNet: Instance-Aware U-Net
- RMT-PPAD: Real-time Multi-task Learning for Panoptic Perception in Autonomous Driving
- Training-Free Class Purification for Open-Vocabulary Semantic Segmentation
- SDMatte: Grafting Diffusion Models for Interactive Matting
- Multimodal Referring Segmentation: A Survey
- UIS-Mamba: Exploring Mamba for Underwater Instance Segmentation via Dynamic Tree Scan and Hidden State Weaken
- MagicRoad: Semantic-Aware 3D Road Surface Reconstruction via Obstacle Inpainting
- Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation
- Segment Anything for Video: A Comprehensive Review of Video Object Segmentation and Tracking from Past to Future
- AlphaDent: A dataset for automated tooth pathology detection
- Exploring Probabilistic Modeling Beyond Domain Generalization for Semantic Segmentation
- ObjectGS: Object-aware Scene Reconstruction and Scene Understanding via Gaussian Splatting
- Lightweight Transformer-Driven Segmentation of Hotspots and Snail Trails in Solar PV Thermal Imagery
- Conditional Video Generation for High-Efficiency Video Compression
- SeeDiff: Off-the-Shelf Seeded Mask Generation from Diffusion Models
- Latest Object Memory Management for Temporally Consistent Video Instance Segmentation
- SurgPIS: Surgical-instrument-level Instances and Part-level Semantics for Weakly-supervised Part-aware Instance Segmentation
- MixA-Q: Revisiting Activation Sparsity for Vision Transformers from a Mixed-Precision Quantization Perspective
- Synthetic Data Augmentation for Enhanced Chicken Carcass Instance Segmentation
- Privacy-Preserving Semantic Segmentation from Ultra-Low-Resolution RGB Inputs
- Advancing Complex Video Object Segmentation via Progressive Concept Construction
- Object segmentation in the wild with foundation models: application to vision assisted neuro-prostheses for upper limbs
- GVCCS : Ground Visible Camera Contrail Sequences
- GTPBD: A Fine-Grained Global Terraced Parcel and Boundary Dataset
- Moving Object Detection from Moving Camera Using Focus of Expansion Likelihood and Segmentation
- IndoorBEV: Joint Detection and Footprint Completion of Objects via Mask-based Prediction in Indoor Scenarios for Bird's-Eye View Perception
- Leveraging Pathology Foundation Models for Panoptic Segmentation of Melanoma in H&E Images
- Test-time Prompt Refinement for Text-to-Image Models
- Part Segmentation of Human Meshes via Multi-View Human Parsing
- Comparative validation of surgical phase recognition, instrument keypoint estimation, and instrument instance segmentation in endoscopy: Results of the PhaKIR 2024 challenge
- Advancing Visual Large Language Model for Multi-granular Versatile Perception
- SCORE: Scene Context Matters in Open-Vocabulary Remote Sensing Instance Segmentation
- NLI4VolVis: Natural Language Interaction for Volume Visualization via LLM Multi-Agents and Editable 3D Gaussian Splatting
- AD-GS: Object-Aware B-Spline Gaussian Splatting for Self-Supervised Autonomous Driving
- Frequency-Dynamic Attention Modulation for Dense Prediction
- From Binary to Semantic: Utilizing Large-Scale Binary Occupancy Data for 3D Semantic Occupancy Prediction
- Spatial Frequency Modulation for Semantic Segmentation
- SGLoc: Semantic Localization System for Camera Pose Estimation from 3D Gaussian Splatting Representation
- DEARLi: Decoupled Enhancement of Recognition and Localization for Semi-supervised Panoptic Segmentation
- Stereo-based 3D Anomaly Object Detection for Autonomous Driving: A New Dataset and Baseline
- Advancing Medical Image Segmentation via Self-supervised Instance-adaptive Prototype Learning
- Objectomaly: Objectness-Aware Refinement for OoD Segmentation with Structural Consistency and Boundary Precision
- Rethinking Query-based Transformer for Continual Image Segmentation
- A multi-modal dataset for insect biodiversity with imagery and DNA at the trap and individual level
- SemRaFiner: Panoptic Segmentation in Sparse and Noisy Radar Point Clouds
- What Demands Attention in Urban Street Scenes? From Scene Understanding towards Road Safety: A Survey of Vision-driven Datasets and Studies
- AnthroTAP: Learning Point Tracking with Real-World Motion
- SPADE: Spatial-Aware Denoising Network for Open-vocabulary Panoptic Scene Graph Generation with Long- and Local-range Context Reasoning
- Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation
- All in One: Visual-Description-Guided Unified Point Cloud Segmentation
- From Vision To Language through Graph of Events in Space and Time: An Explainable Self-supervised Approach
- MOSU: Autonomous Long-range Robot Navigation with Multi-modal Scene Understanding
- VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs
Related