Semantic Understanding of Scenes Through the ADE20K Dataset
2018/12/07 by Bolei Zhou, Hang Zhao, Xavier Puig +4 · 200 citations
Computer Science · #Advanced Image and Video Retrieval Techniques #Advanced Neural Network Applications #Multimodal Machine Learning Applications
paper · pdf · doi:10.1007/s11263-018-1140-0
openalex publication_date 2018/12/07 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/31
Citations
Cited by
- SAM-MI: A Mask-Injected Framework for Enhancing Open-Vocabulary Semantic Segmentation with SAM
- Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous Driving
- Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
- Target Refocusing via Attention Redistribution for Open-Vocabulary Semantic Segmentation: An Explainability Perspective
- InfoCLIP: Bridging Vision-Language Pretraining and Open-Vocabulary Semantic Segmentation via Information-Theoretic Alignment Transfer
- Upsample Anything: A Simple and Hard to Beat Baseline for Feature Upsampling
- MaskAnyNet: Rethinking Masked Image Regions as Valuable Information in Supervised Learning
- NERVE: Neighbourhood & Entropy-guided Random-walk for training free open-Vocabulary sEgmentation
- Another BRIXEL in the Wall: Towards Cheaper Dense Features
- CPO: Condition Preference Optimization for Controllable Image Generation
- Differentiable Hierarchical Visual Tokenization
- Predicting Household Water Consumption Using Satellite and Street View Images in Two Indian Cities
- Unveiling the Spatial-temporal Effective Receptive Fields of Spiking Neural Networks
- Accelerating Vision Transformers with Adaptive Patch Sizes
- Segmentation as A Plug-and-Play Capability for Frozen Multimodal LLMs
- Latent Diffusion Model without Variational Autoencoder
- Neuro-Symbolic Spatial Reasoning in Segmentation
- Exploring Image Representation with Decoupled Classical Visual Descriptors
- AnyUp: Universal Feature Upsampling
- UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and Generation
- Unified Open-World Segmentation with Multi-Modal Prompts
- SaFiRe: Saccade-Fixation Reiteration with Mamba for Referring Image Segmentation
- Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and Grounding
- Datasets for Valence and Arousal Inference: A Survey
- GeoSURGE: Geo-localization using Semantic Fusion with Hierarchy of Geographic Embeddings
- Cutting the Skip: Training Residual-Free Transformers
- Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training
- Multi-modal Data Spectrum: Multi-modal Datasets are Multi-dimensional
- UniVid: Unifying Vision Tasks with Pre-trained Video Generation Models
- SiNGER: A Clearer Voice Distills Vision Transformers Further
- Edge Prediction for Roof Wireframe Reconstruction with Transformers
- Towards biologically plausible phosphene simulation for the differentiable optimization of visual cortical prostheses
- 3D-Aware Vision-Language Models Fine-Tuning with Geometric Distillation
- Learning to Detect Label Errors by Making Them: A Method for Segmentation and Object Detection Datasets
- BEiT: BERT Pre-Training of Image Transformers
- Seg4Diff: Unveiling Open-Vocabulary Segmentation in Text-to-Image Diffusion Transformers
- Towards Interpretable and Efficient Attention: Compressing All by Contracting a Few
- SPATIALGEN: Layout-guided 3D Indoor Scene Generation
- Survey on semantic segmentation using deep learning techniques
- From Pixels to Urban Policy-Intelligence: Recovering Legacy Effects of Redlining with a Multimodal LLM
- Lost in Translation? Vocabulary Alignment for Source-Free Adaptation in Open-Vocabulary Semantic Segmentation
- Improving Generalized Visual Grounding with Instance-aware Joint Learning
- Gaussian Alignment for Relative Camera Pose Estimation via Single-View Reconstruction
- SAGA: Selective Adaptive Gating for Efficient and Expressive Linear Attention
- Progressively Diffused Networks for Semantic Image Segmentation
- Text4Seg++: Advancing Image Segmentation via Generative Language Modeling
- CARDIE: clustering algorithm on relevant descriptors for image enhancement
- PanopticFusion: Online Volumetric Semantic Mapping at the Level of Stuff and Things
- Towards Open World Detection: A Survey
- Beyond Self-attention: External Attention using Two Linear Layers for Visual Tasks
- ViTAE: Vision Transformer Advanced by Exploring Intrinsic Inductive Bias
- EVA-02: A visual representation for neon genesis
- Novel Category Discovery with X-Agent Attention for Open-Vocabulary Semantic Segmentation
- CAD2DMD-SET: Synthetic Generation Tool of Digital Measurement Device CAD Model Datasets for fine-tuning Large Vision-Language Models
- E-ConvNeXt: A Lightweight and Efficient ConvNeXt Variant with Cross-Stage Partial Connections
- Plug-in Feedback Self-adaptive Attention in CLIP for Training-free Open-Vocabulary Segmentation
- Harnessing Meta-Learning for Controllable Full-Frame Video Stabilization
- FOCUS: Frequency-Optimized Conditioning of DiffUSion Models for mitigating catastrophic forgetting during Test-Time Adaptation
- ICNet for Real-Time Semantic Segmentation on High-Resolution Images
- Switchable Whitening for Deep Representation Learning
- DatasetGAN: Efficient Labeled Data Factory with Minimal Human Effort
- UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic Encoding
- Generalized Decoupled Learning for Enhancing Open-Vocabulary Dense Perception
- EvTurb: Event Camera Guided Turbulence Removal
- Stable Diffusion Models are Secretly Good at Visual In-Context Learning
- Evaluating visual experience in heritage-rich areas: A case study of the Inner Ring Road of West Lake, Hangzhou
- UniConvNet: Expanding Effective Receptive Field while Maintaining Asymptotically Gaussian Distribution for ConvNets of Any Scale
- 3DFroMLLM: 3D Prototype Generation only from Pretrained Multimodal LLMs
- Spatial-Separated Curve Rendering Network for Efficient and High-Resolution Image Harmonization
- Affective Image Content Analysis: Two Decades Review and New Perspectives
- Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation
- SynSeg: Feature Synergy for Multi-Category Contrastive Learning in End-to-End Open-Vocabulary Semantic Segmentation
- Adapting Vision-Language Models Without Labels: A Comprehensive Survey
- SPEX: A Vision-Language Model for Land Cover Extraction on Spectral Remote Sensing Images
- Discovering and using Spelke segments
- Context-based Motion Retrieval using Open Vocabulary Methods for Autonomous Driving
- Training-Free Class Purification for Open-Vocabulary Semantic Segmentation
- Self-Guided Masked Autoencoder
- Improving the HardNet Descriptor
- Implicit Integration of Superpixel Segmentation into Fully Convolutional Networks
- Iwin Transformer: Hierarchical Vision Transformer using Interleaved Windows
- Domain2Vec: Domain Embedding for Unsupervised Domain Adaptation
- Indoor hierarchy relation graph construction method based on RGB‐D
- S-BEV: Semantic Birds-Eye View Representation for Weather and Lighting Invariant 3-DoF Localization
- A Picture May Be Worth a Hundred Words for Visual Question Answering
- Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling
- Advancing Visual Large Language Model for Multi-granular Versatile Perception
- Multi-dataset Pretraining: A Unified Model for Semantic Segmentation
- Compact Vision Transformer by Reduction of Kernel Complexity
- Scale-aware Insertion of Virtual Objects in Monocular Videos
- Leveraging Depth and Language for Open-Vocabulary Domain-Generalized Semantic Segmentation
- Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation
- All Tokens Matter: Token Labeling for Training Better Vision Transformers
- Leveraging Outdoor Webcams for Local Descriptor Learning
- Look at here : Utilizing supervision to attend subtle key regions
- Development of an Autonomous Mobile Robotic System for Efficient and Precise Disinfection
- Latent Diffusion Models with Masked AutoEncoders
- When Schrödinger Bridge Meets Real-World Image Dehazing with Unpaired Training
- Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation
- Domain Adaptation for Semantic Segmentation via Class-Balanced Self-Training
- MUXConv: Information Multiplexing in Convolutional Neural Networks
- Vision-Language Models Can't See the Obvious
- Medical Datasets Collections for Artificial Intelligence-based Medical Image Analysis
- From Pixels to Damage Severity: Estimating Earthquake Impacts Using Semantic Segmentation of Social Media Images
- High-Fidelity Differential-information Driven Binary Vision Transformer
- Depth Anything at Any Condition
- Perception-Oriented Latent Coding for High-Performance Compressed Domain Semantic Inference
- Learning Dense Feature Matching via Lifting Single 2D Image to 3D Space
- GameTileNet: A Semantic Dataset for Low-Resolution Game Art in Procedural Content Generation
- CARAFE++: Unified Content-Aware ReAssembly of FEatures
- ReME: A Data-Centric Framework for Training-Free Open-Vocabulary Segmentation
- Norm×Direction: Restoring the Missing Query Norm in Vision Linear Attention
- Sparse Spatial Attention Network for Semantic Segmentation
- Trajectory Prediction in Dynamic Object Tracking: A Critical Study
- Synthetic Visual Genome
- A Scaling Law for Synthetic-to-Real Transfer: How Much Is Your Pre-training Effective?
- Rethinking Semantic Segmentation from a Sequence-to-Sequence Perspective with Transformers
- LBMamba: Locally Bi-directional Mamba
- Rethinking Semantic Segmentation Evaluation for Explainability and Model Selection
- From Generation to Generalization: Emergent Few-Shot Learning in Video Diffusion Models
- Leader360V: The Large-scale, Real-world 360 Video Dataset for Multi-task Learning in Diverse Environment
- Interactive video retrieval in the age of effective joint embedding deep models: lessons from the 11th VBS
- Revisiting Transformers with Insights from Image Filtering and Boosting
- K-Net: Towards Unified Image Segmentation
- ELBO-T2IAlign: A Generic ELBO-Based Method for Calibrating Pixel-level Text-Image Alignment in Diffusion Models
- Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration
- Hyb-KAN ViT: Hybrid Kolmogorov-Arnold Networks Augmented Vision Transformer
- Histo-Miner: Deep Learning based Tissue Features Extraction Pipeline from H&E Whole Slide Images of Cutaneous Squamous Cell Carcinoma
- Are Synthetic Corruptions A Reliable Proxy For Real-World Corruptions?
- DeCLIP: Decoupled Learning for Open-Vocabulary Dense Perception
- OV-COAST: Cost Aggregation with Optimal Transport for Open-Vocabulary Semantic Segmentation
- AetherVision-Bench: An Open-Vocabulary RGB-Infrared Benchmark for Multi-Angle Segmentation across Aerial and Ground Perspectives
- ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads
- BargainNet: Background-Guided Domain Translation for Image Harmonization
- Attacking Attention of Foundation Models Disrupts Downstream Tasks
- MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping
- SAB3R: Semantic-Augmented Backbone in 3D Reconstruction
- Re-distributing Biased Pseudo Labels for Semi-supervised Semantic Segmentation: A Baseline Investigation
- un2CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP
- MMAFFBen: A Multilingual and Multimodal Affective Analysis Benchmark for Evaluating LLMs and VLMs
- TextRegion: Text-Aligned Region Tokens from Frozen Image-Text Models
- Towards urban scenes understanding through polarization cues
- DA-VPT: Semantic-Guided Visual Prompt Tuning for Vision Transformers
- CAST: Contrastive Adaptation and Distillation for Semi-Supervised Instance Segmentation
- SIR: Self-supervised Image Rectification via Seeing the Same Scene from Multiple Different Lenses
- A Survey on Training-free Open-Vocabulary Semantic Segmentation
- SANSA: Unleashing the Hidden Semantics in SAM2 for Few-Shot Segmentation
- PARTONOMY: Large Multimodal Models with Part-Level Visual Understanding
- Vision Transformers with Self-Distilled Registers
- Fully Transformer Networks for Semantic Image Segmentation
- LlamaSeg: Image Segmentation via Autoregressive Mask Generation
- EOVSAM: Efficient Open-Vocabulary Segmentation with SAM 3 in One Pass
- Automated River Substrate Mapping From Sonar Imagery With Machine Learning
- InstructPart: Task-Oriented Part Segmentation with Instruction Reasoning
- REN: Fast and Efficient Region Encodings from Patch-Based Image Encoders
- LookWhere? Efficient Visual Recognition by Learning Where to Look and What to See from Self-Supervision
- OpenSeg-R: Improving Open-Vocabulary Segmentation via Step-by-Step Visual Reasoning
- ALTo: Adaptive-Length Tokenizer for Autoregressive Mask Generation
- Native Segmentation Vision Transformers
- Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels
- CoT-Edit: Let CoT Guide Instruction Video Editing
- FractalMamba++: Scaling Vision Mamba Across Resolutions via Hilbert Fractal Geometry
- SounDiT: Geo-Contextual Soundscape-to-Landscape Generation
- DeepLab2: A TensorFlow Library for Deep Labeling
- VA-RED2: Video Adaptive Redundancy Reduction
- Seeing Sound, Hearing Sight: Uncovering Modality Bias and Conflict of AI models in Sound Localization
- Unlocking the Full Potential of Small Data with Diverse Supervision
- UncertainSAM: Fast and Efficient Uncertainty Quantification of the Segment Anything Model
- AdaDINO: Context-Adaptive DINO-Distilled Vision Foundation Models for Efficient Open-Vocabulary Edge Inference
- Local-to-Global Self-Attention in Vision Transformers
- RWKVQuant: Quantizing the RWKV Family with Proxy Guided Hybrid of Scalar and Vector Quantization
- Neural Inverse Knitting: From Images to Manufacturing Instructions
- Image Harmonization Dataset iHarmony4: HCOCO, HAdobe5k, HFlickr, and Hday2night
- DeepPerimeter: Indoor Boundary Estimation from Posed Monocular Sequences
- Sim2Real for Self-Supervised Monocular Depth and Segmentation
- Better, Stronger, Faster, and Broader: Structured All-Mask Prediction for MLLM-Based Segmentation
- PixelUp: Zero-Shot Semantic Feature Upsampling for Fine-Grained Vision Tasks
- Is This The Right Place? Geometric-Semantic Pose Verification for Indoor Visual Localization
- Spiral RoPE: Rotate Your Rotary Positional Embeddings in the 2D Plane
- DeepLandscape: Adversarial Modeling of Landscape Video
- Environment Semantics Aided Wireless Communications: A Case Study of mmWave Beam Prediction and Blockage Prediction
- Calibrating Self-supervised Monocular Depth Estimation
- Vision Transformers Need More Than Registers
- SpaRE: Enhancing Spatial Reasoning in Vision-Language Models with Synthetic Data
- Routing the Lottery: Adaptive Subnetworks for Heterogeneous Data
- SynergyAmodal: Deocclude Anything with Text Control
- Pixel-Grounded Retrieval for Knowledgeable Large Multimodal Models
- C-RADIOv4 (Tech Report)
- Improving Open-World Object Localization by Discovering Background
- Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes
- Dense Semantic Forecasting in Video by Joint Regression of Features and Feature Motion
- SegICP: Integrated deep semantic segmentation and pose estimation
- Per-Pixel Feedback for improving Semantic Segmentation
- Evaluating the Impact of Semantic Segmentation and Pose Estimation on Dense Semantic SLAM
- Multi-Scale Tensorial Summation and Dimensional Reduction Guided Neural Network for Edge Detection
- M2GAN: A Multi-Stage Self-Attention Network for Image Rain Removal on Autonomous Vehicles
- CondNet: Conditional Classifier for Scene Segmentation
- SCI-CLIP: Segment-Centric Inference with Reference Memory for Training-Free Open-Vocabulary Segmentation
- How Infrastructure and Streetscape Shape E-Scooter Route Choice: Evidence from Washington, DC
- FLOSS: Free Lunch in Open-vocabulary Semantic Segmentation
- Hypergraph Vision Transformers: Images are More than Nodes, More than Edges
- Distilling Knowledge from Heterogeneous Architectures for Semantic Segmentation
- ColorBench: Can VLMs See and Understand the Colorful World? A Comprehensive Benchmark for Color Perception, Reasoning, and Robustness
- Studying Image Diffusion Features for Zero-Shot Video Object Segmentation
- 3DM-WeConvene: Learned Image Compression with 3D Multi-Level Wavelet-Domain Convolution and Entropy Model