Fully Convolutional Networks for Semantic Segmentation
2016/05/20 by Evan Shelhamer, Jonathan Long, Shelhamer, Evan +3 · 175 citations
Computer Science · #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications #cs.CV
paper · pdf · doi:10.48550/arxiv.1605.06211
to appear in PAMI (accepted May, 2016); journal edition of arXiv:1411.4038
arxiv created 2016/05/20 · openalex publication_date 2016/05/20 · arxiv updated 2016/05/23 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Convolutional networks are powerful visual models that yield hierarchies of features. We show that convolutional networks by themselves, trained end-to-end, pixels-to-pixels, improve on the previous best result in semantic segmentation. Our key insight is to build "fully convolutional" networks that take input of arbitrary size and produce correspondingly-sized output with efficient inference and learning. We define and detail the space of fully convolutional networks, explain their application to spatially dense prediction tasks, and draw connections to prior models. We adapt contemporary classification networks (AlexNet, the VGG net, and GoogLeNet) into fully convolutional networks and transfer their learned representations by fine-tuning to the segmentation task. We then define a skip architecture that combines semantic information from a deep, coarse layer with appearance information from a shallow, fine layer to produce accurate and detailed segmentations. Our fully convolutional network achieves improved segmentation of PASCAL VOC (30% relative improvement to 67.2% mean IU on 2012), NYUDv2, SIFT Flow, and PASCAL-Context, while inference takes one tenth of a second for a typical image.
Citations
Cited by
- Generalized Deep Image to Image Regression
- Amulet: Aggregating Multi-level Convolutional Features for Salient Object Detection
- PPR-FCN: Weakly Supervised Visual Relation Detection via Parallel Pairwise R-FCN
- Face Parsing via Recurrent Propagation
- Recurrent Multimodal Interaction for Referring Image Segmentation
- Generalizing to Unseen Domains via Adversarial Data Augmentation
- A Pixel-Based Framework for Data-Driven Clothing
- CDC: Convolutional-De-Convolutional Networks for Precise Temporal Action Localization in Untrimmed Videos
- CNN-based Segmentation of Medical Imaging Data
- Understand Scene Categories by Objects: A Semantic Regularized Scene Classifier Using Convolutional Neural Networks
- Deep Convolutional Features for Image Based Retrieval and Scene Categorization
- Development and validation of an interpretable deep learning framework for Alzheimer’s disease classification
- No More Discrimination: Cross City Adaptation of Road Scene Segmenters
- Learning Sparse High Dimensional Filters: Image Filtering, Dense CRFs and Bilateral Neural Networks
- Distantly Supervised Road Segmentation
- Learning Multi-Domain Convolutional Neural Networks for Visual Tracking
- Convolutional Channel Features
- Multi-style Generative Network for Real-time Transfer
- Learning to Segment Every Thing
- BPGrad: Towards Global Optimality in Deep Learning via Branch and Pruning
- Neuron-level Selective Context Aggregation for Scene Segmentation
- Weakly Supervised Object Localization Using Things and Stuff Transfer
- SegFlow: Joint Learning for Video Object Segmentation and Optical Flow
- Variable Rate Image Compression with Recurrent Neural Networks
- Learning in an Uncertain World: Representing Ambiguity Through Multiple Hypotheses
- Super-BPD: Super Boundary-to-Pixel Direction for Fast Image Segmentation
- 3D Object Reconstruction from a Single Depth View with Adversarial Learning
- SRN: Side-output Residual Network for Object Symmetry Detection in the Wild
- Attend in groups: a weakly-supervised deep learning framework for learning from web data
- Learning High Dynamic Range from Outdoor Panoramas
- End-to-End Learning of Geometry and Context for Deep Stereo Regression
- DeepMVS: Learning Multi-view Stereopsis
- Fast-SCNN: Fast Semantic Segmentation Network
- Plant diseases and pests detection based on deep learning: a review
- Deep Learning Face Attributes in the Wild
- Agile Amulet: Real-Time Salient Object Detection with Contextual Attention
- Making EfficientNet More Efficient: Exploring Batch-Independent Normalization, Group Convolutions and Reduced Resolution Training
- Domain adaptation for holistic skin detection
- Inside-Outside Net: Detecting Objects in Context with Skip Pooling and Recurrent Neural Networks
- Investigating Transfer Learning Capabilities of Vision Transformers and CNNs by Fine-Tuning a Single Trainable Block
- Data Distillation: Towards Omni-Supervised Learning
- Scene Graph Generation via Conditional Random Fields
- Split-Merge Pooling
- Significance-aware Information Bottleneck for Domain Adaptive Semantic Segmentation
- Multi-Oriented Text Detection with Fully Convolutional Networks
- Contour Detection Using Cost-Sensitive Convolutional Neural Networks
- Deep Interactive Object Selection
- Makeup like a superstar: Deep Localized Makeup Transfer Network
- SpherePHD: Applying CNNs on a Spherical PolyHeDron Representation of 360 degree Images
- Fully Connected Deep Structured Networks
- Natural Language Guided Visual Relationship Detection
- Beyond Planar Symmetry: Modeling human perception of reflection and rotation symmetries in the wild
- Analysis on DeepLabV3+ Performance for Automatic Steel Defects Detection
- FCNs in the Wild: Pixel-level Adversarial and Constraint-based Adaptation
- Context Encoding for Semantic Segmentation
- Monocular Object Instance Segmentation and Depth Ordering with CNNs
- Weakly- and Semi-Supervised Learning of a DCNN for Semantic Image Segmentation
- TorontoCity: Seeing the World with a Million Eyes
- Convolutional neural networks with low-rank regularization
- Image Completion on CIFAR-10
- Learning Neural Parsers with Deterministic Differentiable Imitation Learning
- A Localisation-Segmentation Approach for Multi-label Annotation of Lumbar Vertebrae using Deep Nets
- Deep Watershed Transform for Instance Segmentation
- Video Propagation Networks
- EPINET: A Fully-Convolutional Neural Network Using Epipolar Geometry for Depth from Light Field Images
- ProNet: Learning to Propose Object-specific Boxes for Cascaded Neural Networks
- Joint Sequence Learning and Cross-Modality Convolution for 3D Biomedical Segmentation
- Deep Variation-structured Reinforcement Learning for Visual Relationship and Attribute Detection
- LCNN: Lookup-based Convolutional Neural Network
- A 3D Coarse-to-Fine Framework for Volumetric Medical Image Segmentation
- Visual Discovery at Pinterest
- SketchParse : Towards Rich Descriptions for Poorly Drawn Sketches using Multi-Task Hierarchical Deep Networks
- Wider or Deeper: Revisiting the ResNet Model for Visual Recognition
- A Survey On 3D Inner Structure Prediction from its Outer Shape
- Detecting Small, Densely Distributed Objects with Filter-Amplifier Networks and Loss Boosting
- Visualizing and Understanding Deep Texture Representations
- A New Convolutional Network-in-Network Structure and Its Applications in Skin Detection, Semantic Segmentation, and Artifact Reduction
- Boundary-aware Instance Segmentation
- Electricity Theft Detection with self-attention
- Learning to Segment Object Candidates via Recursive Neural Networks
- ZM-Net: Real-time Zero-shot Image Manipulation Network
- Depth Assisted Full Resolution Network for Single Image-based View Synthesis
- Toward Streaming Synapse Detection with Compositional ConvNets
- Track Facial Points in Unconstrained Videos
- Adaptive Weighting Multi-Field-of-View CNN for Semantic Segmentation in Pathology
- Lattice Long Short-Term Memory for Human Action Recognition
- Recovering Realistic Texture in Image Super-resolution by Deep Spatial Feature Transform
- Zoom and Learn: Generalizing Deep Stereo Matching to Novel Domains
- Weakly Supervised Instance Segmentation using Class Peak Response
- A survey of Object Classification and Detection based on 2D/3D data
- Bidirectional Attention Network for Monocular Depth Estimation
- Few-shot Object Detection via Feature Reweighting
- Dynamic Video Segmentation Network
- Dense Captioning with Joint Inference and Visual Context
- Unsupervised learning with sparse space-and-time autoencoders
- Augmentation Inside the Network
- Deep Direct Regression for Multi-Oriented Scene Text Detection
- Learning to detect and localize many objects from few examples
- Using Cross-Model EgoSupervision to Learn Cooperative Basketball Intention
- MAttNet: Modular Attention Network for Referring Expression Comprehension
- ACE: Adapting to Changing Environments for Semantic Segmentation
- PAD-Net: Multi-Tasks Guided Prediction-and-Distillation Network for Simultaneous Depth Estimation and Scene Parsing
- ResUNet-a: a deep learning framework for semantic segmentation of remotely sensed data
- Automatic Liver and Tumor Segmentation of CT and MRI Volumes using Cascaded Fully Convolutional Neural Networks
- Object Detection, Tracking, and Motion Segmentation for Object-level Video Segmentation
- CNN based texture synthesize with Semantic segment
- A Multiscale Patch Based Convolutional Network for Brain Tumor Segmentation
- TAFE-Net: Task-Aware Feature Embeddings for Low Shot Learning
- ShiftCNN: Generalized Low-Precision Architecture for Inference of Convolutional Neural Networks
- Label-Driven Reconstruction for Domain Adaptation in Semantic Segmentation
- Auxiliary Learning for Deep Multi-task Learning
- Proposal-free Network for Instance-level Object Segmentation
- Automating Carotid Intima-Media Thickness Video Interpretation with Convolutional Neural Networks
- Learning a Discriminative Feature Network for Semantic Segmentation
- Multisource and Multitemporal Data Fusion in Remote Sensing
- Monocular 3D Human Pose Estimation In The Wild Using Improved CNN Supervision
- Learning deep structured active contours end-to-end
- FCN-rLSTM: Deep Spatio-Temporal Neural Networks for Vehicle Counting in City Cameras
- Exploiting saliency for object segmentation from image level labels
- Authoring image decompositions with generative models
- Automatic localization and decoding of honeybee markers using deep convolutional neural networks
- Detail Preserving Depth Estimation from a Single Image Using Attention Guided Networks
- Self-supervised Low Light Image Enhancement and Denoising
- Semi and Weakly Supervised Semantic Segmentation Using Generative Adversarial Network
- Colorectal Polyp Segmentation by U-Net with Dilation Convolution
- Adversarial Dropout Regularization
- Noise2Void - Learning Denoising from Single Noisy Images
- Fast Online Object Tracking and Segmentation: A Unifying Approach
- Tube-CNN: Modeling temporal evolution of appearance for object detection in video
- Unsupervised Histopathology Image Synthesis
- Instance Embedding Transfer to Unsupervised Video Object Segmentation
- Surveillance Video Parsing with Single Frame Supervision
- Triplet-based Deep Similarity Learning for Person Re-Identification
- Siamese Cascaded Region Proposal Networks for Real-Time Visual Tracking
- Deep Quantization: Encoding Convolutional Activations with Deep Generative Model
- The Effects of Image Pre- and Post-Processing, Wavelet Decomposition, and Local Binary Patterns on U-Nets for Skin Lesion Segmentation
- Temporally Folded Convolutional Neural Networks for Sequence Forecasting
- Blurring the Line Between Structure and Learning to Optimize and Adapt\n Receptive Fields
- Self-Supervised Feature Learning by Learning to Spot Artifacts
- Training Deep Networks with Structured Layers by Matrix Backpropagation
- Large-Scale 3D Scene Classification With Multi-View Volumetric CNN
- Fully Convolutional Multi-Class Multiple Instance Learning
- DSAC - Differentiable RANSAC for Camera Localization
- Target-Aware Object Discovery and Association for Unsupervised Video Multi-Object Segmentation
- Semi-supervised Semantic Segmentation with Directional Context-aware Consistency
- Learning Better Features for Face Detection with Feature Fusion and Segmentation Supervision
- ArbiText: Arbitrary-Oriented Text Detection in Unconstrained Scene
- Semantic Image Segmentation with Deep Convolutional Nets and Fully Connected CRFs
- Attention Gated Networks: Learning to Leverage Salient Regions in Medical Images
- Learning to cluster in order to transfer across domains and tasks
- Deeply-Recursive Convolutional Network for Image Super-Resolution
- Vehicle Image Generation Going Well with The Surroundings
- Explainable Semantic Mapping for First Responders
- Learning High-level Prior with Convolutional Neural Networks for Semantic Segmentation
- PixelNet: Representation of the pixels, by the pixels, and for the pixels
- Full-Resolution Residual Networks for Semantic Segmentation in Street Scenes
- Dilated Residual Networks
- Dual Encoder Fusion U-Net (DEFU-Net) for Cross-manufacturer Chest X-ray Segmentation
- Deep Learning for Semantic Part Segmentation with High-Level Guidance
- PointSIFT: A SIFT-like Network Module for 3D Point Cloud Semantic Segmentation
- StuffNet: Using 'Stuff' to Improve Object Detection
- Artificial Intelligence in Tumor Subregion Analysis Based on Medical Imaging: A Review
- MonoCap: Monocular Human Motion Capture using a CNN Coupled with a Geometric Prior
- Enhancing Cross-task Black-Box Transferability of Adversarial Examples with Dispersion Reduction
- Semantic Segmentation for Partially Occluded Apple Trees Based on Deep Learning
- SPGNet: Semantic Prediction Guidance for Scene Parsing
- Pushing the Boundaries of Boundary Detection using Deep Learning
- Deep Learning for Time-Series Analysis
- Point and Ask: Incorporating Pointing into Visual Question Answering
- What Can Help Pedestrian Detection?
- Deep Pyramidal Residual Networks
- Good Practice in CNN Feature Transfer
- Super-Resolution with Deep Adaptive Image Resampling
- A Unified Framework for Generalizable Style Transfer: Style and Content Separation
- Deep Layer Aggregation
- Rethinking ImageNet Pre-training
- RethNet: Object-by-Object Learning for Detecting Facial Skin Problems
- A Study on Trees's Knots Prediction from their Bark Outer-Shape
- Semantic Object Parsing with Graph LSTM
- Learning to Segment Instances in Videos with Spatial Propagation Network
- Unsupervised Deep Multi-focus Image Fusion
- Learning a Discriminative Prior for Blind Image Deblurring
- FISHING Net: Future Inference of Semantic Heatmaps In Grids
- Scene Understanding Networks for Autonomous Driving based on Around View Monitoring System
- Error Correction for Dense Semantic Image Labeling
- Perceive Where to Focus: Learning Visibility-aware Part-level Features for Partial Person Re-identification
- Photographic Text-to-Image Synthesis with a Hierarchically-nested Adversarial Network
- Cross-Domain Self-supervised Multi-task Feature Learning using Synthetic Imagery
- EvalAI: Towards Better Evaluation Systems for AI Agents
- Feature-Fused Context-Encoding Network for Neuroanatomy Segmentation
- Gaussian Filter in CRF Based Semantic Segmentation
- Deep learning ensembles for melanoma recognition in dermoscopy images
- Deep Reinforcement Learning for Visual Object Tracking in Videos
- Skin Lesion Segmentation and Classification for ISIC 2018 by Combining Deep CNN and Handcrafted Features
- A Systematic Comparison of Deep Learning Architectures in an Autonomous Vehicle
- Vehicle Detection from 3D Lidar Using Fully Convolutional Network
- Structured Binary Neural Networks for Accurate Image Classification and Semantic Segmentation
- Real-Time Well Log Prediction From Drilling Data Using Deep Learning
- CloudifierNet - Deep Vision Models for Artificial Image Processing
- Residual Attention Network for Image Classification
- A PCB Dataset for Defects Detection and Classification
- Frustum VoxNet for 3D object detection from RGB-D or Depth images
- Evolution of Image Segmentation using Deep Convolutional Neural Network: A Survey
- Learning a Discriminative Filter Bank within a CNN for Fine-grained Recognition
Related