A ConvNet for the 2020s
2022/06/01 by Zhuang Liu, Hanzi Mao, Chao-Yuan Wu +3 · 251 citations
Computer Science · #Advanced Image and Video Retrieval Techniques #Advanced Neural Network Applications #Domain Adaptation and Few-Shot Learning
paper · doi:10.1109/cvpr52688.2022.01167
openalex publication_date 2022/06/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/30
Abstract
The “Roaring 20s” of visual recognition began with the introduction of Vision Transformers (ViTs), which quickly superseded ConvNets as the state-of-the-art image classification model. A vanilla ViT, on the other hand, faces difficulties when applied to general computer vision tasks such as object detection and semantic segmentation. It is the hierarchical Transformers (e.g., Swin Transformers) that reintroduced several ConvNet priors, making Transformers practically viable as a generic vision backbone and demonstrating remarkable performance on a wide variety of vision tasks. However, the effectiveness of such hybrid approaches is still largely credited to the intrinsic superiority of Transformers, rather than the inherent inductive biases of convolutions. In this work, we reexamine the design spaces and test the limits of what a pure ConvNet can achieve. We gradually “modernize” a standard ResNet toward the design of a vision Transformer, and discover several key components that contribute to the performance difference along the way. The outcome of this exploration is a family of pure ConvNet models dubbed ConvNeXt. Constructed entirely from standard ConvNet modules, ConvNeXts compete favorably with Transformers in terms of accuracy and scalability, achieving 87.8% ImageNet top-1 accuracy and outperforming Swin Transformers on COCO detection and ADE20K segmentation, while maintaining the simplicity and efficiency of standard ConvNets.
Citations
Cited by
- Applications of Machine Learning and Artificial Intelligence in Tropospheric Ozone Research
- Amortized Moment Matching for Visual Generation
- A point and density map hybrid network for crowd counting and localization based on unmanned aerial vehicles
- MAPS: A Synthetic Dataset for Probing Vision Models in a Controlled 3D Scene Space
- Recognition of European mammals and birds in camera trap images using deep neural networks
- LSKNet: A Foundation Lightweight Backbone for Remote Sensing
- IBNorm: Information-Bottleneck Inspired Normalization for Representation Learning
- Neighborhood Feature Pooling for Remote Sensing Image Classification
- PitchFlower: A flow-based neural audio codec with pitch controllability
- Kernelized Sparse Fine-Tuning with Bi-level Parameter Competition for Vision Models
- Beyond Inference Intervention: Identity-Decoupled Diffusion for Face Anonymization
- Deep Feature Optimization for Enhanced Fish Freshness Assessment
- UHKD: A Unified Framework for Heterogeneous Knowledge Distillation via Frequency-Domain Representations
- Enhancing Pre-trained Representation Classifiability can Boost its Interpretability
- PLDC‐Net: A Domain‐Specific Base Model for Plant Leaf Disease Classification Domain Adaptation Tasks
- TTST: A Top-<i>k</i> Token Selective Transformer for Remote Sensing Image Super-Resolution
- Deep learning-based ecological analysis of camera trap images is impacted by training data quality and quantity
- SARATR-X: Toward Building a Foundation Model for SAR Target Recognition
- Anomaly Detection for Medical Images Using Heterogeneous Auto-Encoder
- FocalTransNet: A Hybrid Focal-Enhanced Transformer Network for Medical Image Segmentation
- Perceptual Quality Assessment of 360° Images Based on Generative Scanpath Representation
- Being Ranked in a Material World: The visual originality of an artwork and its effects on the artist’s canonization
- RankSEG-RMA: An Efficient Segmentation Algorithm via Reciprocal Moment Approximation
- Model-Behavior Alignment under Flexible Evaluation: When the Best-Fitting Model Isn't the Right One
- SeeDNorm: Self-Rescaled Dynamic Normalization
- Cross-view Localization and Synthesis -- Datasets, Challenges and Opportunities
- From Pixels to Views: Learning Angular-Aware and Physics-Consistent Representations for Light Field Microscopy
- SARVLM: A Vision Language Foundation Model for Semantic Understanding in SAR Imagery
- PromptReverb: Multimodal Room Impulse Response Generation Through Latent Rectified Flow Matching
- Unified token representations for sequential decision models
- Sensing and Storing Less: A MARL-based Solution for Energy Saving in Edge Internet of Things
- Smule Renaissance Small: Efficient General-Purpose Vocal Restoration
- PointMapPolicy: Structured Point Cloud Processing for Multi-Modal Imitation Learning
- What Does It Take to Build a Performant Selective Classifier?
- Attentive Convolution: Unifying the Expressivity of Self-Attention with Convolutional Efficiency
- Dynamic Weight Adjustment for Knowledge Distillation: Leveraging Vision Transformer for High-Accuracy Lung Cancer Detection and Real-Time Deployment
- ConvXformer: Differentially Private Hybrid ConvNeXt-Transformer for Inertial Navigation
- MobiAct: Efficient MAV Action Recognition Using MobileNetV4 with Contrastive Learning and Knowledge Distillation
- Integrated representational signatures strengthen specificity in brains and models
- VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models
- ZACH-ViT: A Zero-Token Vision Transformer with ShuffleStrides Data Augmentation for Robust Lung Ultrasound Classification
- ReefNet: A Large scale, Taxonomically Enriched Dataset and Benchmark for Hard Coral Classification
- Automated C-Arm Positioning via Conformal Landmark Localization
- Intermittent File Encryption in Ransomware: Measurement, Modeling, and Detection
- Cross-Layer Feature Self-Attention Module for Multi-Scale Object Detection
- Vision-Centric Activation and Coordination for Multimodal Large Language Models
- NTIRE 2025 Challenge on Low Light Image Enhancement: Methods and Results
- CoDS: Enhancing Collaborative Perception in Heterogeneous Scenarios via Domain Separation
- Removing Cost Volumes from Optical Flow Estimators
- Prompt-based Adaptation in Large-scale Vision Models: A Survey
- Multi-Scale High-Resolution Logarithmic Grapher Module for Efficient Vision GNNs
- Assessing the Potential for Catastrophic Failure in Dynamic Post-Training Quantization
- DIANet: A Phase-Aware Dual-Stream Network for Micro-Expression Recognition via Dynamic Images
- DiSTAR: Diffusion over a Scalable Token Autoregressive Representation for Speech Generation
- UALM: Unified Audio Language Model for Understanding, Generation and Reasoning
- PanoTPS-Net: Panoramic Room Layout Estimation via Thin Plate Spline Transformation
- Joint Discriminative-Generative Modeling via Dual Adversarial Training
- Exploring and Leveraging Class Vectors for Classifier Editing
- Deep semi-supervised approach based on consistency regularization and similarity learning for weeds classification
- Unified Open-World Segmentation with Multi-Modal Prompts
- Learning Model Representations Using Publicly Available Model Hubs
- DREAM: A Benchmark Study for Deepfake REalism AssessMent
- Leveraging Prior Knowledge of Diffusion Model for Person Search
- One Pass Is Not Enough: Recursive Latent Refinement for Generative Models
- Impact of Scanner Manufacturer, Endorectal Coil Use, and Clinical Variables on Deep Learning–assisted Prostate Cancer Classification Using Multiparametric MRI
- Probabilistic bias adjustment of seasonal predictions of Arctic Sea Ice Concentration
- Vision Language Models: A Survey of 26K Papers
- Scalable Offline Metrics for Autonomous Driving
- Resolution scaling governs DINOv3 transfer performance in chest radiograph classification
- HARP-NeXt: High-Speed and Accurate Range-Point Fusion Network for 3D LiDAR Semantic Segmentation
- NPN: Non-Linear Projections of the Null-Space for Imaging Inverse Problems
- Shaken or Stirred? An Analysis of MetaFormer's Token Mixing for Medical Imaging
- SFANet: Spatial-Frequency Attention Network for Deepfake Detection
- ERDE: Entropy-Regularized Distillation for Early-exit
- Confidence and Dispersity as Signals: Unsupervised Model Evaluation and Ranking
- CoT Referring: Improving Referring Expression Tasks with Grounded Reasoning
- Hyperparameter Loss Surfaces Are Simple Near their Optima
- NLDSI-BWE: Non Linear Dynamical Systems-Inspired Multi Resolution Discriminators for Speech Bandwidth Extension
- FlexiCodec: A Dynamic Neural Audio Codec for Low Frame Rates
- LAKAN: Landmark-assisted Adaptive Kolmogorov-Arnold Network for Face Forgery Detection
- Assessing Foundation Models for Mold Colony Detection with Limited Training Data
- ProbMed: A Probabilistic Framework for Medical Multimodal Binding
- OmniDFA: A Unified Framework for Open Set Synthesis Image Detection and Few-Shot Attribution
- PRPO: Paragraph-level Policy Optimization for Vision-Language Deepfake Detection
- AttentionViG: Cross-Attention-Based Dynamic Neighbor Aggregation in Vision GNNs
- Bayesian Transformer for Pan-Arctic Sea Ice Concentration Mapping and Uncertainty Estimation using Sentinel-1, RCM, and AMSR2 Data
- VNODE: A Piecewise Continuous Volterra Neural Network
- Towards Foundation Models for Cryo-ET Subtomogram Analysis
- FSDENet: A Frequency and Spatial Domains based Detail Enhancement Network for Remote Sensing Semantic Segmentation
- Conda: Column-Normalized Adam for Training Large Language Models Faster
- A phase transition in diffusion models reveals the hierarchical nature of data
- Physics Priors Offer Useful Accuracy-Carbon Trade-Offs in Spatio-Temporal Forecasting
- Does Weak-to-strong Generalization Happen under Spurious Correlations?
- Multi-modal Data Spectrum: Multi-modal Datasets are Multi-dimensional
- Modeling the language cortex with form-independent and enriched representations of sentence meaning reveals remarkable semantic abstractness
- FracDetNet: Advanced Fracture Detection via Dual-Focus Attention and Multi-scale Calibration in Medical X-ray Imaging
- Enhanced Fracture Diagnosis Based on Critical Regional and Scale Aware in YOLO
- Robust Fine-Tuning from Non-Robust Pretrained Models: Mitigating Suboptimal Transfer With Adversarial Scheduling
- A Weakly Supervised and Self-Supervised Learning Approach for Semantic Segmentation of Land Cover in Satellite Images with National Forest Inventory Data
- Deep Learning for Oral Health: Benchmarking ViT, DeiT, BEiT, ConvNeXt, and Swin Transformer
- Convolutional Set Transformer
- Prospective Evaluation of Real‐Time Artificial Intelligence for the Hill Classification of the Gastroesophageal Junction
- TRUST: Test-Time Refinement using Uncertainty-Guided SSM Traverses
- Introducing Multimodal Paradigm for Learning Sleep Staging PSG via General-Purpose Model
- Low-cost, autonomous microscopy using deep learning and robotics: A crystal morphology case study
- IONext: Unlocking the Next Era of Inertial Odometry
- FreqDebias: Towards Generalizable Deepfake Detection via Consistency-Driven Frequency Debiasing
- Conditional Denoising Diffusion Autoencoders for Wireless Semantic Communications
- Ground-Truthing AI Energy Consumption: Validating CodeCarbon Against External Measurements
- MedVSR: Medical Video Super-Resolution with Cross State-Space Propagation
- Less Precise Can Be More Reliable: A Systematic Evaluation of Quantization's Impact on VLMs Beyond Accuracy
- The Unanticipated Asymmetry Between Perceptual Optimization and Assessment
- FSMODNet: A Closer Look at Few-Shot Detection in Multispectral Data
- Does the Manipulation Process Matter? RITA: Reasoning Composite Image Manipulations via Reversely-Ordered Incremental-Transition Autoregression
- Frequency-domain Multi-modal Fusion for Language-guided Medical Image Segmentation
- ICONIC-444: A 3.1-Million-Image Dataset for OOD Detection Research
- VCP-DCN: Beyond Visual Concealed Property via Depth Collaborative Network for Camouflaged Object Detection
- ZMIS-SAM: Segment Anything Model Enhanced with Wavelet Transform for Zooplankton Microscopy Image Instance Segmentation
- BlindPSNR: A No-Reference Fidelity Predictor for Low-Light Image Enhancement
- Compact deep neural network models of the visual cortex
- DS@GT ARC at ImageCLEFmedical 2026: Architectural Diversity for Concept Detection and Foundation-Model Scaling for Caption Prediction in Medical Image Analysis
- <scp>FFM</scp> ‐ <scp>ViT</scp> : an efficient fish species classification method based on deep features and transformers
- Hybrid Quantum-MambaVision: A Quantum-Enhanced State Space Model for Calibrated Mixed-Type Wafer Defect Detection
- Local and Global Feature-Aware Dual-Branch Networks for Plant Disease Recognition
- Designing Practical Models for Isolated Word Visual Speech Recognition
- OA-CNNs: Omni-Adaptive Sparse CNNs for 3D Semantic Segmentation
- Debugging Concept Bottleneck Models through Removal and Retraining
- Rethinking the joint estimation of magnitude and phase for time-frequency domain neural vocoders
- Machine learning approach to single-shot multiparameter estimation for the non-linear Schrödinger equation
- Latent Danger Zone: Distilling Unified Attention for Cross-Architecture Black-box Attacks
- Towards Application Aligned Synthetic Surgical Image Synthesis
- RCTDistill: Cross-Modal Knowledge Distillation Framework for Radar-Camera 3D Object Detection with Temporal Fusion
- LoT-Pass: Long-term-robust Image Watermarking for Image to Video Generation
- Towards Sharper Object Boundaries in Self-Supervised Depth Estimation
- Automatic Classification of Magnetic Chirality of Solar Filaments from H-Alpha Observations
- A Novel Metric for Detecting Memorization in Generative Models for Brain MRI Synthesis
- Improving the quality of respiratory signals extracted from the segmented mask area
- CAMBench-QR : A Structure-Aware Benchmark for Post-Hoc Explanations with QR Understanding
- UniMRSeg: Unified Modality-Relax Segmentation via Hierarchical Self-Supervised Compensation
- Saccadic Vision for Fine-Grained Visual Classification
- Region-Aware Deformable Convolutions
- Leveraging Geometric Visual Illusions as Perceptual Inductive Biases for Vision Models
- Limitations of Public Chest Radiography Datasets for Artificial Intelligence: Label Quality, Domain Shift, Bias and Evaluation Challenges
- OmniSegmentor: A Flexible Multi-Modal Learning Framework for Semantic Segmentation
- AI-Derived Structural Building Intelligence for Urban Resilience: An Application in Saint Vincent and the Grenadines
- Self Identity Mapping
- Real-Time Detection and Tracking of Foreign Object Intrusions in Power Systems via Feature-Based Edge Intelligence
- A biological vision inspired framework for machine perception of abutting grating illusory contours
- Modelling and analysis of the 8 filters from the "master key filters hypothesis" for depthwise-separable deep networks in relation to idealized receptive fields based on scale-space theory
- Global and local pseudo-label filtering for semi-supervised carotid plaque classification from ultrasound
- MSDNet: Efficient 4D Radar Super-Resolution via Multi-Stage Distillation
- Spiking Vocos: An Energy-Efficient Neural Vocoder
- Swin Transformer-Based Multiscale Attention Model for Landslide Extraction From Large-Scale Area
- The Quest for Universal Master Key Filters in DS-CNNs
- Optimizing Class Distributions for Bias-Aware Multi-Class Learning
- UltraUPConvNet: A UPerNet- and ConvNeXt-Based Multi-Task Network for Ultrasound Tissue Segmentation and Disease Prediction
- Multimodal SAM-adapter for Semantic Segmentation
- Transformer Networks for Continuous Gravitational-wave Searches
- Local Information Matters: A Rethink of Crowd Counting
- Augment to Segment: Tackling Pixel-Level Imbalance in Wheat Disease and Pest Segmentation
- NAT: Learning to Attack Neurons for Enhanced Adversarial Transferability
- Graph Alignment via Dual-Pass Spectral Encoding and Latent Space Communication
- CoAtNeXt:An Attention-Enhanced ConvNeXtV2-Transformer Hybrid Model for Gastric Tissue Classification
- Noise-Robust Topology Estimation of 2D Image Data via Neural Networks and Persistent Homology
- AWM-Fuse: Multi-Modality Image Fusion for Adverse Weather via Global and Local Text Perception
- Segment Transformer: AI-Generated Music Detection via Music Structural Analysis
- MedicalPatchNet: A Patch-Based Self-Explainable AI Architecture for Chest X-ray Classification
- Bias in Gender Bias Benchmarks: How Spurious Features Distort Evaluation
- GLEAM: Learning to Match and Explain in Cross-View Geo-Localization
- MRI-Based Brain Tumor Detection through an Explainable EfficientNetV2 and MLP-Mixer-Attention Architecture
- AI-driven Remote Facial Skin Hydration and TEWL Assessment from Selfie Images: A Systematic Solution
- When Language Model Guides Vision: Grounding DINO for Cattle Muzzle Detection
- YOLOv9: Learning What You Want to Learn Using Programmable Gradient Information
- Back To The Drawing Board: Rethinking Scene-Level Sketch-Based Image Retrieval
- Analysis of Transferability Estimation Metrics for Surgical Phase Recognition
- Dual Interaction Network with Cross-Image Attention for Medical Image Segmentation
- Khana: A Comprehensive Indian Cuisine Dataset
- JRN-Geo: A Joint Perception Network based on RGB and Normal images for Cross-view Geo-localization
- A biologically inspired separable learning vision model for real-time traffic object perception in Dark
- Adapt in the Wild: Test-Time Entropy Minimization with Sharpness and Feature Regularization
- VCMamba: Bridging Convolutions with Multi-Directional Mamba for Efficient Visual Representation
- DarkStream: real-time speech anonymization with low latency
- Time-Scaling State-Space Models for Dense Video Captioning
- YOLO-based Bearing Fault Diagnosis With Continuous Wavelet Transform
- A Lightweight Group Multiscale Bidirectional Interactive Network for Real-Time Steel Surface Defect Detection
- Invariant Features for Global Crop Type Classification
- Targeted Physical Evasion Attacks in the Near-Infrared Domain
- An Investigation of Visual Foundation Models Robustness
- Towards deep-learning based detection and quantification of intestinal metaplasia on digitized gastric biopsies: a multi-expert comparative study
- SAR-NAS: Lightweight SAR Object Detection with Neural Architecture Search
- Expandable Residual Approximation for Knowledge Distillation
- Deep Learning-Based Rock Particulate Classification Using Attention-Enhanced ConvNeXt
- REVELIO -- Universal Multimodal Task Load Estimation for Cross-Domain Generalization
- Optical Music Recognition of Jazz Lead Sheets
- Multimodal Deep Learning for Phyllodes Tumor Classification from Ultrasound and Clinical Data
- Solutions for Mitotic Figure Detection and Atypical Classification in MIDOG 2025
- ConvNeXt with Histopathology-Specific Augmentations for Mitotic Figure Classification
- Continuous Determination of Respiratory Rate in Hospitalized Patients using Machine Learning Applied to Electrocardiogram Telemetry
- WaveLLDM: Design and Development of a Lightweight Latent Diffusion Model for Speech Enhancement and Restoration
- Webly-Supervised Image Manipulation Localization via Category-Aware Auto-Annotation
- E-ConvNeXt: A Lightweight and Efficient ConvNeXt Variant with Cross-Stage Partial Connections
- Classifying Mitotic Figures in the MIDOG25 Challenge with Deep Ensemble Learning and Rule Based Refinement
- More Reliable Pseudo-labels, Better Performance: A Generalized Approach to Single Positive Multi-label Learning
- Dual-Model Weight Selection and Self-Knowledge Distillation for Medical Image Classification
- Bridging Domain Gaps for Fine-Grained Moth Classification Through Expert-Informed Adaptation and Foundation Model Priors
- Image Quality Assessment for Machines: Paradigm, Large-scale Database, and Models
- UNIFORM: Unifying Knowledge from Large-scale and Diverse Pre-trained Models
- Bladder Cancer Diagnosis with Deep Learning: A Multi-Task Framework and Online Platform
- Generative Super-Resolution of Turbulent Flows via Stochastic Interpolants
- CLoE: Curriculum Learning on Endoscopic Images for Robust MES Classification
- FNH-TTS: A Fast, Natural, and Human-Like Speech Synthesis System with advanced prosodic modeling based on Mixture of Experts
- LKFMixer: Exploring Large Kernel Feature For Efficient Image Super-Resolution
- Generalized Decoupled Learning for Enhancing Open-Vocabulary Dense Perception
- Processing and acquisition traces in visual encoders: What does CLIP know about your camera?
- PSScreen: Partially Supervised Multiple Retinal Disease Screening
- PQ-DAF: Pose-driven Quality-controlled Data Augmentation for Data-scarce Driver Distraction Detection
- Semantic-Aware Reconstruction Error for Detecting AI-Generated Images
- UniConvNet: Expanding Effective Receptive Field while Maintaining Asymptotically Gaussian Distribution for ConvNets of Any Scale
- Label Smoothing is a Pragmatic Information Bottleneck
- A Guide to Robust Generalization: The Impact of Architecture, Pre-training, and Optimization Strategy
- ASM-UNet: Adaptive Scan Mamba Integrating Group Commonalities and Individual Variations for Fine-Grained Segmentation
- eMotions: A Large-Scale Dataset and Audio-Visual Fusion Network for Emotion Analysis in Short-form Videos
- An antimicrobial drug recommender system using MALDI-TOF MS and dual-branch neural networks
- AdaptInfer: Adaptive Token Pruning for Vision-Language Model Inference with Dynamical Text Guidance
- Deep Distillation Gradient Preconditioning for Inverse Problems
- BridgeDepth: Bridging Monocular and Stereo Reasoning with Latent Alignment
- Segment Any Vehicle: Semantic and Visual Context Driven SAM and A Benchmark
- DocVCE: Diffusion-based Visual Counterfactual Explanations for Document Image Classification
- What Holds Back Open-Vocabulary Segmentation?
- DP-DocLDM: Differentially Private Document Image Generation using Latent Diffusion Models
- Boosting Adversarial Transferability via Residual Perturbation Attack
- Prototype-Driven Structure Synergy Network for Remote Sensing Images Segmentation
- AttZoom: Attention Zoom for Better Visual Features
- GRASPing Anatomy to Improve Pathology Segmentation
- Evaluation and Analysis of Deep Neural Transformers and Convolutional Neural Networks on Modern Remote Sensing Datasets
- FAIR-Pruner: Leveraging Tolerance of Difference for Flexible Automatic Layer-Wise Neural Network Pruning
- After the Party: Navigating the Mapping From Color to Ambient Lighting
- DiffusionFF: A Diffusion-based Framework for Joint Face Forgery Detection and Fine-Grained Artifact Localization
- Foundation Models for Bioacoustics -- a Comparative Review
- Deep Learning for Pavement Condition Evaluation Using Satellite Imagery
- COSTARR: Consolidated Open Set Technique with Attenuation for Robust Recognition
- DBLP: Noise Bridge Consistency Distillation For Efficient And Reliable Adversarial Purification
- Representation Shift: Unifying Token Compression with FlashAttention
- Gaussian Splatting Feature Fields for Privacy-Preserving Visual Localization
- 3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object Detection
- I Am Big, You Are Little; I Am Right, You Are Wrong
- Beyond Linear Bottlenecks: Spline-Based Knowledge Distillation for Culturally Diverse Art Style Classification
- EMedNeXt: An Enhanced Brain Tumor Segmentation Framework for Sub-Saharan Africa using MedNeXt V2 with Deep Supervision
- UAVScenes: A Multi-Modal Dataset for UAVs
- Spatial-Temporal-Spectral Mamba with Sparse Deformable Token Sequence for Enhanced MODIS Time Series Classification
- Brain Tumor Segmentation in Sub-Sahara Africa with Advanced Transformer and ConvNet Methods: Fine-Tuning, Data Mixing and Ensembling
Related