ShuffleNet V2: Practical Guidelines for Efficient CNN Architecture Design
2018/07/30 by Ningning Ma, Ma, Ningning, Xiangyu Zhang +5 · 137 citations
Computer Science · Engineering · #Advanced Neural Network Applications #Adversarial Robustness in Machine Learning #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Ferroelectric and Negative Capacitance Devices #cs.CV
paper · pdf · doi:10.48550/arxiv.1807.11164
arxiv created 2018/07/30 · openalex publication_date 2018/07/30 · arxiv updated 2018/07/31 · openalex created_date 2019/06/27 · openalex updated_date 2026/07/28
Abstract
Currently, the neural network architecture design is mostly guided by the indirect metric of computation complexity, i.e., FLOPs. However, the direct metric, e.g., speed, also depends on the other factors such as memory access cost and platform characterics. Thus, this work proposes to evaluate the direct metric on the target platform, beyond only considering FLOPs. Based on a series of controlled experiments, this work derives several practical guidelines for efficient network design. Accordingly, a new architecture is presented, called ShuffleNet V2. Comprehensive ablation experiments verify that our model is the state-of-the-art in terms of speed and accuracy tradeoff.
Citations
Cited by
- UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models
- Context-Aware Network Based on Multi-scale Spatio-temporal Attention for Action Recognition in Videos
- Towards Ancient Plant Seed Classification: A Benchmark Dataset and Baseline Model
- Causal-Tune: Mining Causal Factors from Vision Foundation Models for Domain Generalized Semantic Segmentation
- Bolmo: Byteifying the Next Generation of Language Models
- Exploring Deep-to-Shallow Transformable Neural Networks for Intelligent Embedded Systems
- ArcGen: Generalizing Neural Backdoor Detection Across Diverse Architectures
- AMD-HookNet++: Evolution of AMD-HookNet with Hybrid CNN-Transformer Feature Enhancement for Glacier Calving Front Segmentation
- GlimmerNet: A Lightweight Grouped Dilated Depthwise Convolutions for UAV-Based Emergency Monitoring
- CLUENet: Cluster Attention Makes Neural Networks Have Eyes
- HSCP: A Two-Stage Spectral Clustering Framework for Resource-Constrained UAV Identification
- Structured Context Learning for Generic Event Boundary Detection
- RISC-V Based TinyML Accelerator for Depthwise Separable Convolutions in Edge AI
- EfficientSAM3: Progressive Hierarchical Distillation for Video Concept Segmentation from SAM1, 2, and 3
- BackWeak: Backdooring Knowledge Distillation Simply with Weak Triggers and Fine-tuning
- Heterogeneous Complementary Distillation
- CompressNAS : A Fast and Efficient Technique for Model Compression using Decomposition
- Beyond ImageNet: Understanding Cross-Dataset Robustness of Lightweight Vision Models
- ODP-Bench: Benchmarking Out-of-Distribution Performance Prediction
- A Study on Inference Latency for Vision Transformers on Mobile Devices
- Lightweight CycleGAN Models for Cross-Modality Image Transformation and Experimental Quality Assessment in Fluorescence Microscopy
- Knowledge-Informed Neural Network for Complex-Valued SAR Image Recognition
- Learning Task-Agnostic Representations through Multi-Teacher Distillation
- Real-Time Crowd Counting for Embedded Systems with Lightweight Architecture
- CSI-4CAST: A Hybrid Deep Learning Model for CSI Prediction with Comprehensive Robustness and Generalization Testing
- KTBox: A Modular LaTeX Framework for Semantic Color, Structured Highlighting, and Scholarly Communication
- Detection of retinal diseases using an accelerated reused convolutional network
- CoLLM-NAS: Collaborative Large Language Models for Efficient Knowledge-Guided Neural Architecture Search
- Taught Well Learned Ill: Towards Distillation-conditional Backdoor Attack
- AntiFLipper: A Secure and Efficient Defense Against Label-Flipping Attacks in Federated Learning
- HierLight-YOLO: A Hierarchical and Lightweight Object Detection Network for UAV Photography
- Johnson-Lindenstrauss Lemma Guided Network for Efficient 3D Medical Segmentation
- MS-YOLO: Infrared Object Detection for Edge Deployment via MobileNetV4 and SlideLoss
- Deep Lookup Network
- EvHand-FPV: Efficient Event-Based 3D Hand Tracking from First-Person View
- GhostNetV3-Small: A Tailored Architecture and Comparative Study of Distillation Strategies for Tiny Images
- DeepTraverse: A Depth-First Search Inspired Network for Algorithmic Visual Understanding
- NAT: Learning to Attack Neurons for Enhanced Adversarial Transferability
- Unified Start, Personalized End: Progressive Pruning for Efficient 3D Medical Image Segmentation
- NeuroDeX: Unlocking Diverse Support in Decompiling Deep Neural Network Executables
- Tackling Device Data Distribution Real-time Shift via Prototype-based Parameter Editing
- Prior Distribution and Model Confidence
- A biologically inspired separable learning vision model for real-time traffic object perception in Dark
- Leveraging Transfer Learning and Mobile-enabled Convolutional Neural Networks for Improved Arabic Handwritten Character Recognition
- A Lightweight Group Multiscale Bidirectional Interactive Network for Real-Time Steel Surface Defect Detection
- Enhanced Fingerprint-based Positioning With Practical Imperfections: Deep learning-based approaches
- E-ConvNeXt: A Lightweight and Efficient ConvNeXt Variant with Cross-Stage Partial Connections
- VQualA 2025 Challenge on Face Image Quality Assessment: Methods and Results
- TAIGen: Training-Free Adversarial Image Generation via Diffusion Models
- AdapSNE: Adaptive Fireworks-Optimized and Entropy-Guided Dataset Sampling for Edge DNN Training
- UAV Individual Identification via Distilled RF Fingerprints-Based LLM in ISAC Networks
- Towards SISO Bistatic Sensing for ISAC
- Hierarchical knowledge guided fault intensity diagnosis of complex industrial systems
- Deep Learning for Automated Identification of Vietnamese Timber Species: A Tool for Ecological Monitoring and Conservation
- Low-Regret and Low-Complexity Learning for Hierarchical Inference
- MSPT: A Lightweight Face Image Quality Assessment Method with Multi-stage Progressive Training
- Lightweight Multi-Scale Feature Extraction with Fully Connected LMF Layer for Salient Object Detection
- Trained Rank Pruning for Efficient Deep Neural Networks
- ULU: A Unified Activation Function
- Learning from Oblivion: Predicting Knowledge Overflowed Weights via Retrodiction of Forgetting
- Multi-Granularity Feature Calibration via VFM for Domain Generalized Semantic Segmentation
- InfoSyncNet: Information Synchronization Temporal Convolutional Network for Visual Speech Recognition
- NMS: Efficient Edge DNN Training via Near-Memory Sampling on Manifolds
- Rein++: Efficient Generalization and Adaptation for Semantic Segmentation with Vision Foundation Models
- Set Pivot Learning: Redefining Generalized Segmentation with Vision Foundation Models
- Calibrated Prediction Set in Fault Detection with Risk Guarantees via Significance Tests
- SBP-YOLO:A Lightweight Real-Time Model for Detecting Speed Bumps and Potholes toward Intelligent Vehicle Suspension Systems
- Minimum Data, Maximum Impact: 20 annotated samples for explainable lung nodule classification
- Local Dense Logit Relations for Enhanced Knowledge Distillation
- Quantum Machine Learning for Secure Cooperative Multi-Layer Edge AI with Proportional Fairness
- IRLAS: Inverse Reinforcement Learning for Architecture Search
- ChamNet: Towards Efficient Network Design through Platform-Aware Model Adaptation
- Design and Implementation of an Annotation-Driven Drone Autonomy Tool Using YOLOv8–V11 Architectures for Real-Time Object Detection and Distance Estimation
- CARS: Continuous Evolution for Efficient Neural Architecture Search
- UGPL: Uncertainty-Guided Progressive Learning for Evidence-Based Classification in Computed Tomography
- Convolutional neural networks compression with low rank and sparse tensor decompositions
- Prototypical Progressive Alignment and Reweighting for Generalizable Semantic Segmentation
- FasTUSS: Faster Task-Aware Unified Source Separation
- Dynamic Multi-path Neural Network
- Small, Accurate, and Fast Vehicle Re-ID on the Edge: the SAFR Approach
- Critical dynamics governs deep learning
- StixelNExT++: Lightweight Monocular Scene Segmentation and Representation for Collective Perception
- Zero-Shot Neural Architecture Search with Weighted Response Correlation
- Efficient SAR Vessel Detection for FPGA-Based On-Satellite Sensing
- TinyProto: Communication-Efficient Federated Learning with Sparse Prototypes in Resource-Constrained Environments
- Heterogeneous Federated Learning with Prototype Alignment and Upscaling
- FNA++: Fast Network Adaptation via Parameter Remapping and Architecture Search
- Progressive DARTS: Bridging the Optimization Gap for NAS in the Wild
- ADAptation: Reconstruction-based Unsupervised Active Learning for Breast Ultrasound Diagnosis
- Neural Architecture Search using Deep Neural Networks and Monte Carlo Tree Search
- LASFNet: A Lightweight Attention-Guided Self-Modulation Feature Fusion Network for Multimodal Object Detection
- Feature Hallucination for Self-supervised Action Recognition
- Building Lightweight Semantic Segmentation Models for Aerial Images Using Dual Relation Distillation
- Biased Teacher, Balanced Student
- Dual-Forward Path Teacher Knowledge Distillation: Bridging the Capacity Gap Between Teacher and Student
- TEM3-Learning: Time-Efficient Multimodal Multi-Task Learning for Advanced Assistive Driving
- NestQuant: Post-Training Integer-Nesting Quantization for On-Device DNN
- AQUA20: A Benchmark Dataset for Underwater Species Classification under Challenging Conditions
- GroupNL: Low-Resource and Robust CNN Design over Cloud and Device
- ECLIP: Energy-efficient and Practical Co-Location of ML Inference on Spatially Partitioned GPUs
- Towards Undistillable Models by Minimizing Conditional Mutual Information
- OCNet: Object Context Network for Scene Parsing
- A Layered Self-Supervised Knowledge Distillation Framework for Efficient Multimodal Learning on the Edge
- LGM-Pose: A Lightweight Global Modeling Network for Real-time Human Pose Estimation
- LPRNet: Lightweight Deep Network by Low-rank Pointwise Residual Convolution
- Measuring the Algorithmic Efficiency of Neural Networks
- Knowledge Distillation: A Survey
- Progressive Class-level Distillation
- Comparative Analysis of Lightweight CNNs for Resource-Constrained Devices: Predictive Performance, Efficiency Trade-offs, and Initialization Effects
- LeMoRe: Learn More Details for Lightweight Semantic Segmentation
- Cross-filter compression for CNN inference acceleration
- RePaViT: Scalable Vision Transformer Acceleration via Structural Reparameterization on Feedforward Network Layers
- Expressive yet Efficient Feature Expansion with Adaptive Cross-Hadamard Products
- A Lightweight Hybrid Dual Channel Speech Enhancement System under Low-SNR Conditions
- Location-guided lesions representation learning via image generation for assessing plant leaf diseases severity
- Searching for Low-Bit Weights in Quantized Neural Networks
- Taming Diffusion for Dataset Distillation with High Representativeness
- HCRMP: A LLM-Hinted Contextual Reinforcement Learning Framework for Autonomous Driving
- Intra-class Patch Swap for Self-Distillation
- Expert-Like Reparameterization of Heterogeneous Pyramid Receptive Fields in Efficient CNNs for Fair Medical Image Classification
- AGI-Elo: How Far Are We From Mastering A Task?
- Energy-Aware Deep Learning on Resource-Constrained Hardware
- FiGKD: Fine-Grained Knowledge Distillation via High-Frequency Detail Transfer
- Understanding Nonlinear Implicit Bias via Region Counts in Input Space
- Depth-Sensitive Soft Suppression with RGB-D Inter-Modal Stylization Flow for Domain Generalization Semantic Segmentation
- AutoShuffleNet: Learning Permutation Matrices via an Exact Lipschitz Continuous Penalty in Deep Convolutional Neural Networks
- AdaDINO: Context-Adaptive DINO-Distilled Vision Foundation Models for Efficient Open-Vocabulary Edge Inference
- GLOBE: Trajectory-Aligned Gradient Matching with Structured SparseOptimization for Coreset Selection
- CliffordNet: All You Need is Geometric Algebra
- xEdgeFace: Efficient Cross-Spectral Face Recognition for Edge Devices
- Swapped Logit Distillation via Bi-level Teacher Alignment
- FX-DARTS: Designing Topology-unconstrained Architectures with Differentiable Architecture Search and Entropy-based Super-network Shrinking
- Conformal Segmentation in Industrial Surface Defect Detection with Statistical Guarantees
- Seeking Flat Minima over Diverse Surrogates for Improved Adversarial Transferability: A Theoretical Framework and Algorithmic Instantiation
- ISTD-YOLO: A Multi-Scale Lightweight High-Performance Infrared Small Target Detection Algorithm
- Hardware-Guided Symbiotic Training for Compact, Accurate, yet Execution-Efficient LSTM
- MULTI-LF: A Continuous Learning Framework for Real-Time Malicious Traffic Detection in Multi-Environment Networks
- GFT: Gradient Focal Transformer
- Compound and Parallel Modes of Tropical Convolutional Neural Networks
Related