Efficient Processing of Deep Neural Networks: A Tutorial and Survey
2017/11/20 by Vivienne Sze, Yu-Hsin Chen, Yu‐Hsin Chen +3 · 226 citations
Computer Science · Engineering · #Advanced Neural Network Applications #Advanced Memory and Neural Computing #Ferroelectric and Negative Capacitance Devices
paper · doi:10.1109/jproc.2017.2761740
Abstract
Deep neural networks (DNNs) are currently widely used for many artificial intelligence (AI) applications including computer vision, speech recognition, and robotics. While DNNs deliver state-of-the-art accuracy on many AI tasks, it comes at the cost of high computational complexity. Accordingly, techniques that enable efficient processing of DNNs to improve energy efficiency and throughput without sacrificing application accuracy or increasing hardware cost are critical to the wide deployment of DNNs in AI systems. This article aims to provide a comprehensive tutorial and survey about the recent advances toward the goal of enabling efficient processing of DNNs. Specifically, it will provide an overview of DNNs, discuss various hardware platforms and architectures that support DNNs, and highlight key trends in reducing the computation cost of DNNs either solely via hardware design changes or via joint hardware design and DNN algorithm changes. It will also summarize various development resources that enable researchers and practitioners to quickly get started in this field, and highlight important benchmarking metrics and design considerations that should be used for evaluating the rapidly growing number of DNN hardware designs, optionally including algorithmic codesigns, being proposed in academia and industry. The reader will take away the following concepts from this article: understand the key design considerations for DNNs; be able to evaluate different DNN hardware implementations with benchmarks and comparison metrics; understand the tradeoffs between various hardware architectures and platforms; be able to evaluate the utility of various DNN design techniques for efficient processing; and understand recent implementation trends and opportunities.
Citations
Cited by
- Deep Learning With Edge Computing: A Review
- Efficient Deep Learning: A Survey on Making Deep Learning Models Smaller, Faster, and Better
- Sparse by Command: Task-Conditional Compute Skipping for Multi-Task Inference Accelerators
- Monolithically Integrated VO2 Mott Oscillators for Energy-Efficient Spiking Neurons
- DiffAxE: Diffusion-driven Hardware Accelerator Generation and Design Space Exploration
- Programmable Probabilistic Computer with 1,000,000 p-bits
- Pruning-Aware Merging for Efficient Multitask Inference
- Fast GPU Linear Algebra via Compile Time Expression Fusion
- Physics-informed neural network for creep-fatigue life prediction of Inconel 617 and interpretation of influencing factors
- Neural Group Testing to Accelerate Deep Learning
- Towards Fully 8-bit Integer Inference for the Transformer Model
- Single-shot Channel Pruning Based on Alternating Direction Method of Multipliers
- NeuroMAX: A High Throughput, Multi-Threaded, Log-Based Accelerator for Convolutional Neural Networks
- Kernel Forge: An Agent Harness for LLM-based Generation and Optimization of CUDA Kernels
- Stacked Intelligent Metasurface-Aided Wave-Domain Signal Processing: From Communications to Sensing and Computing
- Linear Delay-cell Design for Low-energy Delay Multiplication and\n Accumulation
- An Energy-Efficient Adiabatic Capacitive Neural Network Chip
- Asymptotic Soft Filter Pruning for Deep Convolutional Neural Networks
- Learning Dynamics in Memristor-Based Equilibrium Propagation
- Neural Collapse in Test-Time Adaptation
- RIFT: A Scalable Methodology for LLM Accelerator Fault Assessment using Reinforcement Learning
- SHARe-KAN: Post-Training Vector Quantization for Cache-Resident KAN Inference
- A Comprehensive Survey of Machine Learning Applied to Radar Signal Processing
- Vision and Causal Learning Based Channel Estimation for THz Communications
- Beyond Classification: Directly Training Spiking Neural Networks for Semantic Segmentation
- Efficient Kernel Mapping and Comprehensive System Evaluation of LLM Acceleration on a CGLA
- TWEO: Transformers Without Extreme Outliers Enables FP8 Training And Quantization For Dummies
- The Final Frontier: Deep Learning in Space
- Physics-Informed Spiking Neural Networks via Conservative Flux Quantization
- Optimally Scheduling CNN Convolutions for Efficient Memory Access
- Pruning and Quantization for Deep Neural Network Acceleration: A Survey
- Federated Learning in Mobile Edge Networks: A Comprehensive Survey
- Equivariant-Aware Structured Pruning for Efficient Edge Deployment: A Comprehensive Framework with Adaptive Fine-Tuning
- Visual Framing of Science Conspiracy Videos: Integrating Machine Learning with Communication Theories to Study the Use of Color and Brightness
- Sex and age determination in European lobsters using AI-Enhanced bioacoustics
- Towards Practical Real-Time Low-Latency Music Source Separation
- Gazelle: A Low Latency Framework for Secure Neural Network Inference
- Bi-Real Net: Binarizing Deep Network Towards Real-Network Performance
- Exploiting Weight Redundancy in CNNs: Beyond Pruning and Quantization
- A comparison of deep machine learning algorithms in COVID-19 disease\n diagnosis
- SynapticCore-X: A Modular Neural Processing Architecture for Low-Cost FPGA Acceleration
- LILogic Net: Compact Logic Gate Networks with Learnable Connectivity for Efficient Hardware Deployment
- CompressNAS : A Fast and Efficient Technique for Model Compression using Decomposition
- OpenMENA: An Open-Source Memristor Interfacing and Compute Board for Neuromorphic Edge-AI Applications
- Always-On 674uW @ 4GOP/s Error Resilient Binary Neural Networks with\n Aggressive SRAM Voltage Scaling on a 22nm IoT End-Node
- Edge Artificial Intelligence for 6G: Vision, Enabling Technologies, and Applications
- SoftNeuro: Fast Deep Inference using Multi-platform Optimization
- Data Trajectory Alignment for LLM Domain Adaptation: A Two-Phase Synthesis Framework for Telecommunications Mathematics
- Synheart Emotion: Privacy-Preserving On-Device Emotion Recognition from Biosignals
- Learning a Decentralized Medium Access Control Protocol for Shared Message Transmission
- SinReQ: Generalized Sinusoidal Regularization for Low-Bitwidth Deep\n Quantized Training
- A Memory-Efficient Retrieval Architecture for RAG-Enabled Wearable Medical LLMs-Agents
- Sparse Optimization for Green Edge AI Inference
- On-Device Machine Learning: An Algorithms and Learning Theory Perspective
- Edge Artificial Intelligence for 6G: Vision, Enabling Technologies, and Applications
- Edge Intelligence: Paving the Last Mile of Artificial Intelligence with Edge Computing
- Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis
- Learning compositional functions via multiplicative weight updates
- Scheduling Massive Camera Streams to Optimize Large-Scale Live Video Analytics
- Bayesian neural networks for flight trajectory prediction and safety assessment
- Escoin: Efficient Sparse Convolutional Neural Network Inference on GPUs
- Anomaly Detection in Residential Video Surveillance on Edge Devices in IoT Framework
- A Spike in Performance: Training Hybrid-Spiking Neural Networks with Quantized Activation Functions
- Improving Efficiency in Neural Network Accelerator Using Operands Hamming Distance optimization
- Lincoln AI Computing Survey (LAICS) and Trends
- TernaryCLIP: Efficiently Compressing Vision-Language Models with Ternary Weights and Distilled Knowledge
- Stealing Neural Networks via Timing Side Channels
- Enhancing Time-Series Anomaly Detection by Integrating Spectral-Residual Bottom-Up Attention with Reservoir Computing
- Multiplier-free In-Memory Vector-Matrix Multiplication Using Distributed Arithmetic
- Nonlinear optical feature generator for machine learning
- A Methodology for Transparent Logic-Based Classification Using a Multi-Task Convolutional Tsetlin Machine
- SQS: Bayesian DNN Compression through Sparse Quantized Sub-distributions
- CoEdge: Cooperative DNN Inference with Adaptive Workload Partitioning over Heterogeneous Edge Devices
- GUIDE: Guided Initialization and Distillation of Embeddings
- Environment-Aware and Training-Free Beam Alignment for mmWave Massive MIMO via Channel Knowledge Map
- Benchmarking Deep Learning Convolutions on Energy-constrained CPUs
- A novel fusion architecture for detecting Parkinson’s Disease using semi-supervised speech embeddings
- Survey of Machine Learning Accelerators
- Pruning Algorithms to Accelerate Convolutional Neural Networks for Edge Applications: A Survey
- High Clockrate Free-space Optical In-Memory Computing
- Communication-Efficient Distributed Deep Learning: A Comprehensive Survey
- Energy-Efficient Processing and Robust Wireless Cooperative Transmission for Edge Inference
- Cryptocurrency Trading: A Comprehensive Survey
- A Theoretical Framework for Stochastic Activity Prediction in Tensor Accelerator Wallace-Tree Multipliers
- Thanks for Nothing: Predicting Zero-Valued Activations with Lightweight Convolutional Neural Networks
- Puzzle: Scheduling Multiple Deep Learning Models on Mobile Device with Heterogeneous Processors
- Evaluating the Energy Efficiency of NPU-Accelerated Machine Learning Inference on Embedded Microcontrollers
- Joint Memory Frequency and Computing Frequency Scaling for Energy-efficient DNN Inference
- MPNA: A Massively-Parallel Neural Array Accelerator with Dataflow\n Optimization for Convolutional Neural Networks
- FPGA-Based Accelerators of Deep Learning Networks for Learning and Classification: A Review
- Deep Hierarchical Learning with Nested Subspace Networks for Large Language Models
- Survey and Benchmarking of Machine Learning Accelerators
- Model Compression and Hardware Acceleration for Neural Networks: A Comprehensive Survey
- A visual introduction to Gaussian Belief Propagation
- eIQ Neutron: Redefining Edge-AI Inference with Integrated NPU and Compiler Innovations
- Orthrus: Dual-Loop Automated Framework for System-Technology Co-Optimization
- MC2RAM: Markov Chain Monte Carlo Sampling in SRAM for Fast Bayesian\n Inference
- Deep Learning-based Techniques for Integrated Sensing and Communication Systems: State-of-the-Art, Challenges, and Opportunities
- Autonomous Driving with Deep Learning: A Survey of State-of-Art Technologies
- PointAR: Efficient Lighting Estimation for Mobile Augmented Reality
- Knowledge Transfer via Dense Cross-Layer Mutual-Distillation
- Training and Inference with Integers in Deep Neural Networks
- In-Loop Filtering Using Learned Look-Up Tables for Video Coding
- Deep Learning for LiDAR Point Clouds in Autonomous Driving: A Review
- Evolutionary Algorithms in Approximate Computing: A Survey
- Communication-Efficient Edge AI: Algorithms and Systems
- Breaking SafetyCore: Exploring the Risks of On-Device AI Deployment
- Scenario-based Decision-making Using Game Theory for Interactive Autonomous Driving: A Survey
- Selective Structural Ablation for Efficient 3D Point Cloud Signal Processing
- Real Time FPGA Based CNNs for Detection, Classification, and Tracking in Autonomous Systems: State of the Art Designs and Optimizations
- Hardwired-Neurons Language Processing Units as General-Purpose Cognitive Substrates
- Application of discrete Ricci curvature in pruning randomly wired neural networks: A case study with chest x-ray classification of COVID-19
- Deep Learning in Mobile and Wireless Networking: A Survey
- Practical GPU Choices for Earth Observation: ResNet-50 Training Throughput on Integrated, Laptop, and Cloud Accelerators
- Neural Network Quantization for Microcontrollers: A Comprehensive Survey of Methods, Platforms, and Applications
- Cross-Layer Design of Vector-Symbolic Computing: Bridging Cognition and Brain-Inspired Hardware Acceleration
- ReActNet: Towards Precise Binary Neural Network with Generalized Activation Functions
- Fast Exploration of Weight Sharing Opportunities for CNN Compression
- NetAdapt: Platform-Aware Neural Network Adaptation for Mobile Applications
- DeepNetQoE: Self-adaptive QoE Optimization Framework of Deep Networks
- Speeding up Deep Learning with Transient Servers
- Quantization through Piecewise-Affine Regularization: Optimization and Statistical Guarantees
- Phantom: A High-Performance Computational Core for Sparse Convolutional Neural Networks
- GrCAN: Gradient Boost Convolutional Autoencoder with Neural Decision Forest
- Memory devices and applications for in-memory computing
- R2F: A Remote Retraining Framework for AIoT Processors with Computing Errors
- Noisy Machines: Understanding Noisy Neural Networks and Enhancing Robustness to Analog Hardware Errors Using Distillation
- AdderSR: Towards Energy Efficient Image Super-Resolution
- Multi-vision Attention Networks for On-line Red Jujube Grading
- Foundation Models for Demand Forecasting via Dual-Strategy Ensembling
- The Carbon Cost of Conversation, Sustainability in the Age of Language Models
- BRIEF: Backward Reduction of CNNs with Information Flow Analysis
- Towards a Robust Deep Neural Network in Texts: A Survey
- RPR: Random Partition Relaxation for Training; Binary and Ternary Weight Neural Networks
- The Impact of GPU DVFS on the Energy and Performance of Deep Learning: an Empirical Study
- Artificial intelligence in the discovery and design of molecular semiconductors: a systematic review
- Carbon-doped GeTe-based ovonic threshold switch for highly reliable artificial neuron devices
- Human mimetic sensory-interfaced neuromorphic devices and their training mechanism
- Structured Pruning for Deep Convolutional Neural Networks: A Survey
- Dynamic Runtime Feature Map Pruning
- Adaptive Block-Scaled Data Types
- PhoneBit: Efficient GPU-Accelerated Binary Neural Network Inference Engine for Mobile Phones
- Yield Loss Reduction and Test of AI and Deep Learning Accelerators
- Benchmarking Physical Performance of Neural Inference Circuits
- All Optical Classification Surpasses Cascaded Diffractive Networks through Dual Wavelength Differential Modulation within a Single Layer Architecture
- The synergy of neuromarketing and artificial intelligence: A comprehensive literature review in the last decade
- High Performance and Portable Convolution Operators for ARM-based Multicore Processors
- MetaPruning: Meta Learning for Automatic Neural Network Channel Pruning
- An Energy-Efficient Mixed-Signal Parallel Multiply-Accumulate (MAC) Engine Based on Stochastic Computing
- Approximate Processing Element Design and Analysis for the Implementation of CNN Accelerators
- Overfitting Mechanism and Avoidance in Deep Neural Networks
- Energy Efficiency in AI for 5G and Beyond: A DeepRx Case Study
- Weight-Sharing Neural Architecture Search: A Battle to Shrink the Optimization Gap
- Kernel Based Progressive Distillation for Adder Neural Networks
- Versatile On-device Adaptation at the Edge by Unifying Few-shot, Zero-shot, Continual, and In-context Learning
- From the logistic-sigmoid to nlogistic-sigmoid: modelling the COVID-19 pandemic growth
- LiveChess2FEN: a Framework for Classifying Chess Pieces based on CNNs
- H-VGRAE: A Hierarchical Stochastic Spatial-Temporal Embedding Method for Robust Anomaly Detection in Dynamic Networks
- Dynamical Isometry: The Missing Ingredient for Neural Network Pruning
- Performance Evaluation of Deep Learning Tools in Docker Containers
- Towards Lossless Binary Convolutional Neural Networks Using Piecewise Approximation
- QPART: Adaptive Model Quantization and Dynamic Workload Balancing for Accuracy-aware Edge Inference
- A Survey on Deep Learning for Multimodal Data Fusion
- A survey of the recent architectures of deep convolutional neural networks
- Characterizing the Deep Neural Networks Inference Performance of Mobile Applications
- ShiftCNN: Generalized Low-Precision Architecture for Inference of Convolutional Neural Networks
- A Programmable Heterogeneous Microprocessor Based on Bit-Scalable In-Memory Computing
- ACNN: a Full Resolution DCNN for Medical Image Segmentation
- Effects of VLSI Circuit Constraints on Temporal-Coding Multilayer Spiking Neural Networks
- Advances in Intelligent Hearing Aids: Deep Learning Approaches to Selective Noise Cancellation
- STONNE: A Detailed Architectural Simulator for Flexible Neural Network Accelerators
- Particle Filter Networks with Application to Visual Localization
- A Low-Compexity Deep Learning Framework For Acoustic Scene\n Classification
- Constrained Generative Adversarial Network Ensembles for Sharable Synthetic Data Generation
- Exploring Weight Importance and Hessian Bias in Model Pruning
- Model compression using knowledge distillation with integrated gradients
- A Coupled CMOS Oscillator Array for 8ns and 55pJ Inference in Convolutional Neural Networks
- Compression of Deep Convolutional Neural Networks under Joint Sparsity Constraints
- K-TanH: Efficient TanH For Deep Learning
- PBGen: Partial Binarization of Deconvolution-Based Generators for Edge Intelligence
- Machine Intelligence on Wireless Edge Networks
- The Cambrian Explosion of Mixed-Precision Matrix Multiplication for Quantized Deep Learning Inference
- Advanced fraud detection using machine learning models: enhancing financial transaction security
- Advances in Small-Footprint Keyword Spotting: A Comprehensive Review of Efficient Models and Algorithms
- Efficient Winograd Convolution via Integer Arithmetic
- Pegasus: A Universal Framework for Scalable Deep Learning Inference on the Dataplane
- Towards Mixed-Precision Quantization of Neural Networks via Constrained Optimization
- RIOT-ML: toolkit for over-the-air secure updates and performance evaluation of TinyML models
- Using machine learning to model the training scalability of convolutional neural networks on clusters of GPUs
- EnGN: A High-Throughput and Energy-Efficient Accelerator for Large Graph Neural Networks
- Fine-Grained Energy and Performance Profiling framework for Deep Convolutional Neural Networks
- FPGA-Enabled Machine Learning Applications in Earth Observation: A Systematic Review
- The History Began from AlexNet: A Comprehensive Survey on Deep Learning Approaches
- A mixed signal architecture for convolutional neural networks
- Test-Time Distillation for Continual Model Adaptation
- Accelerating CNN inference on FPGAs: A Survey
- Device-Circuit-Architecture Co-Exploration for Computing-in-Memory Neural Accelerators
- Shredder: Learning Noise Distributions to Protect Inference Privacy
- A Tutorial-cum-Survey on Self-Supervised Learning for Wi-Fi Sensing: Trends, Challenges, and Outlook
- High performance and energy efficient inference for deep learning on ARM processors
- Reducing Energy Bloat in Large Model Training
- End-to-End Learned Image Compression with Quantized Weights and Activations
- Cross-filter compression for CNN inference acceleration
- Computerized Ultrasonic Imaging Inspection: From Shallow to Deep Learning. [europepmc]
- Deep Learning With Spiking Neurons: Opportunities and Challenges. [europepmc]
- Memory-Efficient Deep Learning on a SpiNNaker 2 Prototype. [europepmc]
- An End-to-End Deep Neural Network for Autonomous Driving Designed for Embedded Automotive Platforms. [europepmc]
- Memristive and CMOS Devices for Neuromorphic Computing. [europepmc]
- SLIM: Simultaneous Logic-in-Memory Computing Exploiting Bilayer Analog OxRAM Devices. [europepmc]
- Deep learning approach to detect seizure using reconstructed phase space images. [europepmc]
- Rapid vessel segmentation and reconstruction of head and neck angiograms using 3D convolutional neural network. [europepmc]
- A bird's-eye view of deep learning in bioimage analysis. [europepmc]
- Deep Learning-Based Screening Test for Cognitive Impairment Using Basic Blood Test Data for Health Examination. [europepmc]
- An optical neural chip for implementing complex-valued neural network. [europepmc]
- Freely scalable and reconfigurable optical hardware for deep learning. [europepmc]
- In situ Parallel Training of Analog Neural Network Using Electrochemical Random-Access Memory. [europepmc]
- Artificial intelligence explainability: the technical and ethical dimensions. [europepmc]
- Enabling Training of Neural Networks on Noisy Hardware. [europepmc]
- Radiomics, deep learning and early diagnosis in oncology. [europepmc]
- Deep learning-a first meta-survey of selected reviews across scientific disciplines, their commonalities, challenges and research impact. [europepmc]
- An optical neural network using less than 1 photon per multiplication. [europepmc]
- Ten quick tips for deep learning in biology. [europepmc]
- Anomaly detection using edge computing in video surveillance system: review. [europepmc]
- Reconfigurable Compute-In-Memory on Field-Programmable Ferroelectric Diodes. [europepmc]
- Hardware-aware training for large-scale and diverse deep learning inference workloads using in-memory computing-based accelerators. [europepmc]
- Powering AI at the edge: A robust, memristor-based binarized neural network with near-memory computing and miniaturized solar cell. [europepmc]
- Domain wall magnetic tunnel junction-based artificial synapses and neurons for all-spin neuromorphic hardware. [europepmc]
- All-optical ultrafast ReLU function for energy-efficient nanophotonic deep learning. [europepmc]
Related