Ternary Weight Networks
2016/05/16 by Fengfu Li, Bin Liu, Li, Fengfu +7 · 152 citations
Computer Science · #Advanced Neural Network Applications #Human Pose and Action Recognition #Multimodal Machine Learning Applications #cs.CV
paper · pdf · doi:10.48550/arxiv.1605.04711
5 pages, 3 fitures, conference
arxiv created 2022/11/20 · arxiv updated 2022/11/22
Abstract
We present a memory and computation efficient ternary weight networks (TWNs) - with weights constrained to +1, 0 and -1. The Euclidian distance between full (float or double) precision weights and the ternary weights along with a scaling factor is minimized in training stage. Besides, a threshold-based ternary function is optimized to get an approximated solution which can be fast and easily computed. TWNs have shown better expressive abilities than binary precision counterparts. Meanwhile, TWNs achieve up to 16× model compression rate and need fewer multiplications compared with the float32 precision counterparts. Extensive experiments on MNIST, CIFAR-10, and ImageNet datasets show that the TWNs achieve much better result than the Binary-Weight-Networks (BWNs) and the classification performance on MNIST and CIFAR-10 is very close to the full precision networks. We also verify our method on object detection task and show that TWNs significantly outperforms BWN by more than 10% mAP on PASCAL VOC dataset. The pytorch version of source code is available at: https://github.com/Thinklab-SJTU/twns.
Citations
Cited by
- R2Q: Towards Robust 2-Bit Large Language Models via Residual Refinement Quantization
- Extreme Model Compression with Structured Sparsity at Low Precision
- Bayesian Bits: Unifying Quantization and Pruning
- A Survey on Methods and Theories of Quantized Neural Networks
- Learning Architectures for Binary Networks
- Training Quantized Nets: A Deeper Understanding
- LUTNet: Rethinking Inference in FPGA Soft Logic
- And the Bit Goes Down: Revisiting the Quantization of Neural Networks
- Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis
- Model compression as constrained optimization, with application to neural nets. Part I: general framework
- A Survey of FPGA-Based Robotic Computing
- Differentiable Soft Quantization: Bridging Full-Precision and Low-Bit Neural Networks
- AirNN: Over-the-Air Computation for Neural Networks via Reconfigurable Intelligent Surfaces
- Mix and Match: A Novel FPGA-Centric Deep Neural Network Quantization Framework
- Neural Networks Weights Quantization: Target None-retraining Ternary (TNT)
- Improving Network Slimming with Nonconvex Regularization
- On Periodic Functions as Regularizers for Quantization of Neural Networks
- DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients
- A Survey of FPGA-Based Neural Network Accelerator
- Differentiable, Bit-shifting, and Scalable Quantization without training neural network from scratch
- DNN Quantization with Attention
- Communication-Efficient Federated Distillation
- TernaryCLIP: Efficiently Compressing Vision-Language Models with Ternary Weights and Distilled Knowledge
- Differentiable Model Compression via Pseudo Quantization Noise
- Deep Learning for Real-Time Crime Forecasting and its Ternarization
- BinaryBERT: Pushing the Limit of BERT Quantization
- Quantization of Fully Convolutional Networks for Accurate Biomedical Image Segmentation
- Scalable Model Compression by Entropy Penalized Reparameterization
- Joint Architecture and Knowledge Distillation in CNN for Chinese Text Recognition
- Tequila: Trapping-free Ternary Quantization for Large Language Models
- PT2-LLM: Post-Training Ternarization for Large Language Models
- Automatic Neural Network Compression by Sparsity-Quantization Joint Learning: A Constrained Optimization-based Approach
- Knowledge Squeezed Adversarial Network Compression
- Piggyback: Adapting a Single Network to Multiple Tasks by Learning to Mask Weights
- Trained Ternary Quantization
- Rethinking floating point for deep learning
- Model Compression and Hardware Acceleration for Neural Networks: A Comprehensive Survey
- Mitigate Parasitic Resistance in Resistive Crossbar-based Convolutional Neural Networks
- TernaryBERT: Distillation-aware Ultra-low Bit BERT
- Training and Inference with Integers in Deep Neural Networks
- Latent-Space Mean-Field Theory for Deep BitNet-like Training: Constrained Gradient Flows with Smooth Quantization and STE Limits
- Ternary Neural Networks with Fine-Grained Quantization
- Efficient Integer-Arithmetic-Only Convolutional Neural Networks
- Blended Coarse Gradient Descent for Full Quantization of Deep Neural Networks
- Mixed Low-precision Deep Learning Inference using Dynamic Fixed Point
- Recent Advances in Efficient Computation of Deep Convolutional Neural Networks
- Neural Network Quantization for Microcontrollers: A Comprehensive Survey of Methods, Platforms, and Applications
- ReActNet: Towards Precise Binary Neural Network with Generalized Activation Functions
- SOTERIA: In Search of Efficient Neural Networks for Private Inference
- Compressive Meta-Learning
- Training DNNs with Hybrid Block Floating Point
- Lightweight Neural Networks
- CPT: Efficient Deep Neural Network Training via Cyclic Precision
- Model compression as constrained optimization, with application to neural nets. Part V: combining compressions
- Exploration of Low Numeric Precision Deep Learning Inference Using Intel FPGAs
- A Lower Bound for the Number of Linear Regions of Ternary ReLU Regression Neural Networks
- Low-memory convolutional neural networks through incremental depth-first processing
- RPR: Random Partition Relaxation for Training; Binary and Ternary Weight Neural Networks
- Reward-Based 1-bit Compressed Federated Distillation on Blockchain
- Localization-aware Channel Pruning for Object Detection
- Dynamic Runtime Feature Map Pruning
- Intermediate Deep Feature Compression: the Next Battlefield of Intelligent Sensing
- Towards Efficient Post-training Quantization of Pre-trained Language Models
- Learned Step Size Quantization
- DFQ-ViT: Data-Free Quantization for Vision Transformers without Fine-tuning
- RTN: Reparameterized Ternary Network
- Efficient and Robust Machine Learning for Real-World Systems
- DecisiveNets: Training Deep Associative Memories to Solve Complex Machine Learning Problems
- Reducing Inference Latency with Concurrent Architectures for Image Recognition
- Pruning Ternary Quantization
- Binary Ensemble Neural Network: More Bits per Network or More Networks per Bit?
- SquishedNets: Squishing SqueezeNet further for edge device scenarios via deep evolutionary synthesis
- Quantization for Rapid Deployment of Deep Neural Networks
- A Survey of FPGA Based Deep Learning Accelerators: Challenges and Opportunities
- Ternary Residual Networks
- An FPGA Accelerated Method for Training Feed-forward Neural Networks Using Alternating Direction Method of Multipliers and LSMR
- Simultaneously Optimizing Weight and Quantizer of Ternary Neural Network using Truncated Gaussian Approximation
- Direct Quantization for Training Highly Accurate Low Bit-width Deep Neural Networks
- MWQ: Multiscale Wavelet Quantized Neural Networks
- Mixed Precision DNN Qunatization for Overlapped Speech Separation and Recognition
- StrassenNets: Deep Learning with a Multiplication Budget
- Towards End-to-End Neural Face Authentication in the Wild -- Quantifying and Compensating for Directional Lighting Effects
- Towards Lossless Binary Convolutional Neural Networks Using Piecewise Approximation
- ShiftCNN: Generalized Low-Precision Architecture for Inference of Convolutional Neural Networks
- Towards Efficient Training for Neural Network Quantization
- Model compression as constrained optimization, with application to neural nets. Part II: quantization
- Distillation Guided Residual Learning for Binary Convolutional Neural Networks
- BinaryRelax: A Relaxation Approach For Training Deep Neural Networks With Quantized Weights
- Adaptive Precision Training for Resource Constrained Devices
- Structured Convolutions for Efficient Neural Network Design
- Binarized Weight Error Networks With a Transition Regularization Term
- Precision Gating: Improving Neural Network Efficiency with Dynamic Dual-Precision Activations
- PBGen: Partial Binarization of Deconvolution-Based Generators for Edge Intelligence
- n-hot: Efficient bit-level sparsity for powers-of-two neural network quantization
- Structured Compression by Weight Encryption for Unstructured Pruning and Quantization
- Quantization of Deep Neural Networks for Accurate Edge Computing
- A Lite Distributed Semantic Communication System for Internet of Things
- Demystifying and Generalizing BinaryConnect
- Towards Mixed-Precision Quantization of Neural Networks via Constrained Optimization
- HOTCAKE: Higher Order Tucker Articulated Kernels for Deeper CNN Compression
- Weight Normalization based Quantization for Deep Neural Network Compression
- Prune Your Model Before Distill It
- Heterogeneous Bitwidth Binarization in Convolutional Neural Networks
- ZeroQ: A Novel Zero Shot Quantization Framework
- FQ-Conv: Fully Quantized Convolution for Efficient and Accurate Inference
- Accelerating CNN inference on FPGAs: A Survey
- Binarized Neural Architecture Search for Efficient Object Recognition
- Low-bit Quantization of Recurrent Neural Network Language Models Using Alternating Direction Methods of Multipliers
- Loss-aware Weight Quantization of Deep Networks
- Towards thinner convolutional neural networks through Gradually Global Pruning
- Resource-Efficient Neural Networks for Embedded Systems
- End-to-End Learned Image Compression with Quantized Weights and Activations
- Deep Learning with Low Precision by Half-wave Gaussian Quantization
- Toward Compact Parameter Representations for Architecture-Agnostic Neural Network Compression
- Cross-filter compression for CNN inference acceleration
- BNAS v2: Learning Architectures for Binary Networks with Empirical Improvements
- Optimizing LLMs for Resource-Constrained Environments: A Survey of Model Compression Techniques
- Learning Recurrent Binary/Ternary Weights
- Musical Chair: Efficient Real-Time Recognition Using Collaborative IoT Devices
- LoTA-QAF: Lossless Ternary Adaptation for Quantization-Aware Fine-Tuning
- QKD: Quantization-aware Knowledge Distillation
- LANCE: Efficient Low-Precision Quantized Winograd Convolution for Neural Networks Based on Graphics Processing Units
- Discrimination-aware Channel Pruning for Deep Neural Networks
- Forward and Backward Information Retention for Accurate Binary Neural Networks
- BitSplit-Net: Multi-bit Deep Neural Network with Bitwise Activation Function
- Entropy-Based Modeling for Estimating Soft Errors Impact on Binarized Neural Network Inference
- An Optimal Control Approach to Deep Learning and Applications to Discrete-Weight Neural Networks
- Attend to Your Own Thoughts: Breaking the Barrier for Post-Training Quantization of Reasoning LLMs through the Lens of 1.58-Bit Quantization
- Optimizing Binary and Ternary Neural Network Inference on RRAM Crossbars using CIM-Explorer
- Attention Based Pruning for Shift Networks
- Quantized Neural Networks via -1, +1 Encoding Decomposition and Acceleration
- Mixed Precision DNNs: All you need is a good parametrization
- WRPN: Wide Reduced-Precision Networks
- DNQ: Dynamic Network Quantization
- Deep Molecular Programming: A Natural Implementation of Binary-Weight ReLU Neural Networks
- Faster Convolution Inference Through Using Pre-Calculated Lookup Tables
- NASB: Neural Architecture Search for Binary Convolutional Neural Networks
- Evaluating the Scalability of Binary and Ternary CNN Workloads on RRAM-based Compute-in-Memory Accelerators
- Bridging the Accuracy Gap for 2-bit Quantized Neural Networks (QNN)
- LCP: A Low-Communication Parallelization Method for Fast Neural Network Inference in Image Recognition
- CAT-Q: Cost-efficient and Accurate Ternary Quantization for LLMs
- Demystifying Parallel and Distributed Deep Learning
- Quantization and Training of Low Bit-Width Convolutional Neural Networks for Object Detection
- An end to end Deep Neural Network for iris segmentation in unconstrained scenarios
- Low-Precision Batch-Normalized Activations
- Differentiable Neural Architecture Learning for Efficient Neural Network Design
- BackSlash: Rate Constrained Optimized Training of Large Language Models
- The ZipML Framework for Training Models with End-to-End Low Precision: The Cans, the Cannots, and a Little Bit of Deep Learning
- Network Quantization with Element-wise Gradient Scaling
- In-situ Stochastic Training of MTJ Crossbar based Neural Networks
- BiDet: An Efficient Binarized Object Detector
- APQF: Agentic Profiling-Guided Structured Pruning and Mixed-Precision Quantization with Adaptive Fine-Tuning
Related