vix.ing · top · new · best · stats · spec

Compressing Low Precision Deep Neural Networks Using Sparsity-Induced\n Regularization in Ternary Networks

2017/09/19 by Julian Faraone, Faraone, Julian, Nicholas C. Fraser +7
Computer Science · #Advanced Image and Video Retrieval Techniques #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Neural Networks and Applications

paper · pdf · doi:10.48550/arxiv.1709.06262

openalex publication_date 2017/09/19 · openalex created_date 2022/10/06 · openalex updated_date 2026/07/28

Abstract

A low precision deep neural network training technique for producing sparse,\nternary neural networks is presented. The technique incorporates hard- ware\nimplementation costs during training to achieve significant model compression\nfor inference. Training involves three stages: network training using L2\nregularization and a quantization threshold regularizer, quantization pruning,\nand finally retraining. Resulting networks achieve improved accuracy, reduced\nmemory footprint and reduced computational complexity compared with\nconventional methods, on MNIST and CIFAR10 datasets. Our networks are up to 98%\nsparse and 5 & 11 times smaller than equivalent binary and ternary models,\ntranslating to significant resource and speed benefits for hardware\nimplementations.\n

Related