2023/09/06 by Shanzhi Yin, Tongda Xu, Yin, Shanzhi +11
Computer Science · Engineering · #68U10(primary) #94A08 68T07(secondary) #Advanced Data Compression Techniques #CCD and CMOS Imaging Sensors #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #I.2.6 #I.4.2 #Image and Video Processing (eess.IV) #Neural Networks and Applications #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2309.02855
openalex publication_date 2023/09/06 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
With neural networks growing deeper and feature maps growing larger, limited communication bandwidth with external memory (or DRAM) and power constraints become a bottleneck in implementing network inference on mobile and edge devices. In this paper, we propose an end-to-end differentiable bandwidth efficient neural inference method with the activation compressed by neural data compression method. Specifically, we propose a transform-quantization-entropy coding pipeline for activation compression with symmetric exponential Golomb coding and a data-dependent Gaussian entropy model for arithmetic coding. Optimized with existing model quantization methods, low-level task of image compression can achieve up to 19x bandwidth reduction with 6.21x energy saving.