2018/05/10 by Zhe Li, Li, Zhe, Li Ji +14 · 1 voice
Computer Science · #Advanced Neural Network Applications #Algorithm #Artificial intelligence #Artificial neural network #Computation #Computer architecture #Computer engineering #Computer hardware #Computer science #Convolutional neural network #Deep learning #Distributed computing #Electric power system #Emerging Technologies (cs.ET) #Error Correcting Code Techniques #FOS: Computer and information sciences #Field-programmable gate array #Hardware acceleration #Mobile device #Neural and Evolutionary Computing (cs.NE) #Operating system #Parallel computing #Power (physics) #Power budget #Scalability #Speedup #Stochastic Gradient Optimization Techniques #Stochastic computing #cs.ET #cs.NE
paper · pdf · doi:10.48550/arxiv.1805.04142
published in arXiv (Cornell University) (Cornell University) · Accepted by IEEE Computer Society Annual Symposium on VLSI 2018
arxiv created 2018/05/10 · openalex publication_date 2018/05/10 · arxiv published 2018/05/10 · arxiv updated 2018/05/14 · openalex created_date 2019/06/27 · openalex updated_date 2026/08/08
Recently, Deep Convolutional Neural Network (DCNN) has achieved tremendous success in many machine learning applications. Nevertheless, the deep structure has brought significant increases in computation complexity. Largescale deep learning systems mainly operate in high-performance server clusters, thus restricting the application extensions to personal or mobile devices. Previous works on GPU and/or FPGA acceleration for DCNNs show increasing speedup, but ignore other constraints, such as area, power, and energy. Stochastic Computing (SC), as a unique data representation and processing technique, has the potential to enable the design of fully parallel and scalable hardware implementations of large-scale deep learning systems. This paper proposed an automatic design allocation algorithm driven by budget requirement considering overall accuracy performance. This systematic method enables the automatic design of a DCNN where all design parameters are jointly optimized. Experimental results demonstrate that proposed algorithm can achieve a joint optimization of all design parameters given the comprehensive budget of a DCNN.