vix.ing · top · new · best · stats · spec

DLAU: A Scalable Deep Learning Accelerator Unit on FPGA

2016/05/23 by Chao Wang, Wang, Chao, Qi Yu +9 · 1 citation
Computer Science · Engineering · #Advanced Image and Video Retrieval Techniques #Advanced Neural Network Applications #CCD and CMOS Imaging Sensors #Distributed #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural and Evolutionary Computing (cs.NE) #Parallel #and Cluster Computing (cs.DC)

paper · pdf · doi:10.48550/arxiv.1605.06894

openalex publication_date 2016/05/23 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

As the emerging field of machine learning, deep learning shows excellent ability in solving complex learning problems. However, the size of the networks becomes increasingly large scale due to the demands of the practical applications, which poses significant challenge to construct a high performance implementations of deep learning neural networks. In order to improve the performance as well to maintain the low power cost, in this paper we design DLAU, which is a scalable accelerator architecture for large-scale deep learning networks using FPGA as the hardware prototype. The DLAU accelerator employs three pipelined processing units to improve the throughput and utilizes tile techniques to explore locality for deep learning applications. Experimental results on the state-of-the-art Xilinx FPGA board demonstrate that the DLAU accelerator is able to achieve up to 36.1x speedup comparing to the Intel Core2 processors, with the power consumption at 234mW.

Cited by

Related