2018/07/11 by Vladimir Rybalkin, Rybalkin, Vladimir, Alessandro Pappalardo +9 · 2 citations
Computer Science · Neuroscience · #Advanced Neural Network Applications #Brain Tumor Detection and Classification #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Hardware Architecture (cs.AR) #Machine Learning (cs.LG) #Neural Networks and Applications
paper · pdf · doi:10.48550/arxiv.1807.04093
openalex publication_date 2018/07/11 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
It is well known that many types of artificial neural networks, including\nrecurrent networks, can achieve a high classification accuracy even with\nlow-precision weights and activations. The reduction in precision generally\nyields much more efficient hardware implementations in regards to hardware\ncost, memory requirements, energy, and achievable throughput. In this paper, we\npresent the first systematic exploration of this design space as a function of\nprecision for Bidirectional Long Short-Term Memory (BiLSTM) neural network.\nSpecifically, we include an in-depth investigation of precision vs. accuracy\nusing a fully hardware-aware training flow, where during training quantization\nof all aspects of the network including weights, input, output and in-memory\ncell activations are taken into consideration. In addition, hardware resource\ncost, power consumption and throughput scalability are explored as a function\nof precision for FPGA-based implementations of BiLSTM, and multiple approaches\nof parallelizing the hardware. We provide the first open source HLS library\nextension of FINN for parameterizable hardware architectures of LSTM layers on\nFPGAs which offers full precision flexibility and allows for parameterizable\nperformance scaling offering different levels of parallelism within the\narchitecture. Based on this library, we present an FPGA-based accelerator for\nBiLSTM neural network designed for optical character recognition, along with\nnumerous other experimental proof points for a Zynq UltraScale+ XCZU7EV MPSoC\nwithin the given design space.\n