2018/01/18 by Jung-Woo Chang, Keon-Woo Kang, Suk-Ju Kang +1 · 1 voice
Computer Science · Engineering · #Advanced Image Processing Techniques #Advanced Vision and Imaging #Image Processing Techniques and Applications #cs.AR #cs.DC #msc:68U10
paper · pdf · doi:10.1109/tcsvt.2018.2888898
Accepted for publication in IEEE Transactions on Circuits and Systems for Video Technology (TCSVT)
arxiv published 2018/01/18 · arxiv created 2018/12/18 · arxiv updated 2018/12/19 · openalex publication_date 2018/12/20 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Convolutional neural networks (CNNs) demonstrate excellent performance in various computer vision applications. In recent years, FPGA-based CNN accelerators have been proposed for optimizing performance and power efficiency. Most accelerators are designed for object detection and recognition algorithms that are performed on low-resolution (LR) images. However, real-time image super-resolution (SR) cannot be implemented on a typical accelerator because of the long execution cycles required to generate high-resolution (HR) images, such as those used in ultra-high-definition (UHD) systems. In this paper, we propose a novel CNN accelerator with efficient parallelization methods for SR applications. First, we propose a new methodology for optimizing the deconvolutional neural networks (DCNNs) used for increasing feature maps. Secondly, we propose a novel method to optimize CNN dataflow so that the SR algorithm can be driven at low power in display applications. Finally, we quantize and compress a DCNN-based SR algorithm into an optimal model for efficient inference using on-chip memory. We present an energy-efficient architecture for SR and validate our architecture on a mobile panel with quad-high-definition (QHD) resolution. Our experimental results show that, with the same hardware resources, the proposed DCNN accelerator achieves a throughput up to 108 times greater than that of a conventional DCNN accelerator. In addition, our SR system achieves an energy efficiency of 144.9 GOPS/W, 293.0 GOPS/W, and 500.2 GOPS/W at SR scale factors of 2, 3, and 4, respectively. Furthermore, we demonstrate that our system can restore HR images to a high quality while greatly reducing the data bit-width and the number of parameters compared to conventional SR algorithms.