vix.ing · top · new · best · stats · spec

Koios: A Deep Learning Benchmark Suite for FPGA Architecture and CAD\n Research

2021/06/13 by Aman Arora, Andrew Boutros, Arora, Aman +21 · 2 citations
Computer Science · Engineering · #FOS: Computer and information sciences #Hardware Architecture (cs.AR) #Low-power high-performance VLSI design #VLSI and Analog Circuit Testing #VLSI and FPGA Design Techniques

paper · pdf · doi:10.48550/arxiv.2106.07087

openalex publication_date 2021/06/13 · openalex created_date 2022/07/25 · openalex updated_date 2026/07/28

Abstract

With the prevalence of deep learning (DL) in many applications, researchers\nare investigating different ways of optimizing FPGA architecture and CAD to\nachieve better quality-of-results (QoR) on DL-based workloads. In this\noptimization process, benchmark circuits are an essential component; the QoR\nachieved on a set of benchmarks is the main driver for architecture and CAD\ndesign choices. However, current academic benchmark suites are inadequate, as\nthey do not capture any designs from the DL domain. This work presents a new\nsuite of DL acceleration benchmark circuits for FPGA architecture and CAD\nresearch, called Koios. This suite of 19 circuits covers a wide variety of\naccelerated neural networks, design sizes, implementation styles, abstraction\nlevels, and numerical precisions. These designs are larger, more data parallel,\nmore heterogeneous, more deeply pipelined, and utilize more FPGA architectural\nfeatures compared to existing open-source benchmarks. This enables researchers\nto pin-point architectural inefficiencies for this class of workloads and\noptimize CAD tools on more realistic benchmarks that stress the CAD algorithms\nin different ways. In this paper, we describe the designs in our benchmark\nsuite, present results of running them through the Verilog-to-Routing (VTR)\nflow using a recent FPGA architecture model, and identify key insights from the\nresulting metrics. On average, our benchmarks have 3.7x more netlist\nprimitives, 1.8x and 4.7x higher DSP and BRAM densities, and 1.7x higher\nfrequency with 1.9x more near-critical paths compared to the widely-used VTR\nsuite. Finally, we present two example case studies showing how architectural\nexploration for DL-optimized FPGAs can be performed using our new benchmark\nsuite.\n

Cited by

Related