2020/04/20 by Dionysios Diamantopoulos, Burkhard Ringlein, Diamantopoulos, Dionysios +7 · 2 citations
Computer Science · #Distributed #Embedded Systems Design Techniques #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural and Evolutionary Computing (cs.NE) #Parallel #Parallel Computing and Optimization Techniques #Teaching and Learning Programming #and Cluster Computing (cs.DC)
paper · pdf · doi:10.48550/arxiv.2004.10854
openalex publication_date 2020/04/20 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Specialized accelerators for tensor-operations, such as blocked-matrix\noperations and multi-dimensional convolutions, have been emerged as powerful\narchitecture choices for high-performance Deep-Learning computing. The rapid\ndevelopment of frameworks, models, and precision options challenges the\nadaptability of such tensor-accelerators since the adaptation to new\nrequirements incurs significant engineering costs. Programmable tensor\naccelerators offer a promising alternative by allowing reconfiguration of a\nvirtual architecture that overlays on top of the physical FPGA configurable\nfabric. We propose an overlay (\τ-VTA) and an optimization method guided by\nagile-inspired auto-tuning techniques. We achieve higher performance and faster\nconvergence than state-of-art.\n