vix.ing · top · new · best · stats

Beyond the Memory Wall: A Case for Memory-centric HPC System for Deep Learning

2019/02/18 by Youngeun Kwon, Kwon, Youngeun, Minsoo Rhu +1 · 4 citations
Computer Science · Engineering · #Advanced Data Storage Technologies #Advanced Memory and Neural Computing #Distributed #FOS: Computer and information sciences #Ferroelectric and Negative Capacitance Devices #Hardware Architecture (cs.AR) #Machine Learning (cs.LG) #Neural and Evolutionary Computing (cs.NE) #Parallel #Parallel Computing and Optimization Techniques #and Cluster Computing (cs.DC) #cs.AR #cs.DC #cs.LG #cs.NE

paper · pdf · doi:10.48550/arxiv.1902.06468

Published as a conference paper at the 51st IEEE/ACM International Symposium on Microarchitecture (MICRO-51), 2018

arxiv created 2019/02/18 · openalex publication_date 2019/02/18 · arxiv updated 2019/02/19 · openalex created_date 2022/07/29 · openalex updated_date 2026/07/28

Abstract

As the models and the datasets to train deep learning (DL) models scale, system architects are faced with new challenges, one of which is the memory capacity bottleneck, where the limited physical memory inside the accelerator device constrains the algorithm that can be studied. We propose a memory-centric deep learning system that can transparently expand the memory capacity available to the accelerators while also providing fast inter-device communication for parallel training. Our proposal aggregates a pool of memory modules locally within the device-side interconnect, which are decoupled from the host interface and function as a vehicle for transparent memory capacity expansion. Compared to conventional systems, our proposal achieves an average 2.8x speedup on eight DL applications and increases the system-wide memory capacity to tens of TBs.

Citations

Cited by

Related