vix.ing · top · new · best · stats · spec

Designing Efficient and High-performance AI Accelerators with Customized\n STT-MRAM

2021/04/05 by Kaniz Mishty, Mishty, Kaniz, Mehdi Sadi +1
Engineering · #Advanced Memory and Neural Computing #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Ferroelectric and Negative Capacitance Devices #Hardware Architecture (cs.AR) #Semiconductor materials and devices

paper · pdf · doi:10.48550/arxiv.2104.02199

openalex publication_date 2021/04/05 · openalex created_date 2022/07/25 · openalex updated_date 2026/07/28

Abstract

In this paper, we demonstrate the design of efficient and high-performance\nAI/Deep Learning accelerators with customized STT-MRAM and a reconfigurable\ncore. Based on model-driven detailed design space exploration, we present the\ndesign methodology of an innovative scratchpad-assisted on-chip STT-MRAM based\nbuffer system for high-performance accelerators. Using analytically derived\nexpression of memory occupancy time of AI model weights and activation maps,\nthe volatility of STT-MRAM is adjusted with process and temperature variation\naware scaling of thermal stability factor to optimize the retention time,\nenergy, read/write latency, and area of STT-MRAM. From the analysis of modern\nAI workloads and accelerator implementation in 14nm technology, we verify the\nefficacy of our designed AI accelerator with STT-MRAM STT-AI. Compared to an\nSRAM-based implementation, the STT-AI accelerator achieves 75% area and 3%\npower savings at iso-accuracy. Furthermore, with a relaxed bit error rate and\nnegligible AI accuracy trade-off, the designed STT-AI Ultra accelerator\nachieves 75.4%, and 3.5% savings in area and power, respectively over regular\nSRAM-based accelerators.\n

Citations

Related