2024/04/18 by Qizhe Wu, Wu, Qizhe, Yuchen Gui +9
Computer Science · Engineering · #FOS: Computer and information sciences #Hardware Architecture (cs.AR) #Low-power high-performance VLSI design #Parallel Computing and Optimization Techniques
paper · pdf · doi:10.48550/arxiv.2404.11887
openalex publication_date 2024/04/18 · openalex created_date 2024/04/20 · openalex updated_date 2026/07/28
Tensor computations, with matrix multiplication being the primary operation, serve as the fundamental basis for data analysis, physics, machine learning, and deep learning. As the scale and complexity of data continue to grow rapidly, the demand for tensor computations has also increased significantly. To meet this demand, several research institutions have started developing dedicated hardware for tensor computations. To further improve the computational performance of tensor process units, we have reexamined the issue of computation reuse that was previously overlooked in existing architectures. As a result, we propose a novel EN-T architecture that can reduce chip area and power consumption. Furthermore, our method is compatible with existing tensor processing units. We evaluated our method on prevalent microarchitectures, the results demonstrate an average improvement in area efficiency of 8.7%, 12.2%, and 11.0% for tensor computing units at computational scales of 256 GOPS, 1 TOPS, and 4 TOPS, respectively. Similarly, there were energy efficiency enhancements of 13.0%, 17.5%, and 15.5%.