Tensor Manipulation Unit (TMU): Reconfigurable, Near-Memory Tensor Manipulation for High-Throughput AI SoC
2025/06/17 by Weiyu Zhou, Zheng Wang, Zhou, Weiyu +11 · 10 voices
Computer Science · Mathematics · #Parallel Computing and Optimization Techniques #Tensor decomposition and applications #Advanced Neural Network Applications
paper · pdf · doi:10.48550/arxiv.2506.14364
Abstract
While recent advances in AI SoC design have focused heavily on accelerating tensor computation, the equally critical task of tensor manipulation, centered on high,volume data movement with minimal computation, remains underexplored. This work addresses that gap by introducing the Tensor Manipulation Unit (TMU), a reconfigurable, near-memory hardware block designed to efficiently execute data-movement-intensive operators. TMU manipulates long datastreams in a memory-to-memory fashion using a RISC-inspired execution model and a unified addressing abstraction, enabling broad support for both coarse- and fine-grained tensor transformations. Integrated alongside a TPU within a high-throughput AI SoC, the TMU leverages double buffering and output forwarding to improve pipeline utilization. Fabricated in SMIC 40nm technology, the TMU occupies only 0.019 mm2 while supporting over 10 representative tensor manipulation operators. Benchmarking shows that TMU alone achieves up to 1413 and 8.54 operator-level latency reduction compared to ARM A72 and NVIDIA Jetson TX2, respectively. When integrated with the in-house TPU, the complete system achieves a 34.6% reduction in end-to-end inference latency, demonstrating the effectiveness and scalability of reconfigurable tensor manipulation in modern AI SoCs.
Citations
Discussions
- Tensor Manipulation Unit (TMU): Reconfigurable, Near-Memory, High-Throughput AI [hn, 58 points, 13 comments]
- Блок манипуляции тензорами (TMU): Перестраиваемый, Близкий к Памяти, Высокопроизводительный ИИ #ai #news [bsky, 0 points, 0 comments]
- Tensor Manipulation Unit (TMU): Reconfigurable, Near-Memory, High-Throughput AI https://arxiv.org/abs/2506.14364 (https://news.ycombinator.com/item?id=44351798) [bsky, 0 points, 0 comments]
- Tensor Manipulation Unit (TMU): Reconfigurable, Near-Memory, High-Throughput AI https://arxiv.org/abs/2506.14364 (https://news.ycombinator.com/item?id=44351798) [bsky, 0 points, 0 comments]
- Tensor Manipulation Unit (TMU): Reconfigurable, Near-Memory, High-Throughput AI [bsky, 0 points, 0 comments]
- Tensor Manipulation Unit (TMU): Reconfigurable, Near-Memory, High-Throughput AI #ai #news [bsky, 0 points, 0 comments]
- Tensor Manipulation Unit (TMU): Reconfigurable, Near-Memory, High-Throughput AI View Article | Join the HN Conversation Summary of HN discussion 🧵👇 #hacker-news [bsky, 0 points, 1 comments]
- Interesting approach to accelerating AI workloads! Reconfigurable compute near memory could significantly boost performance and efficiency. A promising direction. 🤖 #ai Tensor Manipulation Unit (TMU) [bsky, 0 points, 0 comments]
- ⚡ Hackernews Top story: Tensor Manipulation Unit (TMU): Reconfigurable, Near-Memory, High-Throughput AI [bsky, 0 points, 0 comments]
- Tensor Manipulation Unit (TMU): Reconfigurable, Near-Memory, High-Throughput AI https://arxiv.org/abs/2506.14364 [bsky, 0 points, 0 comments]
Related