2021/07/26 by Piyush Subedi, Subedi, Piyush, Jianwei Hao +5
Computer Science · #Cloud Computing and Resource Management #Distributed #FOS: Computer and information sciences #IoT and Edge/Fog Computing #Parallel #Stochastic Gradient Optimization Techniques #and Cluster Computing (cs.DC)
paper · pdf · doi:10.48550/arxiv.2107.12486
openalex publication_date 2021/07/26 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Many real-world applications are widely adopting the edge computing paradigm\ndue to its low latency and better privacy protection. With notable success in\nAI and deep learning (DL), edge devices and AI accelerators play a crucial role\nin deploying DL inference services at the edge of the Internet. While prior\nworks quantified various edge devices' efficiency, most studies focused on the\nperformance of edge devices with single DL tasks. Therefore, there is an urgent\nneed to investigate AI multi-tenancy on edge devices, required by many advanced\nDL applications for edge computing.\n This work investigates two techniques - concurrent model executions and\ndynamic model placements - for AI multi-tenancy on edge devices. With image\nclassification as an example scenario, we empirically evaluate AI multi-tenancy\non various edge devices, AI accelerators, and DL frameworks to identify its\nbenefits and limitations. Our results show that multi-tenancy significantly\nimproves DL inference throughput by up to 3.3x -- 3.8x on Jetson TX2. These AI\nmulti-tenancy techniques also open up new opportunities for flexible deployment\nof multiple DL services on edge devices and AI accelerators.\n