2019/01/21 by Amit Samanta, Suhas Shrinivasan, Samanta, Amit +5 · 1 citation
Computer Science · #Advanced Neural Network Applications #Distributed #FOS: Computer and information sciences #IoT and Edge/Fog Computing #Parallel #Privacy-Preserving Technologies in Data #and Cluster Computing (cs.DC) #cs.DC
paper · pdf · doi:10.48550/arxiv.1901.06887
5 pages
openalex publication_date 2019/01/21 · arxiv created 2019/01/23 · arxiv updated 2019/01/24 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
With the rise of machine learning, inference on deep neural networks (DNNs) has become a core building block on the critical path for many cloud applications. Applications today rely on isolated ad-hoc deployments that force users to compromise on consistent latency, elasticity, or cost-efficiency, depending on workload characteristics. We propose to elevate DNN inference to be a first class cloud primitive provided by a shared multi-tenant system, akin to cloud storage, and cloud databases. A shared system enables cost-efficient operation with consistent performance across the full spectrum of workloads. We argue that DNN inference is an ideal candidate for a multi-tenant system because of its narrow and well-defined interface and predictable resource requirements.