2021/10/11 by Martin Uray, Eduard Hirsch, Uray, Martin +5
Computer Science · Decision Sciences · #Artificial Intelligence (cs.AI) #Cloud Computing and Resource Management #Distributed #Distributed and Parallel Computing Systems #FOS: Computer and information sciences #Parallel #Scientific Computing and Data Management #and Cluster Computing (cs.DC)
paper · pdf · doi:10.48550/arxiv.2110.05156
openalex publication_date 2021/10/11 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Enterprises and labs performing computationally expensive data science applications sooner or later face the problem of scale but unconnected infrastructure. For this up-scaling process, an IT service provider can be hired or in-house personnel can attempt to implement a software stack. The first option can be quite expensive if it is just about connecting several machines. For the latter option often experience is missing with the data science staff in order to navigate through the software jungle. In this technical report, we illustrate the decision process towards an on-premises infrastructure, our implemented system architecture, and the transformation of the software stack towards a scaleable GPU cluster system.