2021/08/25 by Chansup Byun, William Arcand, David Bestor +16 · 9 citations
Computer Science · Engineering · #Big data #Cloud Computing and Resource Management #Cloud computing #Computer science #Distributed and Parallel Computing Systems #Distributed computing #Engineering #Job queue #Job scheduler #Node (physics) #Operating system #Parallel Computing and Optimization Techniques #Parallel computing #Processor scheduling #Schedule #Scheduling (production processes) #Supercomputer #cs.DC
paper · pdf · doi:10.1109/hpec49654.2021.9622870
IEEE HPEC 2021
arxiv created 2021/08/25 · openalex publication_date 2021/09/20 · arxiv updated 2021/12/13 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
Diverse workloads such as interactive supercomputing, big data analysis, and large-scale AI algorithm development, requires a high-performance scheduler. This paper presents a novel node-based scheduling approach for large scale simulations of short running jobs on MIT SuperCloud systems, that allows the resources to be fully utilized for both long running batch jobs while simultaneously providing fast launch and release of large-scale short running jobs. The node-based scheduling approach has demonstrated up to 100 times faster scheduler performance that other state-of-the-art systems.