vix.ing · top · new · best · stats · spec

Failure Analysis and Quantification for Contemporary and Future Supercomputers

2019/11/05 by Li Tan, Tan, Li, Nathan DeBardeleben +1
Computer Science · Engineering · #68M15 #68M20 #68N20 #Distributed #Distributed systems and fault tolerance #FOS: Computer and information sciences #Parallel #Radiation Effects in Electronics #Reliability and Maintenance Optimization #and Cluster Computing (cs.DC)

paper · pdf · doi:10.48550/arxiv.1911.02118

openalex publication_date 2019/11/05 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Large-scale computing systems today are assembled by numerous computing units for massive computational capability needed to solve problems at scale, which enables failures common events in supercomputing scenarios. Considering the demanding resilience requirements of supercomputers today, we present a quantitative study on fine-grained failure modeling for contemporary and future large-scale computing systems. We integrate various types of failures from different system hierarchical levels and system components, and summarize the overall system failure rates formally. Given that nowadays system-wise failure rate needs to be capped under a threshold value for reliability and cost-efficiency purposes, we quantitatively discuss different scenarios of system resilience, and analyze the impacts of resilience to different error types on the variation of system failure rates, and the correlation of hierarchical failure rates. Moreover, we formalize and showcase the resilience efficiency of failure-bounded supercomputers today.

Related