2017/08/28 by Nils Kohl, Kohl, Nils, Johannes Hötzer +13
Computer Science · Decision Sciences · #Advanced Data Storage Technologies #Distributed #Distributed and Parallel Computing Systems #FOS: Computer and information sciences #Parallel #Simulation Techniques and Applications #and Cluster Computing (cs.DC)
paper · pdf · doi:10.48550/arxiv.1708.08286
openalex publication_date 2017/08/28 · openalex created_date 2022/10/05 · openalex updated_date 2026/07/28
Realistic simulations in engineering or in the materials sciences can consume\nenormous computing resources and thus require the use of massively parallel\nsupercomputers. The probability of a failure increases both with the runtime\nand with the number of system components. For future exascale systems it is\ntherefore considered critical that strategies are developed to make software\nresilient against failures. In this article, we present a scalable,\ndistributed, diskless, and resilient checkpointing scheme that can create and\nrecover snapshots of a partitioned simulation domain. We demonstrate the\nefficiency and scalability of the checkpoint strategy for simulations with up\nto 40 billion computational cells executing on more than 400 billion\nfloating point values. A checkpoint creation is shown to require only a few\nseconds and the new checkpointing scheme scales almost perfectly up to more\nthan 260 ,000 (218) processes. To recover from a diskless checkpoint\nduring runtime, we realize the recovery algorithms using ULFM MPI. The\ncheckpointing mechanism is fully integrated in a state-of-the-art\nhigh-performance multi-physics simulation framework. We demonstrate the\nefficiency and robustness of the method with a realistic phase-field simulation\noriginating in the material sciences and with a lattice Boltzmann method\nimplementation.\n