On Checkpoint/Restart Viability for Exascale High Performance Computing Open Access
Lafrance, Matthew (Spring 2026)
Abstract
In 2004, foundational research evaluated the efficacy of checkpoint-based rollback-recovery for high-performance computing, predicting that extreme fault-tolerance overheads would render the paradigm unsustainable at scale. Since then, the supercomputing landscape has undergone a fundamental paradigm shift, with highly dense GPU-accelerated systems largely replacing traditional CPU-only architectures. Consequently, it is unknown whether traditional coordinated checkpointing remains a viable resilience strategy for modern machines. To address this gap, we re-evaluate the limits of rollback-recovery by applying a normalized k-means clustering algorithm to the Top500 dataset to derive four distinct modern architectural classes. Using a parallel discrete-event simulator, we scale these system classes to 1.0 ExaFLOP and beyond, analyzing their Useful Work Factor (U ) and Parallel Efficiency (E) under varying reliability parameters. We find that architectural compute density funda- mentally dictates system resilience. While traditional CPU-only systems succumb entirely to recovery overheads at exascale, modern GPU architectures successfully sustain operations by drastically minimizing the required node count. However, our extreme scaling experiments up to 1024×exascale demonstrate that this density only provides a temporary reprieve; as systems approach the zettascale boundary, the optimal checkpoint interval converges with the checkpoint latency itself, proving that rollback-recovery will ultimately break down.
Table of Contents
1 Introduction 1
2 Background 3
2.1 Principles of Checkpoint-Based Rollback-Recovery 3
2.2 Useful Work and Parallel Efficiency 4
2.3 Simulation Framework 6
2.4 Related Work 7
2.5 Summary of Simulation Parameters 8
3 Methodology 10
3.1 Data Acquisition and Preprocessing 10
3.2 System Classification 11
3.3 Derivation of Simulation Parameters 12
3.4 Simulation Execution and Experimental Parameters 13
4 Experiments 14
4.1 System Classification Results 14
4.2 Derived Exascale Parameters 14
4.3 Experimental Parameters 15
4.4 Baseline Exascale Viability 17
4.5 Sensitivity to Reliability and Overhead Parameters 20
4.6 Extreme Scale Sustainability 26
5 Conclusion and Future Work 29
5.1 Summary of Findings 29
5.2 Limitations 30
5.3 Future Work 31
A Appendix 32
Bibliography 33
About this Honors Thesis
| School | |
|---|---|
| Department | |
| Degree | |
| Submission | |
| Language |
|
| Research Field | |
| Keyword | |
| Committee Chair / Thesis Advisor | |
| Committee Members |
Primary PDF
| Thumbnail | Title | Date Uploaded | Actions |
|---|---|---|---|
|
|
On Checkpoint/Restart Viability for Exascale High Performance Computing () | 2026-04-22 13:52:22 -0400 |
|
Supplemental Files
| Thumbnail | Title | Date Uploaded | Actions |
|---|