Distributed, Parallel, and Cluster Computing · Computer Science
Optimizing Checkpoint-Restart Mechanisms for HPC with DMTCP in Containers at NERSC
Madan Timalsina, Lisa Gerhardt, Nicholas Tyler, Johannes P. Blaschke +1
2024-07-30
Distributed, Parallel, and Cluster Computing · Computer Science
DMTCP: Transparent Checkpointing for Cluster Computations and the Desktop
Jason Ansel, Kapil Arya, Gene Cooperman
2009-02-24
Distributed, Parallel, and Cluster Computing · Computer Science
A Study of Checkpointing in Large Scale Training of Deep Neural Networks
Elvis Rojas, Albert Njoroge Kahira, Esteban Meneses, Leonardo Bautista Gomez +1
2021-03-30
Distributed, Parallel, and Cluster Computing · Computer Science
Checkpoint/restart approaches for a thread-based MPI runtime
Julien Adam, Maxime Kermarquer, Jean-Baptiste Besnard, Leonardo Bautista-Gomez +5
2019-06-13
Computational Physics · Physics
Use of checkpoint-restart for complex HEP software on traditional architectures and Intel MIC
Kapil Arya, Gene Cooperman, Andrea Dotti, Peter Elmer
2014-10-24
Operating Systems · Computer Science
Transparent Checkpoint-Restart over InfiniBand
Jiajun Cao, Gregory Kerr, Kapil Arya, Gene Cooperman
2014-02-03
Distributed, Parallel, and Cluster Computing · Computer Science
Performance Evaluation of an Algorithm-based Asynchronous Checkpoint-Restart Fault Tolerant Application Using Mixed MPI/GPI-2
Adrian Bazaga, Michal Pitonak
2018-05-08
Distributed, Parallel, and Cluster Computing · Computer Science
Checkpoint and Restart: An Energy Consumption Characterization in Clusters
Marina Moran, Javier Balladini, Dolores Rexachs, Emilio Luque
2024-09-05
Distributed, Parallel, and Cluster Computing · Computer Science
CRAFT: A library for easier application-level Checkpoint/Restart and Automatic Fault Tolerance
Faisal Shahzad, Jonas Thies, Moritz Kreutzer, Thomas Zeiser +2
2017-08-08
Operating Systems · Computer Science
Adapting the DMTCP Plugin Model for Checkpointing of Hardware Emulation
Rohan Garg, Kapil Arya, Jiajun Cao, Gene Cooperman +4
2017-03-03
Distributed, Parallel, and Cluster Computing · Computer Science
Extending the OpenCHK Model with Advanced Checkpoint Features
Marcos Maroñas, Sergi Mateo, Kai Keller, Leonardo Bautista-Gomez +2
2020-07-02
Distributed, Parallel, and Cluster Computing · Computer Science
System-level Scalable Checkpoint-Restart for Petascale Computing
Jiajun Cao, Kapil Arya, Rohan Garg, Shawn Matott +4
2016-09-27
Distributed, Parallel, and Cluster Computing · Computer Science
Dependency-Aware Rollback and Checkpoint-Restart for Distributed Task-Based Runtimes
Kiril Dichev, Herbert Jordan, Konstantinos Tovletoglou, Thomas Heller +3
2017-05-30
Distributed, Parallel, and Cluster Computing · Computer Science
Improving Grid Computing Performance by Optimally Reducing Checkpointing Effect
Garba Aliyu, Kana A. F. D., Abdullahi Mohammed, Idris Abdulmumin +2
2020-06-05
Distributed, Parallel, and Cluster Computing · Computer Science
A Utilization Model for Optimization of Checkpoint Intervals in Distributed Stream Processing Systems
Sachini Jayasekara, Aaron Harwood, Shanika Karunasekera
2020-04-21
Distributed, Parallel, and Cluster Computing · Computer Science
A cooperative partial snapshot algorithm for checkpoint-rollback recovery of large-scale and dynamic distributed systems and experimental evaluations
Junya Nakamura, Yonghwan Kim, Yoshiaki Katayama, Toshimitsu Masuzawa
2021-03-30
Distributed, Parallel, and Cluster Computing · Computer Science
Understanding LLM Checkpoint/Restore I/O Strategies and Patterns
Mikaila J. Gossman, Avinash Maurya, Bogdan Nicolae, Jon C. Calhoun
2026-01-01
Distributed, Parallel, and Cluster Computing · Computer Science
Checkpoint-Restart Libraries Must Become More Fault Tolerant
Anthony Skjellum, Derek Schafer
2021-12-22
Distributed, Parallel, and Cluster Computing · Computer Science
PartRePer-MPI: Combining Fault Tolerance and Performance for MPI Applications
Sarthak Joshi, Sathish Vadhiyar
2023-10-26
Distributed, Parallel, and Cluster Computing · Computer Science
Reinit++: Evaluating the Performance of Global-Restart Recovery Methods For MPI Fault Tolerance
Giorgis Georgakoudis, Luanzheng Guo, Ignacio Laguna
2021-02-16
Distributed, Parallel, and Cluster Computing · Computer Science
Distributed Order Recording Techniques for Efficient Record-and-Replay of Multi-threaded Programs
Xiang Fu, Shiman Meng, Weiping Zhang, Luanzheng Guo +5
2026-02-19
Distributed, Parallel, and Cluster Computing · Computer Science
ReStore: In-Memory REplicated STORagE for Rapid Recovery in Fault-Tolerant Algorithms
Lukas Hübner, Demian Hespe, Peter Sanders, Alexandros Stamatakis
2023-01-26