Training large language models faces frequent interruptions due to various faults, demanding robust fault-tolerance. Existing backup-free methods, such as redundant computation, dynamic parallelism, and data rerouting, each incur performance penalties, whether from ongoing overhead, lengthy reconfigurations, or post-recovery inefficiencies. We propose Chameleon, an adaptive fault-tolerant system that intelligently selects optimal recovery strategies when a failure occurs. Chameleon achieves this through a unified performance model, expedient execution plan search, accurate performance estimation, and efficient communication optimizations. Experiments on a 32-card cluster show that Chameleon maintains a performance gap of within 11.00% between post-recovery and failure-free training, while preserving model convergence and efficient memory usage. Compared to state-of-the-art methods, Chameleon achieves up to 1.229x and 1.355x higher average throughput than Oobleck and Recycle, respectively.
@article{arxiv.2508.21613,
title = {Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection},
author = {Yuhang Zhou and Zhibin Wang and Peng Jiang and Haoran Xia and Junhe Lu and Qianyu Jiang and Rong Gu and Hengxi Xu and Xinjing Huang and Guanghuan Fang and Zhiheng Hu and Jingyi Zhang and Yongjin Cai and Jian He and Chen Tian},
journal= {arXiv preprint arXiv:2508.21613},
year = {2026}
}