English

CapTrack: Multifaceted Evaluation of Forgetting in LLM Post-Training

Machine Learning 2026-05-25 v2

Abstract

Large language model (LLM) post-training enhances latent skills, unlocks value alignment, improves performance, and enables domain adaptation. Unfortunately, post-training is known to induce forgetting, especially in the ubiquitous use-case of leveraging third-party pre-trained models, which is typically understood as a loss of parametric or factual knowledge. We argue that this accuracy-centric view is insufficient for modern foundation models and instead define forgetting as systematic model drift that degrades behavior and user experience. In this context, we introduce CapTrack, a capability-centric framework for analyzing forgetting in LLMs that combines a behavioral taxonomy with an evaluation suite centered on capability-specific metrics. Using CapTrack, we conduct a large-scale empirical study across post-training algorithms, domains, and model families, including models up to 80B parameters. We find that forgetting extends beyond parametric knowledge, with pronounced drift in robustness and default behaviors. Instruction fine-tuning induces the strongest relative drift, while preference optimization is more conservative and can partially recover lost capabilities. Differences across model families persist, and no universal mitigation emerges.

Keywords

Cite

@article{arxiv.2603.06610,
  title  = {CapTrack: Multifaceted Evaluation of Forgetting in LLM Post-Training},
  author = {Lukas Thede and Stefan Winzeck and Zeynep Akata and Jonathan Richard Schwarz},
  journal= {arXiv preprint arXiv:2603.06610},
  year   = {2026}
}
R2 v1 2026-07-01T11:07:32.287Z