English

GIFT: Reconciling Post-Training Objectives via Finite-Temperature Gibbs Initialization

Machine Learning 2026-03-19 v2 Artificial Intelligence Computation and Language

Abstract

The prevailing post-training paradigm for Large Reasoning Models (LRMs) - Supervised Fine-Tuning (SFT) followed by Reinforcement Learning (RL) - suffers from an intrinsic optimization mismatch: the rigid supervision inherent in SFT induces distributional collapse, thereby exhausting the exploration space necessary for subsequent RL. In this paper, we reformulate SFT to reconcile post-training objectives and propose Gibbs Initialization with Finite Temperature (GIFT). We characterize standard SFT as a degenerate zero-temperature limit that suppresses base priors. Conversely, GIFT incorporates supervision as a finite-temperature energy potential, establishing a distributional bridge that promotes objective consistency throughout the post-training pipeline. Our experiments demonstrate that GIFT significantly outperforms standard SFT and other competitive baselines when utilized for RL initialization, providing a mathematically principled pathway to preserve exploration and align the two post-training stages. Our code is available at https://github.com/zzy1127/GIFT.

Keywords

Cite

@article{arxiv.2601.09233,
  title  = {GIFT: Reconciling Post-Training Objectives via Finite-Temperature Gibbs Initialization},
  author = {Zhengyang Zhao and Lu Ma and Yizhen Jiang and Xiaochen Ma and Zimo Meng and Chengyu Shen and Lexiang Tang and Haoze Sun and Peng Pei and Wentao Zhang},
  journal= {arXiv preprint arXiv:2601.09233},
  year   = {2026}
}
R2 v1 2026-07-01T09:03:55.792Z