English

Bridging Offline and Online Reinforcement Learning for LLMs

Computation and Language 2025-06-27 v1

Abstract

We investigate the effectiveness of reinforcement learning methods for finetuning large language models when transitioning from offline to semi-online to fully online regimes for both verifiable and non-verifiable tasks. Our experiments cover training on verifiable math as well as non-verifiable instruction following with a set of benchmark evaluations for both. Across these settings, we extensively compare online and semi-online Direct Preference Optimization and Group Reward Policy Optimization objectives, and surprisingly find similar performance and convergence between these variants, which all strongly outperform offline methods. We provide a detailed analysis of the training dynamics and hyperparameter selection strategies to achieve optimal results. Finally, we show that multi-tasking with verifiable and non-verifiable rewards jointly yields improved performance across both task types.

Keywords

Cite

@article{arxiv.2506.21495,
  title  = {Bridging Offline and Online Reinforcement Learning for LLMs},
  author = {Jack Lanchantin and Angelica Chen and Janice Lan and Xian Li and Swarnadeep Saha and Tianlu Wang and Jing Xu and Ping Yu and Weizhe Yuan and Jason E Weston and Sainbayar Sukhbaatar and Ilia Kulikov},
  journal= {arXiv preprint arXiv:2506.21495},
  year   = {2025}
}
R2 v1 2026-07-01T03:34:55.481Z