English

$\pi_\texttt{RL}$: Online RL Fine-tuning for Flow-based Vision-Language-Action Models

Machine Learning 2026-01-30 v3

Abstract

Vision-Language-Action (VLA) models enable robots to understand and perform complex tasks from multimodal input. Although recent work explores using reinforcement learning (RL) to automate the laborious data collection process in scaling supervised fine-tuning (SFT), applying RL to large-scale flow-based VLAs (\eg, π0\pi_0, π0.5\pi_{0.5}) remains challenging due to intractable action log-likelihoods raised from flow matching. We address this challenge with πRL\pi_{\texttt{RL}}, featuring two technical approaches: (1) \textbf{Flow-Noise} models the denoising process as a discrete-time MDP with a learnable noise network for exact log-likelihood computation. (2) \textbf{Flow-SDE} integrates denoising with agent-environment interaction, formulating a two-layer MDP that employs ODE-to-SDE conversion for efficient RL exploration. We evaluate πRL\pi_{\texttt{RL}} across various benchmarks, with experiments demonstrating that RL yields significant performance improvements in both in-distribution and out-of-distribution settings.

Keywords

Cite

@article{arxiv.2510.25889,
  title  = {$\pi_\texttt{RL}$: Online RL Fine-tuning for Flow-based Vision-Language-Action Models},
  author = {Kang Chen and Zhihao Liu and Tonghe Zhang and Zhen Guo and Si Xu and Hao Lin and Hongzhi Zang and Xiang Li and Quanlu Zhang and Zhaofei Yu and Guoliang Fan and Tiejun Huang and Yu Wang and Chao Yu},
  journal= {arXiv preprint arXiv:2510.25889},
  year   = {2026}
}
R2 v1 2026-07-01T07:12:42.255Z