English

UniRL-Zero: Reinforcement Learning on Unified Models with Joint Language Model and Diffusion Model Experts

Machine Learning 2025-10-22 v1 Artificial Intelligence

Abstract

We present UniRL-Zero, a unified reinforcement learning (RL) framework that boosts, multimodal language model understanding and reasoning, diffusion model multimedia generation, and their beneficial interaction capabilities within a unified model. Our work defines six scenarios for unified model reinforcement learning, providing systematic baselines for reinforcement learning of unified understanding and generation model. Our code is available at https://github.com/G-U-N/UniRL.

Keywords

Cite

@article{arxiv.2510.17937,
  title  = {UniRL-Zero: Reinforcement Learning on Unified Models with Joint Language Model and Diffusion Model Experts},
  author = {Fu-Yun Wang and Han Zhang and Michael Gharbi and Hongsheng Li and Taesung Park},
  journal= {arXiv preprint arXiv:2510.17937},
  year   = {2025}
}
R2 v1 2026-07-22T20:52:23.769Z