English

Self-Rewarding Rubric-Based Reinforcement Learning for Open-Ended Reasoning

Computation and Language 2025-10-01 v1 Artificial Intelligence Machine Learning

Abstract

Open-ended evaluation is essential for deploying large language models in real-world settings. In studying HealthBench, we observe that using the model itself as a grader and generating rubric-based reward signals substantially improves reasoning performance. Remarkably, the trained model also becomes a stronger grader. Motivated by this, we introduce Self-Rewarding Rubric-Based Reinforcement Learning for Open-Ended Reasoning, a lightweight framework that enables faster and more resource-efficient training while surpassing baselines. Remarkably, on Qwen3-32B, training with just the 4000-sample HealthBench Easy subset is sufficient to obtain a model that exceeds GPT-5 on HealthBench Hard. Incorporating a small amount of teacher-graded data further enhances performance for less capable models.

Keywords

Cite

@article{arxiv.2509.25534,
  title  = {Self-Rewarding Rubric-Based Reinforcement Learning for Open-Ended Reasoning},
  author = {Zhiling Ye and Yun Yue and Haowen Wang and Xudong Han and Jiadi Jiang and Cheng Wei and Lei Fan and Jiaxin Liang and Shuowen Zhang and Ji Li and Chunxiao Guo and Jian Wang and Peng Wei and Jinjie Gu},
  journal= {arXiv preprint arXiv:2509.25534},
  year   = {2025}
}
R2 v1 2026-07-01T06:06:20.380Z