English

PersRM-R1: Enhance Personalized Reward Modeling with Reinforcement Learning

Machine Learning 2025-08-21 v1 Artificial Intelligence

Abstract

Reward models (RMs), which are central to existing post-training methods, aim to align LLM outputs with human values by providing feedback signals during fine-tuning. However, existing RMs struggle to capture nuanced, user-specific preferences, especially under limited data and across diverse domains. Thus, we introduce PersRM-R1, the first reasoning-based reward modeling framework specifically designed to identify and represent personal factors from only one or a few personal exemplars. To address challenges including limited data availability and the requirement for robust generalization, our approach combines synthetic data generation with a two-stage training pipeline consisting of supervised fine-tuning followed by reinforcement fine-tuning. Experimental results demonstrate that PersRM-R1 outperforms existing models of similar size and matches the performance of much larger models in both accuracy and generalizability, paving the way for more effective personalized LLMs.

Keywords

Cite

@article{arxiv.2508.14076,
  title  = {PersRM-R1: Enhance Personalized Reward Modeling with Reinforcement Learning},
  author = {Mengdi Li and Guanqiao Chen and Xufeng Zhao and Haochen Wen and Shu Yang and Di Wang},
  journal= {arXiv preprint arXiv:2508.14076},
  year   = {2025}
}
R2 v1 2026-07-01T04:57:16.673Z