English

Claim-Level Rubric Rewards for Video Caption Reinforcement Learning

Computer Vision and Pattern Recognition 2026-07-06 v1

Abstract

In this paper, we introduce Claim-Level Rubric Rewards (CuRe), a structured reward framework designed to address the reward-design bottleneck in reinforcement learning for dense video captioning. Existing reward designs generally fall into two categories: holistic response-level judgment across heterogeneous criteria, or alignment-based evaluation against reference captions. However, both paradigms suffer from fundamental limitations. Holistic rewards struggle to ensure factual accuracy and are prone to stylistic reward hacking, while reference-based rewards overly rely on rigid textual alignment, failing to preserve the completeness and diversity inherent to open-ended generation tasks. To address these challenges, CuRe reformulates reward modeling as fine-grained claim-level verification. Specifically, CuRe decomposes captions into category-aware atomic claims through a structured rubric, converting holistic evaluation into simpler and more reliable claim-level verification.

Cite

@article{arxiv.2607.05150,
  title  = {Claim-Level Rubric Rewards for Video Caption Reinforcement Learning},
  author = {Mingqi Gao and Hongyuan Dong and Yifei Chen and Zhisheng Zhong and Zheng Ruan and Wenjin Hou and Yu Chen and Han Hu and Yansong Tang},
  journal= {arXiv preprint arXiv:2607.05150},
  year   = {2026}
}