English

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback

Information Retrieval 2026-08-01 v1

Abstract

In recommendation systems, users interact with only a small fraction of a vast item catalog, producing feedback that is both sparse and noisy. This challenges post-training generative recommenders: reward models trained from logged interactions often fail to generalize, while directly optimizing imperfect rewards can lead to reward over-optimization. We propose Exponential reward-weighted fine-tuning (Exp-RSFT), where each logged interaction is weighted by exp(r/λ)\exp(r/\lambda), avoids this failure by optimizing directly on the logged rewards, with the temperature λ\lambda regularizing against their noise. We theoretically show that Exp-RSFT's suboptimality decomposes into two costs: a coverage cost arising from limitations of the logging policy and a noise cost from imperfect feedback. The temperature λ\lambda balances these competing effects, yielding an optimal tradeoff between exploiting high-reward behavior and robustness to noise. Across three public benchmarks and a large-scale industrial dataset, we verify this theoretical prediction: performance follows an inverted-U trend as a function of λ\lambda, while PPO and DPO often over-optimize unreliable reward models and degrade recommendation quality. Exp-RSFT consistently improves ranking performance without requiring online exploration or preference data.

Keywords

Cite

@article{arxiv.2608.00816,
  title  = {Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback},
  author = {Keertana Chidambaram and Sanath Kumar Krishnamurthy and Qiuling Xu and Ko-Jen Hsiao and Moumita Bhattacharya},
  journal= {arXiv preprint arXiv:2608.00816},
  year   = {2026}
}