English

On the Convergent Validity of Offline Evaluation Designs for Recommender Systems

Information Retrieval 2026-07-27 v1

Abstract

Offline evaluation on historical interaction logs is the most common evaluation methodology for recommender systems. However, such evaluations depend on sparse, incomplete, or biased data, which raises concerns about whether commonly used evaluation setups reliably reflect true user preferences. In this work, we study how offline evaluation design choices affect the validity of recommender system comparisons. We evaluate a set of recommendation models across several evaluation setups that vary key factors such as data filtering thresholds and candidate set construction. To assess the validity of these configurations, we measure the correlation between model rankings obtained from conventional train-test splits on sparse interaction data and rankings from evaluations based on dense ground-truth user feedback. We use this agreement as an indication of their validity with respect to true user preferences. Our results show that the validity of sparse evaluation depends on the dataset and the specific dense evaluation targets, and that there is no uniformly best offline evaluation design.

Keywords

Cite

@article{arxiv.2607.25097,
  title  = {On the Convergent Validity of Offline Evaluation Designs for Recommender Systems},
  author = {Sushobhan Parajuli and Samira Vaez Barenji and Michael D. Ekstrand},
  journal= {arXiv preprint arXiv:2607.25097},
  year   = {2026}
}

Comments

9 pages, Accepted at ACM RecSys 2026 Main Track