English

How robust is MovieLens? A dataset analysis for recommender systems

Information Retrieval 2019-09-30 v1 Machine Learning

Abstract

Research publication requires public datasets. In recommender systems, some datasets are largely used to compare algorithms against a --supposedly-- common benchmark. Problem: for various reasons, these datasets are heavily preprocessed, making the comparison of results across papers difficult. This paper makes explicit the variety of preprocessing and evaluation protocols to test the robustness of a dataset (or lack of flexibility). While robustness is good to compare results across papers, for flexible datasets we propose a method to select a preprocessing protocol and share results more transparently.

Keywords

Cite

@article{arxiv.1909.12799,
  title  = {How robust is MovieLens? A dataset analysis for recommender systems},
  author = {Anne-Marie Tousch},
  journal= {arXiv preprint arXiv:1909.12799},
  year   = {2019}
}

Comments

2 pages ; accepted at REVEAL workshop, RecSys 2019

R2 v1 2026-06-23T11:28:24.639Z