English

The 2021 RecSys Challenge Dataset: Fairness is not optional

Social and Information Networks 2021-09-23 v3

Abstract

After the success the RecSys 2020 Challenge, we are describing a novel and bigger dataset that was released in conjunction with the ACM RecSys Challenge 2021. This year's dataset is not only bigger (~ 1B data points, a 5 fold increase), but for the first time it take into consideration fairness aspects of the challenge. Unlike many static datsets, a lot of effort went into making sure that the dataset was synced with the Twitter platform: if a user deleted their content, the same content would be promptly removed from the dataset too. In this paper, we introduce the dataset and challenge, highlighting some of the issues that arise when creating recommender systems at Twitter scale.

Keywords

Cite

@article{arxiv.2109.08245,
  title  = {The 2021 RecSys Challenge Dataset: Fairness is not optional},
  author = {Luca Belli and Alykhan Tejani and Frank Portman and Alexandre Lung-Yut-Fong and Ben Chamberlain and Yuanpu Xie and Kristian Lum and Jonathan Hunt and Michael Bronstein and Vito Walter Anelli and Saikishore Kalloori and Bruce Ferwerda and Wenzhe Shi},
  journal= {arXiv preprint arXiv:2109.08245},
  year   = {2021}
}
R2 v1 2026-06-24T06:03:19.843Z