English

Minimally Invasive Randomization for Collecting Unbiased Preferences from Clickthrough Logs

Information Retrieval 2007-05-23 v1 Machine Learning

Abstract

Clickthrough data is a particularly inexpensive and plentiful resource to obtain implicit relevance feedback for improving and personalizing search engines. However, it is well known that the probability of a user clicking on a result is strongly biased toward documents presented higher in the result set irrespective of relevance. We introduce a simple method to modify the presentation of search results that provably gives relevance judgments that are unaffected by presentation bias under reasonable assumptions. We validate this property of the training data in interactive real world experiments. Finally, we show that using these unbiased relevance judgments learning methods can be guaranteed to converge to an ideal ranking given sufficient data.

Keywords

Cite

@article{arxiv.cs/0605037,
  title  = {Minimally Invasive Randomization for Collecting Unbiased Preferences from Clickthrough Logs},
  author = {Filip Radlinski and Thorsten Joachims},
  journal= {arXiv preprint arXiv:cs/0605037},
  year   = {2007}
}

Comments

7 pages. Proceedings of the 21st National Conference on Artificial Intelligence (AAAI), 2006

R2 v1 2026-07-22T12:25:39.453Z