English

Impact of Shallow vs. Deep Relevance Judgments on BERT-based Reranking Models

Information Retrieval 2025-07-01 v1

Abstract

This paper investigates the impact of shallow versus deep relevance judgments on the performance of BERT-based reranking models in neural Information Retrieval. Shallow-judged datasets, characterized by numerous queries each with few relevance judgments, and deep-judged datasets, involving fewer queries with extensive relevance judgments, are compared. The research assesses how these datasets affect the performance of BERT-based reranking models trained on them. The experiments are run on the MS MARCO and LongEval collections. Results indicate that shallow-judged datasets generally enhance generalization and effectiveness of reranking models due to a broader range of available contexts. The disadvantage of the deep-judged datasets might be mitigated by a larger number of negative training examples.

Keywords

Cite

@article{arxiv.2506.23191,
  title  = {Impact of Shallow vs. Deep Relevance Judgments on BERT-based Reranking Models},
  author = {Gabriel Iturra-Bocaz and Danny Vo and Petra Galuscakova},
  journal= {arXiv preprint arXiv:2506.23191},
  year   = {2025}
}

Comments

Accepted at ICTIR'25