Impact of Shallow vs. Deep Relevance Judgments on BERT-based Reranking Models
Abstract
This paper investigates the impact of shallow versus deep relevance judgments on the performance of BERT-based reranking models in neural Information Retrieval. Shallow-judged datasets, characterized by numerous queries each with few relevance judgments, and deep-judged datasets, involving fewer queries with extensive relevance judgments, are compared. The research assesses how these datasets affect the performance of BERT-based reranking models trained on them. The experiments are run on the MS MARCO and LongEval collections. Results indicate that shallow-judged datasets generally enhance generalization and effectiveness of reranking models due to a broader range of available contexts. The disadvantage of the deep-judged datasets might be mitigated by a larger number of negative training examples.
Cite
@article{arxiv.2506.23191,
title = {Impact of Shallow vs. Deep Relevance Judgments on BERT-based Reranking Models},
author = {Gabriel Iturra-Bocaz and Danny Vo and Petra Galuscakova},
journal= {arXiv preprint arXiv:2506.23191},
year = {2025}
}
Comments
Accepted at ICTIR'25