English

Blending Learning to Rank and Dense Representations for Efficient and Effective Cascades

Information Retrieval 2025-10-21 v1 Machine Learning Performance

Abstract

We investigate the exploitation of both lexical and neural relevance signals for ad-hoc passage retrieval. Our exploration involves a large-scale training dataset in which dense neural representations of MS-MARCO queries and passages are complemented and integrated with 253 hand-crafted lexical features extracted from the same corpus. Blending of the relevance signals from the two different groups of features is learned by a classical Learning-to-Rank (LTR) model based on a forest of decision trees. To evaluate our solution, we employ a pipelined architecture where a dense neural retriever serves as the first stage and performs a nearest-neighbor search over the neural representations of the documents. Our LTR model acts instead as the second stage that re-ranks the set of candidates retrieved by the first stage to enhance effectiveness. The results of reproducible experiments conducted with state-of-the-art dense retrievers on publicly available resources show that the proposed solution significantly enhances the end-to-end ranking performance while relatively minimally impacting efficiency. Specifically, we achieve a boost in nDCG@10 of up to 11% with an increase in average query latency of only 4.3%. This confirms the advantage of seamlessly combining two distinct families of signals that mutually contribute to retrieval effectiveness.

Keywords

Cite

@article{arxiv.2510.16393,
  title  = {Blending Learning to Rank and Dense Representations for Efficient and Effective Cascades},
  author = {Franco Maria Nardini and Raffaele Perego and Nicola Tonellotto and Salvatore Trani},
  journal= {arXiv preprint arXiv:2510.16393},
  year   = {2025}
}
R2 v1 2026-07-01T06:44:46.730Z