Neural IR architectures, particularly cross-encoders, are highly effective models whose internal mechanisms are mostly unknown. Most works trying to explain their behavior focused on high-level processes (e.g., what in the input influences the prediction, does the model adhere to known IR axioms) but fall short of describing the matching process. Instead of Mechanistic Interpretability approaches which specifically aim at explaining the hidden mechanisms of neural models, we demonstrate that more straightforward methods can already provide valuable insights. In this paper, we first focus on the attention process and extract causal insights highlighting the crucial roles of some attention heads in this process. Second, we provide an interpretation of the mechanism underlying matching detection.
@article{arxiv.2507.14604,
title = {Understanding Matching Mechanisms in Cross-Encoders},
author = {Mathias Vast and Basile Van Cooten and Laure Soulier and Benjamin Piwowarski},
journal= {arXiv preprint arXiv:2507.14604},
year = {2025}
}
Comments
Accepted at Workshop on Explainability in Information Retrieval at SIGIR 25 (WExIR25)