English

Sparsifying Sparse Representations for Passage Retrieval by Top-$k$ Masking

Information Retrieval 2021-12-20 v1 Computation and Language

Abstract

Sparse lexical representation learning has demonstrated much progress in improving passage retrieval effectiveness in recent models such as DeepImpact, uniCOIL, and SPLADE. This paper describes a straightforward yet effective approach for sparsifying lexical representations for passage retrieval, building on SPLADE by introducing a top-kk masking scheme to control sparsity and a self-learning method to coax masked representations to mimic unmasked representations. A basic implementation of our model is competitive with more sophisticated approaches and achieves a good balance between effectiveness and efficiency. The simplicity of our methods opens the door for future explorations in lexical representation learning for passage retrieval.

Keywords

Cite

@article{arxiv.2112.09628,
  title  = {Sparsifying Sparse Representations for Passage Retrieval by Top-$k$ Masking},
  author = {Jheng-Hong Yang and Xueguang Ma and Jimmy Lin},
  journal= {arXiv preprint arXiv:2112.09628},
  year   = {2021}
}

Comments

8 pages, 1 figure

R2 v1 2026-06-24T08:22:17.147Z