English

Siamese BERT-based Model for Web Search Relevance Ranking Evaluated on a New Czech Dataset

Information Retrieval 2021-12-06 v1 Computation and Language

Abstract

Web search engines focus on serving highly relevant results within hundreds of milliseconds. Pre-trained language transformer models such as BERT are therefore hard to use in this scenario due to their high computational demands. We present our real-time approach to the document ranking problem leveraging a BERT-based siamese architecture. The model is already deployed in a commercial search engine and it improves production performance by more than 3%. For further research and evaluation, we release DaReCzech, a unique data set of 1.6 million Czech user query-document pairs with manually assigned relevance levels. We also release Small-E-Czech, an Electra-small language model pre-trained on a large Czech corpus. We believe this data will support endeavours both of search relevance and multilingual-focused research communities.

Keywords

Cite

@article{arxiv.2112.01810,
  title  = {Siamese BERT-based Model for Web Search Relevance Ranking Evaluated on a New Czech Dataset},
  author = {Matěj Kocián and Jakub Náplava and Daniel Štancl and Vladimír Kadlec},
  journal= {arXiv preprint arXiv:2112.01810},
  year   = {2021}
}

Comments

Accepted at the Thirty-Fourth Annual Conference on Innovative Applications of Artificial Intelligence (IAAI-22). IAAI Innovative Application Award. 9 pages, 3 figures, 8 tables

R2 v1 2026-06-24T08:02:55.856Z