English
Related papers

Related papers: ORCAS: 18 Million Clicked Query-Document Pairs for…

200 papers

Click logs are valuable resources for a variety of information retrieval (IR) tasks. This includes query understanding/analysis, as well as learning effective IR models particularly when the models require large amounts of training data. We…

Information Retrieval · Computer Science 2021-04-29 Navid Rekabsaz , Oleg Lesota , Markus Schedl , Jon Brassey , Carsten Eickhoff

The Deep Learning Track is a new track for TREC 2019, with the goal of studying ad hoc ranking in a large data regime. It is the first track with large human-labeled training sets, introducing two sets corresponding to two tasks, each with…

Information Retrieval · Computer Science 2020-03-19 Nick Craswell , Bhaskar Mitra , Emine Yilmaz , Daniel Campos , Ellen M. Voorhees

The TREC Deep Learning (DL) Track studies ad hoc search in the large data regime, meaning that a large set of human-labeled training data is available. Results so far indicate that the best models with large data may be deep neural…

Information Retrieval · Computer Science 2021-04-20 Nick Craswell , Bhaskar Mitra , Emine Yilmaz , Daniel Campos , Ellen M. Voorhees , Ian Soboroff

Users' clicks on Web search results are one of the key signals for evaluating and improving web search quality and have been widely used as part of current state-of-the-art Learning-To-Rank(LTR) models. With a large volume of search logs…

Information Retrieval · Computer Science 2021-05-24 Jianghong Zhou , Sayyed M. Zahiri , Simon Hughes , Khalifeh Al Jadda , Surya Kallumadi , Eugene Agichtein

This is the fourth year of the TREC Deep Learning track. As in previous years, we leverage the MS MARCO datasets that made hundreds of thousands of human annotated training labels available for both passage and document ranking tasks. In…

Information Retrieval · Computer Science 2025-07-16 Nick Craswell , Bhaskar Mitra , Emine Yilmaz , Daniel Campos , Jimmy Lin , Ellen M. Voorhees , Ian Soboroff

Search engine companies collect the "database of intentions", the histories of their users' search queries. These search logs are a gold mine for researchers. Search engine companies, however, are wary of publishing search logs in order not…

Databases · Computer Science 2011-05-13 Michaela Goetz , Ashwin Machanavajjhala , Guozhang Wang , Xiaokui Xiao , Johannes Gehrke

We present CWRCzech, Click Web Ranking dataset for Czech, a 100M query-document Czech click dataset for relevance ranking with user behavior data collected from search engine logs of Seznam$.$cz. To the best of our knowledge, CWRCzech is…

Information Retrieval · Computer Science 2024-07-16 Josef Vonášek , Milan Straka , Rostislav Krč , Lenka Lasoňová , Ekaterina Egorova , Jana Straková , Jakub Náplava

We study the problem of deep recall model in industrial web search, which is, given a user query, retrieve hundreds of most relevance documents from billions of candidates. The common framework is to train two encoding models based on…

Information Retrieval · Computer Science 2020-07-06 Yusi Zhang , Chuanjie Liu , Angen Luo , Hui Xue , Xuan Shan , Yuxiang Luo , Yiqian Xia , Yuanchi Yan , Haidong Wang

We present DepCC, the largest-to-date linguistically analyzed corpus in English including 365 million documents, composed of 252 billion tokens and 7.5 billion of named entity occurrences in 14.3 billion sentences from a web-scale crawl of…

Computation and Language · Computer Science 2018-03-01 Alexander Panchenko , Eugen Ruppert , Stefano Faralli , Simone Paolo Ponzetto , Chris Biemann

This is the third year of the TREC Deep Learning track. As in previous years, we leverage the MS MARCO datasets that made hundreds of thousands of human annotated training labels available for both passage and document ranking tasks. In…

Information Retrieval · Computer Science 2025-07-14 Nick Craswell , Bhaskar Mitra , Emine Yilmaz , Daniel Campos , Jimmy Lin

The number of RDF knowledge graphs available on the Web grows constantly. Gathering these graphs at large scale for downstream applications hence requires the use of crawlers. Although Data Web crawlers exist, and general Web crawlers could…

The centralized collection of search interaction logs for training ranking models raises significant privacy concerns. Federated Online Learning to Rank (FOLTR) offers a privacy-preserving alternative by enabling collaborative model…

Information Retrieval · Computer Science 2025-08-19 Marcel Gregoriadis , Jingwei Kang , Johan Pouwelse

Click models are an important tool for leveraging user feedback, and are used by commercial search engines for surfacing relevant search results. However, existing click models are lacking in two aspects. First, they do not share…

Information Retrieval · Computer Science 2014-01-03 Dinesh Govindaraj , Tao Wang , S. V. N. Vishwanathan

Users often fail to formulate their complex information needs in a single query. As a consequence, they may need to scan multiple result pages or reformulate their queries, which may be a frustrating experience. Alternatively, systems can…

Computation and Language · Computer Science 2019-07-16 Mohammad Aliannejadi , Hamed Zamani , Fabio Crestani , W. Bruce Croft

Despite its troubled past, the AOL Query Log continues to be an important resource to the research community -- particularly for tasks like search personalisation. When using the query log these ranking experiments, little attention is…

Information Retrieval · Computer Science 2022-01-24 Sean MacAvaney , Craig Macdonald , Iadh Ounis

This is the second year of the TREC Deep Learning Track, with the goal of studying ad hoc ranking in the large training data regime. We again have a document retrieval task and a passage retrieval task, each with hundreds of thousands of…

Information Retrieval · Computer Science 2021-02-16 Nick Craswell , Bhaskar Mitra , Emine Yilmaz , Daniel Campos

In the era of big data, we continuously - and at times unknowingly - leave behind digital traces, by browsing, sharing, posting, liking, searching, watching, and listening to online content. When aggregated, these digital traces can provide…

Information Retrieval · Computer Science 2021-02-23 David Graus

Modeling contextual information in a search session has drawn more and more attention when understanding complex user intents. Recent methods are all data-driven, i.e., they train different models on large-scale search log data to identify…

Information Retrieval · Computer Science 2024-07-08 Haonan Chen , Zhicheng Dou , Yutao Zhu , Ji-Rong Wen

At the very beginning of compiling a bibliography, usually only basic information, such as title, authors and publication date of an item are known. In order to gather additional information about a specific item, one typically has to…

Digital Libraries · Computer Science 2012-12-18 Johann Schaible , Philipp Mayr

Evaluating retrieval performance without editorial relevance judgments is challenging, but instead, user interactions can be used as relevance signals. Living labs offer a way for small-scale platforms to validate information retrieval…

Information Retrieval · Computer Science 2023-10-12 Timo Breuer , Norbert Fuhr , Philipp Schaer
‹ Prev 1 2 3 10 Next ›