English
Related papers

Related papers: CWRCzech: 100M Query-Document Czech Click Dataset …

200 papers

Web search engines focus on serving highly relevant results within hundreds of milliseconds. Pre-trained language transformer models such as BERT are therefore hard to use in this scenario due to their high computational demands. We present…

Information Retrieval · Computer Science 2021-12-06 Matěj Kocián , Jakub Náplava , Daniel Štancl , Vladimír Kadlec

Traditional machine-learned ranking systems for web search are often trained to capture stationary relevance of documents to queries, which has limited ability to track non-stationary user intention in a timely manner. In recency search,…

Information Retrieval · Computer Science 2011-03-22 Taesup Moon , Wei Chu , Lihong Li , Zhaohui Zheng , Yi Chang

Publication databases rely on accurate metadata extraction from diverse web sources, yet variations in web layouts and data formats present challenges for metadata providers. This paper introduces CRAWLDoc, a new method for contextual…

Computation and Language · Computer Science 2025-06-05 Fabian Karl , Ansgar Scherp

As stakeholders' pressure on corporates for disclosing their corporate social responsibility operations grows, it is crucial to understand how efficient corporate disclosure systems are in bridging the gap between corporate social…

General Economics · Economics 2023-01-10 Xhesilda Vogli , Erion Çano

Users of Web search engines reveal their information needs through queries and clicks, making click logs a useful asset for information retrieval. However, click logs have not been publicly released for academic use, because they can be too…

Information Retrieval · Computer Science 2020-08-19 Nick Craswell , Daniel Campos , Bhaskar Mitra , Emine Yilmaz , Bodo Billerbeck

Data plays a vital role in machine learning studies. In the research of recommendation, both user behaviors and side information are helpful to model users. So, large-scale real scenario datasets with abundant user behaviors will contribute…

Information Retrieval · Computer Science 2021-06-14 Bin Hao , Min Zhang , Weizhi Ma , Shaoyun Shi , Xinxing Yu , Houzhi Shan , Yiqun Liu , Shaoping Ma

This paper describes a novel dataset consisting of sentences with semantic similarity annotations. The data originate from the journalistic domain in the Czech language. We describe the process of collecting and annotating the data in…

Computation and Language · Computer Science 2022-01-24 Jakub Sido , Michal Seják , Ondřej Pražák , Miloslav Konopík , Václav Moravec

We present BenCzechMark (BCM), the first comprehensive Czech language benchmark designed for large language models, offering diverse tasks, multiple task formats, and multiple evaluation metrics. Its duel scoring system is grounded in…

In this paper, we describe the Czech Court Decision Corpus (CzCDC). CzCDC is a dataset of 237,723 decisions published by the Czech apex (or top-tier) courts, namely the Supreme Court, the Supreme Administrative Court and the Constitutional…

Computation and Language · Computer Science 2019-10-22 Tereza Novotná , Jakub Harašta

In this paper, we examine several methods of acquiring Czech data for automated fact-checking, which is a task commonly modeled as a classification of textual claim veracity w.r.t. a corpus of trusted ground truths. We attempt to collect…

Computation and Language · Computer Science 2023-12-19 Herbert Ullrich , Jan Drchal , Martin Rýpar , Hana Vincourová , Václav Moravec

Context information in search sessions has proven to be useful for capturing user search intent. Existing studies explored user behavior sequences in sessions in different ways to enhance query suggestion or document ranking. However, a…

Information Retrieval · Computer Science 2021-08-25 Yutao Zhu , Jian-Yun Nie , Zhicheng Dou , Zhengyi Ma , Xinyu Zhang , Pan Du , Xiaochen Zuo , Hao Jiang

Users' clicks on Web search results are one of the key signals for evaluating and improving web search quality and have been widely used as part of current state-of-the-art Learning-To-Rank(LTR) models. With a large volume of search logs…

Information Retrieval · Computer Science 2021-05-24 Jianghong Zhou , Sayyed M. Zahiri , Simon Hughes , Khalifeh Al Jadda , Surya Kallumadi , Eugene Agichtein

In this paper, we introduce a new Czech subjectivity dataset of 10k manually annotated subjective and objective sentences from movie reviews and descriptions. Our prime motivation is to provide a reliable dataset that can be used with the…

Computation and Language · Computer Science 2022-05-02 Pavel Přibáň , Josef Steinberger

Pre-trained models for Czech Natural Language Processing are often evaluated on purely linguistic tasks (POS tagging, parsing, NER) and relatively simple classification tasks such as sentiment classification or article classification from a…

Computation and Language · Computer Science 2023-07-21 Hynek Kydlíček , Jindřich Libovický

We present the Verifee Dataset: a novel dataset of news articles with fine-grained trustworthiness annotations. We develop a detailed methodology that assesses the texts based on their parameters encompassing editorial transparency,…

Computation and Language · Computer Science 2022-12-19 Matyáš Boháček , Michal Bravanský , Filip Trhlík , Václav Moravec

Click logs are valuable resources for a variety of information retrieval (IR) tasks. This includes query understanding/analysis, as well as learning effective IR models particularly when the models require large amounts of training data. We…

Information Retrieval · Computer Science 2021-04-29 Navid Rekabsaz , Oleg Lesota , Markus Schedl , Jon Brassey , Carsten Eickhoff

Click-through rate (CTR) is a key signal of relevance for search engine results, both organic and sponsored. CTR of a result has two core components: (a) the probability of examination of a result by a user, and (b) the perceived relevance…

Machine Learning · Computer Science 2018-10-22 Muhammad Asiful Islam , Ramakrishnan Srikant , Sugato Basu

An effective ranking model usually requires a large amount of training data to learn the relevance between documents and queries. User clicks are often used as training data since they can indicate relevance and are cheap to collect, but…

Information Retrieval · Computer Science 2023-02-21 Xiaojie Sun , Lulu Yu , Yiting Wang , Keping Bi , Jiafeng Guo

Topic localization aims to identify spans of text that express a given topic defined by a name and description. To study this task, we introduce a human-annotated benchmark based on Czech historical documents, containing human-defined…

Computation and Language · Computer Science 2026-03-05 Martin Kostelník , Michal Hradiš , Martin Dočekal

The ranking incentives of many authors of Web pages play an important role in the Web dynamics. That is, authors who opt to have their pages highly ranked for queries of interest, often respond to rankings for these queries by manipulating…

Information Retrieval · Computer Science 2020-06-29 Ziv Vasilisky , Moshe Tennenholtz , Oren Kurland
‹ Prev 1 2 3 10 Next ›