中文
相关论文

相关论文: CWRCzech: 100M Query-Document Czech Click Dataset …

200 篇论文

Web search engines focus on serving highly relevant results within hundreds of milliseconds. Pre-trained language transformer models such as BERT are therefore hard to use in this scenario due to their high computational demands. We present…

信息检索 · 计算机科学 2021-12-06 Matěj Kocián , Jakub Náplava , Daniel Štancl , Vladimír Kadlec

Traditional machine-learned ranking systems for web search are often trained to capture stationary relevance of documents to queries, which has limited ability to track non-stationary user intention in a timely manner. In recency search,…

信息检索 · 计算机科学 2011-03-22 Taesup Moon , Wei Chu , Lihong Li , Zhaohui Zheng , Yi Chang

Publication databases rely on accurate metadata extraction from diverse web sources, yet variations in web layouts and data formats present challenges for metadata providers. This paper introduces CRAWLDoc, a new method for contextual…

计算与语言 · 计算机科学 2025-06-05 Fabian Karl , Ansgar Scherp

As stakeholders' pressure on corporates for disclosing their corporate social responsibility operations grows, it is crucial to understand how efficient corporate disclosure systems are in bridging the gap between corporate social…

综合经济学 · 经济学 2023-01-10 Xhesilda Vogli , Erion Çano

Users of Web search engines reveal their information needs through queries and clicks, making click logs a useful asset for information retrieval. However, click logs have not been publicly released for academic use, because they can be too…

信息检索 · 计算机科学 2020-08-19 Nick Craswell , Daniel Campos , Bhaskar Mitra , Emine Yilmaz , Bodo Billerbeck

Data plays a vital role in machine learning studies. In the research of recommendation, both user behaviors and side information are helpful to model users. So, large-scale real scenario datasets with abundant user behaviors will contribute…

信息检索 · 计算机科学 2021-06-14 Bin Hao , Min Zhang , Weizhi Ma , Shaoyun Shi , Xinxing Yu , Houzhi Shan , Yiqun Liu , Shaoping Ma

This paper describes a novel dataset consisting of sentences with semantic similarity annotations. The data originate from the journalistic domain in the Czech language. We describe the process of collecting and annotating the data in…

计算与语言 · 计算机科学 2022-01-24 Jakub Sido , Michal Seják , Ondřej Pražák , Miloslav Konopík , Václav Moravec

We present BenCzechMark (BCM), the first comprehensive Czech language benchmark designed for large language models, offering diverse tasks, multiple task formats, and multiple evaluation metrics. Its duel scoring system is grounded in…

In this paper, we describe the Czech Court Decision Corpus (CzCDC). CzCDC is a dataset of 237,723 decisions published by the Czech apex (or top-tier) courts, namely the Supreme Court, the Supreme Administrative Court and the Constitutional…

计算与语言 · 计算机科学 2019-10-22 Tereza Novotná , Jakub Harašta

In this paper, we examine several methods of acquiring Czech data for automated fact-checking, which is a task commonly modeled as a classification of textual claim veracity w.r.t. a corpus of trusted ground truths. We attempt to collect…

计算与语言 · 计算机科学 2023-12-19 Herbert Ullrich , Jan Drchal , Martin Rýpar , Hana Vincourová , Václav Moravec

Context information in search sessions has proven to be useful for capturing user search intent. Existing studies explored user behavior sequences in sessions in different ways to enhance query suggestion or document ranking. However, a…

信息检索 · 计算机科学 2021-08-25 Yutao Zhu , Jian-Yun Nie , Zhicheng Dou , Zhengyi Ma , Xinyu Zhang , Pan Du , Xiaochen Zuo , Hao Jiang

Users' clicks on Web search results are one of the key signals for evaluating and improving web search quality and have been widely used as part of current state-of-the-art Learning-To-Rank(LTR) models. With a large volume of search logs…

信息检索 · 计算机科学 2021-05-24 Jianghong Zhou , Sayyed M. Zahiri , Simon Hughes , Khalifeh Al Jadda , Surya Kallumadi , Eugene Agichtein

In this paper, we introduce a new Czech subjectivity dataset of 10k manually annotated subjective and objective sentences from movie reviews and descriptions. Our prime motivation is to provide a reliable dataset that can be used with the…

计算与语言 · 计算机科学 2022-05-02 Pavel Přibáň , Josef Steinberger

Pre-trained models for Czech Natural Language Processing are often evaluated on purely linguistic tasks (POS tagging, parsing, NER) and relatively simple classification tasks such as sentiment classification or article classification from a…

计算与语言 · 计算机科学 2023-07-21 Hynek Kydlíček , Jindřich Libovický

We present the Verifee Dataset: a novel dataset of news articles with fine-grained trustworthiness annotations. We develop a detailed methodology that assesses the texts based on their parameters encompassing editorial transparency,…

计算与语言 · 计算机科学 2022-12-19 Matyáš Boháček , Michal Bravanský , Filip Trhlík , Václav Moravec

Click logs are valuable resources for a variety of information retrieval (IR) tasks. This includes query understanding/analysis, as well as learning effective IR models particularly when the models require large amounts of training data. We…

信息检索 · 计算机科学 2021-04-29 Navid Rekabsaz , Oleg Lesota , Markus Schedl , Jon Brassey , Carsten Eickhoff

Click-through rate (CTR) is a key signal of relevance for search engine results, both organic and sponsored. CTR of a result has two core components: (a) the probability of examination of a result by a user, and (b) the perceived relevance…

机器学习 · 计算机科学 2018-10-22 Muhammad Asiful Islam , Ramakrishnan Srikant , Sugato Basu

An effective ranking model usually requires a large amount of training data to learn the relevance between documents and queries. User clicks are often used as training data since they can indicate relevance and are cheap to collect, but…

信息检索 · 计算机科学 2023-02-21 Xiaojie Sun , Lulu Yu , Yiting Wang , Keping Bi , Jiafeng Guo

Topic localization aims to identify spans of text that express a given topic defined by a name and description. To study this task, we introduce a human-annotated benchmark based on Czech historical documents, containing human-defined…

计算与语言 · 计算机科学 2026-03-05 Martin Kostelník , Michal Hradiš , Martin Dočekal

The ranking incentives of many authors of Web pages play an important role in the Web dynamics. That is, authors who opt to have their pages highly ranked for queries of interest, often respond to rankings for these queries by manipulating…

信息检索 · 计算机科学 2020-06-29 Ziv Vasilisky , Moshe Tennenholtz , Oren Kurland
‹ 上一页 1 2 3 10 下一页 ›