中文
相关论文

相关论文: Variations in Relevance Judgments and the Shelf Li…

200 篇论文

Information retrieval systems have been evaluated using the Cranfield paradigm for many years. This paradigm allows a systematic, fair, and reproducible evaluation of different retrieval methods in fixed experimental environments. However,…

信息检索 · 计算机科学 2024-07-02 Jüri Keller , Timo Breuer , Philipp Schaer

Evaluating recommender systems remains a long-standing challenge, as offline methods based on historical user interactions and train-test splits often yield unstable and inconsistent results due to exposure bias, popularity bias, sampled…

Relevance judgment of human assessors is inherently subjective and dynamic when evaluation datasets are created for Information Retrieval (IR) systems. However, a small group of experts' relevance judgment results are usually taken as…

信息检索 · 计算机科学 2022-08-09 Dengya Zhu , Shastri L Nimmagadda , Kok Wai Wong , Torsten Reiners

Relevance judgments are crucial for evaluating information retrieval systems, but traditional human-annotated labels are time-consuming and expensive. As a result, many researchers turn to automatic alternatives to accelerate method…

信息检索 · 计算机科学 2025-07-15 Naghmeh Farzi , Laura Dietz

Large-scale test collections play a crucial role in Information Retrieval (IR) research. However, according to the Cranfield paradigm and the research into publicly available datasets, the existing information retrieval research studies are…

信息检索 · 计算机科学 2025-01-28 Hossein A. Rahmani , Xi Wang , Emine Yilmaz , Nick Craswell , Bhaskar Mitra , Paul Thomas

Cranfield-style retrieval evaluations with too few or too many relevant documents or with low inter-assessor agreement on relevance can reduce the reliability of observations. In evaluations with human assessors, information needs are often…

信息检索 · 计算机科学 2026-04-30 Jüri Keller , Maik Fröbe , Björn Engelmann , Fabian Haak , Timo Breuer , Birger Larsen , Philipp Schaer

Evaluation of search engines relies on assessments of search results for selected test queries, from which we would ideally like to draw conclusions in terms of relevance of the results for general (e.g., future, unknown) users. In practice…

信息检索 · 计算机科学 2015-11-24 Thomas Demeester , Robin Aly , Djoerd Hiemstra , Dong Nguyen , Chris Develder

Making the relevance judgments for a TREC-style test collection can be complex and expensive. A typical TREC track usually involves a team of six contractors working for 2-4 weeks. Those contractors need to be trained and monitored.…

信息检索 · 计算机科学 2025-03-27 Ian Soboroff

Using large language models (LLMs) to annotate relevance is an increasingly important technique in the information retrieval community. While some studies demonstrate that LLMs can achieve high user agreement with ground truth (human)…

信息检索 · 计算机科学 2026-01-15 Watheq Mansour , J. Shane Culpepper , Joel Mackenzie , Andrew Yates

Robust test collections are crucial for Information Retrieval research. Recently there is a growing interest in evaluating retrieval systems for domain-specific retrieval tasks, however these tasks often lack a reliable test collection with…

信息检索 · 计算机科学 2022-08-16 Sophia Althammer , Sebastian Hofstätter , Suzan Verberne , Allan Hanbury

Offline evaluation of search systems depends on test collections. These benchmarks provide the researchers with a corpus of documents, topics and relevance judgements indicating which documents are relevant for each topic. While test…

信息检索 · 计算机科学 2025-07-23 David Otero , Javier Parapar , Álvaro Barreiro

Large Language Models (LLMs) have been used as relevance assessors for Information Retrieval (IR) evaluation collection creation due to reduced cost and increased scalability as compared to human assessors. While previous research has…

信息检索 · 计算机科学 2026-01-06 Samaneh Mohtadi , Gianluca Demartini

The creation of relevance assessments by human assessors (often nowadays crowdworkers) is a vital step when building IR test collections. Prior works have investigated assessor quality & behaviour, though into the impact of a document's…

信息检索 · 计算机科学 2023-04-24 Nirmal Roy , Agathe Balayn , David Maxwell , Claudia Hauff

Evaluation is crucial in Information Retrieval. The development of models, tools and methods has significantly benefited from the availability of reusable test collections formed through a standardized and thoroughly tested methodology,…

信息检索 · 计算机科学 2017-09-07 Dan Li , Evangelos Kanoulas

While test collections provide the cornerstone for Cranfield-based evaluation of information retrieval (IR) systems, it has become practically infeasible to rely on traditional pooling techniques to construct test collections at the scale…

信息检索 · 计算机科学 2017-09-20 Mucahid Kutlu , Tamer Elsayed , Matthew Lease

Human relevance assessment is time-consuming and cognitively intensive, limiting the scalability of Information Retrieval evaluation. This has led to growing interest in using large language models (LLMs) as proxies for human judges.…

信息检索 · 计算机科学 2026-04-28 Chuting Yu , Hang Li , Guido Zuccon , Joel Mackenzie , Teerapong Leelanupab

This paper investigates the impact of shallow versus deep relevance judgments on the performance of BERT-based reranking models in neural Information Retrieval. Shallow-judged datasets, characterized by numerous queries each with few…

信息检索 · 计算机科学 2025-07-01 Gabriel Iturra-Bocaz , Danny Vo , Petra Galuscakova

The Cranfield paradigm has served as a foundational approach for developing test collections, with relevance judgments typically conducted by human assessors. However, the emergence of large language models (LLMs) has introduced new…

信息检索 · 计算机科学 2024-06-12 Gabriel de Jesus , Sérgio Nunes

In any ranking system, the retrieval model outputs a single score for a document based on its belief on how relevant it is to a given search query. While retrieval models have continued to improve with the introduction of increasingly…

信息检索 · 计算机科学 2021-05-12 Daniel Cohen , Bhaskar Mitra , Oleg Lesota , Navid Rekabsaz , Carsten Eickhoff

Neural retrieval models are generally regarded as fundamentally different from the retrieval techniques used in the late 1990's when the TREC ad hoc test collections were constructed. They thus provide the opportunity to empirically test…

信息检索 · 计算机科学 2022-01-27 Ellen M. Voorhees , Ian Soboroff , Jimmy Lin
‹ 上一页 1 2 3 10 下一页 ›