中文
相关论文

相关论文: SARA: A Collection of Sensitivity-Aware Relevance …

200 篇论文

The amount of data for processing and categorization grows at an ever increasing rate. At the same time the demand for collaboration and transparency in organizations, government and businesses, drives the release of data from internal…

机器学习 · 计算机科学 2020-08-26 Jan Neerbek

To create a new IR test collection at low cost, it is valuable to carefully select which documents merit human relevance judgments. Shared task campaigns such as NIST TREC pool document rankings from many participating systems (and often…

信息检索 · 计算机科学 2020-08-06 Md Mustafizur Rahman , Mucahid Kutlu , Tamer Elsayed , Matthew Lease

Sentiment analysis (SA) is commonly applied to digital textual data, revealing insight into opinions and feelings. Many systematic reviews have summarized existing work, but often overlook discussions of validity and scientific practices.…

计算机与社会 · 计算机科学 2023-10-31 Laura E. M. Ayravainen , Joanne Hinds , Brittany I. Davidson

Considering the limited internal parametric knowledge, retrieval-augmented generation (RAG) has been widely used to extend the knowledge scope of large language models (LLMs). Despite the extensive efforts on RAG research, in existing…

计算与语言 · 计算机科学 2024-11-22 Yuhao Wang , Ruiyang Ren , Junyi Li , Wayne Xin Zhao , Jing Liu , Ji-Rong Wen

Users increasingly expect modern search systems to offer a unified interface that seamlessly retrieves information from diverse data sources and formats. However, current information retrieval (IR) evaluation benchmarks have not kept pace…

信息检索 · 计算机科学 2026-05-13 Mehmet Deniz Türkmen , Suchana Datta , Dwaipayan Roy , Daniel Hienert , Philipp Mayr , Derek Greene

At least ninety countries implement Freedom of Information laws that state that government documents must be made freely available, or opened, to the public. However, many government documents contain sensitive information, such as personal…

计算机与社会 · 计算机科学 2019-07-08 Graham McDonald , Craig Macdonald , Iadh Ounis

There is an extensive body of research on Social Network Analysis (SNA) based on the email archive. The network used in the analysis is generally extracted either by capturing the email communication in From, To, Cc and Bcc email header…

社会与信息网络 · 计算机科学 2017-09-04 Gregory Tsipenyuk , Jon Crowcroft

In this position paper we argue that certain aspects of relevance assessment in the evaluation of IR systems are oversimplified and that human assessments represented by qrels should be augmented to take account of contextual factors and…

信息检索 · 计算机科学 2015-01-27 Laura Hasler , Martin Halvey , Robert Villa

To evaluate Information Retrieval (IR) effectiveness, a possible approach is to use test collections, which are composed of a collection of documents, a set of description of information needs (called topics), and a set of relevant…

信息检索 · 计算机科学 2020-11-03 Kevin Roitero

In recent years, with the rapid advancements in large language models (LLMs), achieving excellent empathetic response capability has become a crucial prerequisite. Consequently, managing and understanding large-scale video datasets has…

计算与语言 · 计算机科学 2024-07-10 Linzhuang Sun , Hao Liang , Jingxuan Wei , Linkun Sun , Bihui Yu , Bin Cui , Wentao Zhang

We use the Enron email corpus to study relationships in a network by applying six different measures of centrality. Our results came out of an in-semester undergraduate research seminar. The Enron corpus is well suited to statistical…

其他统计学 · 统计学 2015-08-06 Johanna Hardin , Ghassan Sarkis , P. C. Urc

The fundamental property of Cranfield-style evaluations, that system rankings are stable even when assessors disagree on individual relevance decisions, was validated on traditional test collections. However, the paradigm shift towards…

Cranfield-style retrieval evaluations with too few or too many relevant documents or with low inter-assessor agreement on relevance can reduce the reliability of observations. In evaluations with human assessors, information needs are often…

信息检索 · 计算机科学 2026-04-30 Jüri Keller , Maik Fröbe , Björn Engelmann , Fabian Haak , Timo Breuer , Birger Larsen , Philipp Schaer

Retrieval Augmented Generation (RAG) has become one of the most popular methods for bringing knowledge-intensive context to large language models (LLM) because of its ability to bring local context at inference time without the cost or data…

信息检索 · 计算机科学 2025-05-02 Michael J. Ryan , Danmei Xu , Chris Nivera , Daniel Campos

Improvements of entity-relationship (E-R) search techniques have been hampered by a lack of test collections, particularly for complex queries involving multiple entities and relationships. In this paper we describe a method for generating…

信息检索 · 计算机科学 2017-06-14 Pedro Saleiro , Natasa Milic-Frayling , Eduarda Mendes Rodrigues , Carlos Soares

Building test collections for Information Retrieval evaluation has traditionally been a resource-intensive and time-consuming task, primarily due to the dependence on manual relevance judgments. While various cost-effective strategies have…

信息检索 · 计算机科学 2025-01-07 Mehmet Deniz Türkmen , Mucahid Kutlu , Bahadir Altun , Gokalp Cosgun

What if Information Retrieval (IR) systems did not just retrieve relevant information that is stored in their indices, but could also "understand" it and synthesise it into a single document? We present a preliminary study that makes a…

信息检索 · 计算机科学 2016-06-28 Christina Lioma , Birger Larsen , Casper Petersen , Jakob Grue Simonsen

The exponential growth of scientific production makes secondary literature abridgements increasingly demanding. We introduce a new open-source framework for systematic reviews that significantly reduces time and workload for collecting and…

数字图书馆 · 计算机科学 2022-02-24 Angelo D'Ambrosio , Hajo Grundmann , Tjibbe Donker

Retrieval-augmented Generation (RAG) extends large language models (LLMs) with external knowledge but faces key challenges: restricted effective context length and redundancy in retrieved documents. Pure compression-based approaches reduce…

计算与语言 · 计算机科学 2025-07-09 Yiqiao Jin , Kartik Sharma , Vineeth Rakesh , Yingtong Dou , Menghai Pan , Mahashweta Das , Srijan Kumar

The emergence of open data portals necessitates more attention to protecting sensitive data before datasets get published and exchanged. To do so effectively, we observe the need to refine and broaden our definitions of sensitive data, and…

密码学与安全 · 计算机科学 2026-03-16 Liang Telkamp , Madelon Hulsebos
‹ 上一页 1 2 3 10 下一页 ›