中文
相关论文

相关论文: ORCAS: 18 Million Clicked Query-Document Pairs for…

200 篇论文

Many Information Retrieval (IR) models make use of offline statistical techniques to score documents for ranking over a single period, rather than use an online, dynamic system that is responsive to users over time. In this paper, we…

信息检索 · 计算机科学 2013-03-22 Marc Sloan , Jun Wang

Deep Learning Hard (DL-HARD) is a new annotated dataset designed to more effectively evaluate neural ranking models on complex topics. It builds on TREC Deep Learning (DL) topics by extensively annotating them with question intent…

信息检索 · 计算机科学 2021-05-18 Iain Mackie , Jeffery Dalton , Andrew Yates

Scholarly data are largely fragmented across siloed databases with divergent metadata and missing linkages among them. We present the Science Data Lake, a locally-deployable infrastructure built on DuckDB and simple Parquet files that…

数字图书馆 · 计算机科学 2026-03-04 Jonas Wilinski

Open datasets play a crucial role in three research domains that intersect data science and education: learning analytics, educational data mining, and artificial intelligence in education. Researchers in these domains apply computational…

计算机与社会 · 计算机科学 2026-04-14 Valdemar Švábenský , Brendan Flanagan , Erwin Daniel López Zapata , Atsushi Shimada

Cybersecurity professionals need hands-on training to prepare for managing the current advanced cyber threats. To practice cybersecurity skills, training participants use numerous software tools in computer-supported interactive learning…

计算机与社会 · 计算机科学 2023-07-18 Valdemar Švábenský , Jan Vykopal , Pavel Čeleda , Lydia Kraus

Online tracking has become of increasing concern in recent years, however our understanding of its extent to date has been limited to snapshots from web crawls. Previous at-tempts to measure the tracking ecosystem, have been done using…

计算机与社会 · 计算机科学 2020-07-07 Arjaldo Karaj , Sam Macbeth , Rémi Berson , Josep M. Pujol

The DRAGUN Track at TREC 2025 targets the growing need for effective support tools that help users evaluate the trustworthiness of online news. We describe the UR_Trecking system submitted for both Task 1 (critical question generation) and…

Deep Learning (DL) techniques are increasingly applied in scientific studies across various domains to address complex research questions. However, the methodological details of these DL models are often hidden in the unstructured text. As…

信息检索 · 计算机科学 2024-11-15 Vamsi Krishna Kommineni , Birgitta König-Ries , Sheeba Samuel

Quality pretraining data is often seen as the key to high-performance language models. However, progress in understanding pretraining data has been slow due to the costly pretraining runs required for data selection experiments. We present…

计算与语言 · 计算机科学 2025-03-11 Tristan Thrush , Christopher Potts , Tatsunori Hashimoto

Contextualisation has proven to be effective in tailoring \linebreak search results towards the users' information need. While this is true for a basic query search, the usage of contextual session information during exploratory search…

信息检索 · 计算机科学 2018-12-10 Zeljko Carevic , Sascha Schüller , Philipp Mayr , Norbert Fuhr

Click-through data has proven to be a valuable resource for improving search-ranking quality. Search engines can easily collect click data, but biases introduced in the data can make it difficult to use the data effectively. In order to…

机器学习 · 计算机科学 2020-02-13 Yingcheng Sun , Richard Kolacinski , Kenneth Loparo

With an expansive and ubiquitously available gold mine of educational data, Massive Open Online courses (MOOCs) have become the an important foci of learning analytics research. The hope is that this new surge of development will bring the…

人机交互 · 计算机科学 2014-07-29 Tanmay Sinha

Educational recommenders have received much less attention in comparison to e-commerce and entertainment-related recommenders, even though efficient intelligent tutors have great potential to improve learning gains. One of the main…

信息检索 · 计算机科学 2021-09-15 Sahan Bulathwela , Maria Perez-Ortiz , Erik Novak , Emine Yilmaz , John Shawe-Taylor

Key to any research involving session search is the understanding of how a user's queries evolve throughout the session. When a user creates a query reformulation, he or she is consciously retaining terms from their original query, removing…

信息检索 · 计算机科学 2016-01-21 Marc Sloan , Hui Yang , Jun Wang

Pre-trained language models have achieved great success in various large-scale information retrieval tasks. However, most of pretraining tasks are based on counterfeit retrieval data where the query produced by the tailored rule is assumed…

信息检索 · 计算机科学 2023-02-28 Xiangsheng Li , Xiaoshu Chen , Kunliang Wei , Bin Hu , Lei Jiang , Zeqian Huang , Zhanhui Kang

Conversational information seeking has evolved rapidly in the last few years with the development of Large Language Models (LLMs), providing the basis for interpreting and responding in a naturalistic manner to user requests. The extended…

信息检索 · 计算机科学 2024-05-07 Mohammad Aliannejadi , Zahra Abbasiantaeb , Shubham Chatterjee , Jeffery Dalton , Leif Azzopardi

There is a complex correlation among the data of scientific papers. The phenomenon reveals the data characteristics, laws, and correlations contained in the data of scientific and technological papers in specific fields, which can realize…

数字图书馆 · 计算机科学 2022-04-26 Jie Song , Meiyu Liang , Zhe Xue , Feifei Kou , Ang Li

World Wide Web is a huge data repository and is growing with the explosive rate of about 1 million pages a day. As the information available on World Wide Web is growing the usage of the web sites is also growing. Web log records each…

信息检索 · 计算机科学 2009-08-03 Ratnesh Kumar Jain , Dr. R. S. Kasana , Dr. Suresh Jain

Retrieval-Augmented Generation (RAG) delivers substantial value in knowledge-intensive applications. However, its generated responses often lack transparent reasoning paths that trace back to source evidence from retrieved documents. This…

计算与语言 · 计算机科学 2026-01-30 Jingyi Ren , Yekun Xu , Xiaolong Wang , Weitao Li , Ante Wang , Weizhi Ma , Yang Liu

Published experiments on spidering the Web suggest that, given training data in the form of a (relatively small) subgraph of the Web containing a subset of a selected class of target pages, it is possible to conduct a directed search and…

信息检索 · 计算机科学 2012-12-12 Joel Young , Thomas L. Dean