中文
相关论文

相关论文: ORCAS: 18 Million Clicked Query-Document Pairs for…

200 篇论文

The Archive Query Log (AQL) is a previously unused, comprehensive query log collected at the Internet Archive over the last 25 years. Its first version includes 356 million queries, 166 million search result pages, and 1.7 billion search…

This paper describes Brown University's submission to the TREC 2019 Deep Learning track. We followed a 2-phase method for producing a ranking of passages for a given input query: In the the first phase, the user's query is expanded by…

信息检索 · 计算机科学 2020-09-10 George Zerveas , Ruochen Zhang , Leila Kim , Carsten Eickhoff

We benchmark Conformer-Kernel models under the strict blind evaluation setting of the TREC 2020 Deep Learning track. In particular, we study the impact of incorporating: (i) Explicit term matching to complement matching based on learned…

信息检索 · 计算机科学 2021-02-15 Bhaskar Mitra , Sebastian Hofstatter , Hamed Zamani , Nick Craswell

Recent years have witnessed great progress on applying pre-trained language models, e.g., BERT, to information retrieval (IR) tasks. Hyperlinks, which are commonly used in Web pages, have been leveraged for designing pre-training…

信息检索 · 计算机科学 2022-09-15 Jiawen Wu , Xinyu Zhang , Yutao Zhu , Zheng Liu , Zikai Guo , Zhaoye Fei , Ruofei Lai , Yongkang Wu , Zhao Cao , Zhicheng Dou

Finding relevant literature underpins the practice of evidence-based medicine. From 2014 to 2016, TREC conducted a clinical decision support track, wherein participants were tasked with finding articles relevant to clinical questions posed…

信息检索 · 计算机科学 2018-01-30 Vincent Nguyen , Sarvnaz Karimi , Sara Falamaki , Cecile Paris

A large number of URLs are made public by various platforms for security analysis, archiving, and paste sharing -- such as VirusTotal, URLScan.io, Hybrid Analysis, the Wayback Machine, and RedHunt. These services may unintentionally expose…

密码学与安全 · 计算机科学 2026-02-26 Tarek Ramadan , AbdelRahman Abdou , Mohammad Mannan , Amr Youssef

ClueWeb22, the newest iteration of the ClueWeb line of datasets, provides 10 billion web pages affiliated with rich information. Its design was influenced by the need for a high quality, large scale web corpus to support a range of academic…

信息检索 · 计算机科学 2022-12-05 Arnold Overwijk , Chenyan Xiong , Xiao Liu , Cameron VandenBerg , Jamie Callan

General purpose Search Engines (SEs) crawl all domains (e.g., Sports, News, Entertainment) of the Web, but sometimes the informational need of a query is restricted to a particular domain (e.g., Medical). We leverage the work of SEs as part…

信息检索 · 计算机科学 2016-05-04 Alexander Nwala , Michael Nelson

Among the many sources of event data available today, a prominent one is user interaction data. User activity may be recorded during the use of an application or website, resulting in a type of user interaction data often called click data.…

数据库 · 计算机科学 2022-04-11 Marco Pegoraro , Merih Seran Uysal , Tom-Hendrik Hülsmann , Wil M. P. van der Aalst

The task of session search focuses on using interaction data to improve relevance for the user's next query at the session level. In this paper, we formulate session search as a personalization task under the framework of learning to rank.…

信息检索 · 计算机科学 2020-09-18 Saad Aloteibi , Stephen Clark

Scholars studying organizations often work with multiple datasets lacking shared identifiers or covariates. In such situations, researchers usually use approximate string ("fuzzy") matching methods to combine datasets. String matching,…

社会与信息网络 · 计算机科学 2025-09-24 Brian Libgober , Connor T. Jerzak

We present SemOpenAlex, an extensive RDF knowledge graph that contains over 26 billion triples about scientific publications and their associated entities, such as authors, institutions, journals, and concepts. SemOpenAlex is licensed under…

数字图书馆 · 计算机科学 2023-08-08 Michael Färber , David Lamprecht , Johan Krause , Linn Aung , Peter Haase

Search clarification has recently attracted much attention due to its applications in search engines. It has also been recognized as a major component in conversational information seeking systems. Despite its importance, the research…

信息检索 · 计算机科学 2020-06-19 Hamed Zamani , Gord Lueck , Everest Chen , Rodolfo Quispe , Flint Luu , Nick Craswell

Web query log data contain information useful to research; however, release of such data can re-identify the search engine users issuing the queries. These privacy concerns go far beyond removing explicitly identifying information such as…

数据库 · 计算机科学 2010-12-06 Amin Milani Fard , Ke Wang

Open Educational Resources (OERs) are openly licensed educational materials that are widely used for learning. Nowadays, many online learning repositories provide millions of OERs. Therefore, it is exceedingly difficult for learners to find…

计算机与社会 · 计算机科学 2021-01-20 Mohammadreza Tavakoli , Mirette Elias , Gábor Kismihók , Sören Auer

This article provides a quantitative analysis of privacy-compromising mechanisms on 1 million popular websites. Findings indicate that nearly 9 in 10 websites leak user data to parties of which the user is likely unaware; more than 6 in 10…

密码学与安全 · 计算机科学 2015-11-03 Timothy Libert

Training statistical dialog models in spoken dialog systems (SDS) requires large amounts of annotated data. The lack of scalable methods for data mining and annotation poses a significant hurdle for state-of-the-art statistical dialog…

计算与语言 · 计算机科学 2016-06-28 Lu Wang , Larry Heck , Dilek Hakkani-Tur

Translating verbose information needs into crisp search queries is a phenomenon that is ubiquitous but hardly understood. Insights into this process could be valuable in several applications, including synthesizing large privacy-friendly…

信息检索 · 计算机科学 2021-06-04 Asia J. Biega , Jana Schmidt , Rishiraj Saha Roy

Large-scale test collections play a crucial role in Information Retrieval (IR) research. However, according to the Cranfield paradigm and the research into publicly available datasets, the existing information retrieval research studies are…

信息检索 · 计算机科学 2025-01-28 Hossein A. Rahmani , Xi Wang , Emine Yilmaz , Nick Craswell , Bhaskar Mitra , Paul Thomas

Nowadays online searches are undeniably the most common form of information gathering, as witnessed by billions of clicks generated each day on search engines. In this work we describe online searches as foraging processes that take place…

物理与社会 · 物理学 2017-04-05 Xiangwen Wang , Michel Pleimling