English
Related papers

Related papers: ORCAS: 18 Million Clicked Query-Document Pairs for…

200 papers

The Archive Query Log (AQL) is a previously unused, comprehensive query log collected at the Internet Archive over the last 25 years. Its first version includes 356 million queries, 166 million search result pages, and 1.7 billion search…

This paper describes Brown University's submission to the TREC 2019 Deep Learning track. We followed a 2-phase method for producing a ranking of passages for a given input query: In the the first phase, the user's query is expanded by…

Information Retrieval · Computer Science 2020-09-10 George Zerveas , Ruochen Zhang , Leila Kim , Carsten Eickhoff

We benchmark Conformer-Kernel models under the strict blind evaluation setting of the TREC 2020 Deep Learning track. In particular, we study the impact of incorporating: (i) Explicit term matching to complement matching based on learned…

Information Retrieval · Computer Science 2021-02-15 Bhaskar Mitra , Sebastian Hofstatter , Hamed Zamani , Nick Craswell

Recent years have witnessed great progress on applying pre-trained language models, e.g., BERT, to information retrieval (IR) tasks. Hyperlinks, which are commonly used in Web pages, have been leveraged for designing pre-training…

Information Retrieval · Computer Science 2022-09-15 Jiawen Wu , Xinyu Zhang , Yutao Zhu , Zheng Liu , Zikai Guo , Zhaoye Fei , Ruofei Lai , Yongkang Wu , Zhao Cao , Zhicheng Dou

Finding relevant literature underpins the practice of evidence-based medicine. From 2014 to 2016, TREC conducted a clinical decision support track, wherein participants were tasked with finding articles relevant to clinical questions posed…

Information Retrieval · Computer Science 2018-01-30 Vincent Nguyen , Sarvnaz Karimi , Sara Falamaki , Cecile Paris

A large number of URLs are made public by various platforms for security analysis, archiving, and paste sharing -- such as VirusTotal, URLScan.io, Hybrid Analysis, the Wayback Machine, and RedHunt. These services may unintentionally expose…

Cryptography and Security · Computer Science 2026-02-26 Tarek Ramadan , AbdelRahman Abdou , Mohammad Mannan , Amr Youssef

ClueWeb22, the newest iteration of the ClueWeb line of datasets, provides 10 billion web pages affiliated with rich information. Its design was influenced by the need for a high quality, large scale web corpus to support a range of academic…

Information Retrieval · Computer Science 2022-12-05 Arnold Overwijk , Chenyan Xiong , Xiao Liu , Cameron VandenBerg , Jamie Callan

General purpose Search Engines (SEs) crawl all domains (e.g., Sports, News, Entertainment) of the Web, but sometimes the informational need of a query is restricted to a particular domain (e.g., Medical). We leverage the work of SEs as part…

Information Retrieval · Computer Science 2016-05-04 Alexander Nwala , Michael Nelson

Among the many sources of event data available today, a prominent one is user interaction data. User activity may be recorded during the use of an application or website, resulting in a type of user interaction data often called click data.…

Databases · Computer Science 2022-04-11 Marco Pegoraro , Merih Seran Uysal , Tom-Hendrik Hülsmann , Wil M. P. van der Aalst

The task of session search focuses on using interaction data to improve relevance for the user's next query at the session level. In this paper, we formulate session search as a personalization task under the framework of learning to rank.…

Information Retrieval · Computer Science 2020-09-18 Saad Aloteibi , Stephen Clark

Scholars studying organizations often work with multiple datasets lacking shared identifiers or covariates. In such situations, researchers usually use approximate string ("fuzzy") matching methods to combine datasets. String matching,…

Social and Information Networks · Computer Science 2025-09-24 Brian Libgober , Connor T. Jerzak

We present SemOpenAlex, an extensive RDF knowledge graph that contains over 26 billion triples about scientific publications and their associated entities, such as authors, institutions, journals, and concepts. SemOpenAlex is licensed under…

Digital Libraries · Computer Science 2023-08-08 Michael Färber , David Lamprecht , Johan Krause , Linn Aung , Peter Haase

Search clarification has recently attracted much attention due to its applications in search engines. It has also been recognized as a major component in conversational information seeking systems. Despite its importance, the research…

Information Retrieval · Computer Science 2020-06-19 Hamed Zamani , Gord Lueck , Everest Chen , Rodolfo Quispe , Flint Luu , Nick Craswell

Web query log data contain information useful to research; however, release of such data can re-identify the search engine users issuing the queries. These privacy concerns go far beyond removing explicitly identifying information such as…

Databases · Computer Science 2010-12-06 Amin Milani Fard , Ke Wang

Open Educational Resources (OERs) are openly licensed educational materials that are widely used for learning. Nowadays, many online learning repositories provide millions of OERs. Therefore, it is exceedingly difficult for learners to find…

Computers and Society · Computer Science 2021-01-20 Mohammadreza Tavakoli , Mirette Elias , Gábor Kismihók , Sören Auer

This article provides a quantitative analysis of privacy-compromising mechanisms on 1 million popular websites. Findings indicate that nearly 9 in 10 websites leak user data to parties of which the user is likely unaware; more than 6 in 10…

Cryptography and Security · Computer Science 2015-11-03 Timothy Libert

Training statistical dialog models in spoken dialog systems (SDS) requires large amounts of annotated data. The lack of scalable methods for data mining and annotation poses a significant hurdle for state-of-the-art statistical dialog…

Computation and Language · Computer Science 2016-06-28 Lu Wang , Larry Heck , Dilek Hakkani-Tur

Translating verbose information needs into crisp search queries is a phenomenon that is ubiquitous but hardly understood. Insights into this process could be valuable in several applications, including synthesizing large privacy-friendly…

Information Retrieval · Computer Science 2021-06-04 Asia J. Biega , Jana Schmidt , Rishiraj Saha Roy

Large-scale test collections play a crucial role in Information Retrieval (IR) research. However, according to the Cranfield paradigm and the research into publicly available datasets, the existing information retrieval research studies are…

Information Retrieval · Computer Science 2025-01-28 Hossein A. Rahmani , Xi Wang , Emine Yilmaz , Nick Craswell , Bhaskar Mitra , Paul Thomas

Nowadays online searches are undeniably the most common form of information gathering, as witnessed by billions of clicks generated each day on search engines. In this work we describe online searches as foraging processes that take place…

Physics and Society · Physics 2017-04-05 Xiangwen Wang , Michel Pleimling