中文
相关论文

相关论文: Building Retrieval Systems for the ClueWeb22-B Cor…

200 篇论文

Query-based document summarization aims to extract or generate a summary of a document which directly answers or is relevant to the search query. It is an important technique that can be beneficial to a variety of applications such as…

人工智能 · 计算机科学 2020-10-29 Mingjun Zhao , Shengli Yan , Bang Liu , Xinwang Zhong , Qian Hao , Haolan Chen , Di Niu , Bowei Long , Weidong Guo

A novel pseudocode search engine is designed to facilitate efficient retrieval and search of academic papers containing pseudocode. By leveraging Elasticsearch, the system enables users to search across various facets of a paper, such as…

信息检索 · 计算机科学 2024-11-20 Levent Toksoz , Mukund Srinath , Gang Tan , C. Lee Giles

In this paper, we provide a detailed overview of the models used for information retrieval in the first and second stages of the typical processing chain. We discuss the current state-of-the-art models, including methods based on terms,…

信息检索 · 计算机科学 2024-02-16 Kailash A. Hambarde , Hugo Proenca

WikiKG90Mv2 in NeurIPS 2022 is a large encyclopedic knowledge graph. Embedding knowledge graphs into continuous vector spaces is important for many practical applications, such as knowledge acquisition, question answering, and…

计算与语言 · 计算机科学 2026-03-31 Feng Nie , Zhixiu Ye , Sifa Xie , Shuang Wu , Xin Yuan , Liang Yao , Jiazhen Peng , Xu Cheng

Natural language-based vehicle retrieval is a task to find a target vehicle within a given image based on a natural language description as a query. This technology can be applied to various areas including police searching for a suspect…

计算机视觉与模式识别 · 计算机科学 2023-08-04 Sangrok Lee , Taekang Woo , Sang Hun Lee

Training deep research agents requires long-horizon trajectories that interleave search, evidence aggregation, and multi-step reasoning. However, existing data collection pipelines typically rely on proprietary web APIs, making large-scale…

信息检索 · 计算机科学 2026-03-24 Zhuofeng Li , Dongfu Jiang , Xueguang Ma , Haoxiang Zhang , Ping Nie , Yuyu Zhang , Kai Zou , Jianwen Xie , Yu Zhang , Wenhu Chen

Large-scale text retrieval technology has been widely used in various practical business scenarios. This paper presents our systems for the TREC 2022 Deep Learning Track. We explain the hybrid text retrieval and multi-stage text ranking…

信息检索 · 计算机科学 2023-08-24 Guangwei Xu , Yangzhao Zhang , Longhui Zhang , Dingkun Long , Pengjun Xie , Ruijie Guo

We introduce WebFAQ 2.0, a new version of the WebFAQ dataset, containing 198 million FAQ-based natural question-answer pairs across 108 languages. Compared to the previous version, it significantly expands multilingual coverage and the…

信息检索 · 计算机科学 2026-02-20 Michael Dinzinger , Laura Caspari , Ali Salman , Irvin Topi , Jelena Mitrović , Michael Granitzer

Recent lay language generation systems have used Transformer models trained on a parallel corpus to increase health information accessibility. However, the applicability of these models is constrained by the limited size and topical breadth…

计算与语言 · 计算机科学 2024-01-26 Yue Guo , Wei Qiu , Gondy Leroy , Sheng Wang , Trevor Cohen

Legal retrieval techniques play an important role in preserving the fairness and equality of the judicial system. As an annually well-known international competition, COLIEE aims to advance the development of state-of-the-art retrieval…

信息检索 · 计算机科学 2024-04-02 Haitao Li , You Chen , Zhekai Ge , Qingyao Ai , Yiqun Liu , Quan Zhou , Shuai Huo

Existing Text-to-SQL generators require the entire schema to be encoded with the user text. This is expensive or impractical for large databases with tens of thousands of columns. Standard dense retrieval techniques are inadequate for…

计算与语言 · 计算机科学 2023-11-03 Mayank Kothyari , Dhruva Dhingra , Sunita Sarawagi , Soumen Chakrabarti

In this technical report, we present Skywork-13B, a family of large language models (LLMs) trained on a corpus of over 3.2 trillion tokens drawn from both English and Chinese texts. This bilingual foundation model is the most extensively…

This paper presents our system for SemEval-2025 Task 7: Multilingual and Crosslingual Fact-Checked Claim Retrieval. In an era where misinformation spreads rapidly, effective fact-checking is increasingly critical. We introduce TriAligner, a…

Interdisciplinary scientific research is increasingly important in knowledge production, funding policies, and academic discussions on scholarly communication. While many studies focus on interdisciplinary corpora defined a priori --…

数字图书馆 · 计算机科学 2025-10-29 Malena Mendez Isla , Agustin Mauro , Diego Kozlowski

Despite the advancements in search engine features, ranking methods, technologies, and the availability of programmable APIs, current-day open-access digital libraries still rely on crawl-based approaches for acquiring their underlying…

信息检索 · 计算机科学 2016-04-19 Sujatha Das Gollapalli , Krutarth Patel , Cornelia Caragea

Statutory law retrieval is a typical problem in legal language processing, that has various practical applications in law engineering. Modern deep learning-based retrieval methods have achieved significant results for this problem. However,…

计算与语言 · 计算机科学 2024-10-17 Hai-Long Nguyen , Tan-Minh Nguyen , Duc-Minh Nguyen , Thi-Hai-Yen Vuong , Ha-Thanh Nguyen , Xuan-Hieu Phan

A classification scheme of a scientific subject gives an overview of its body of knowledge. It can also be used to facilitate access to research articles and other materials related to the subject. For example, the ACM Computing…

Query by Example is a well-known information retrieval task in which a document is chosen by the user as the search query and the goal is to retrieve relevant documents from a large collection. However, a document often covers multiple…

信息检索 · 计算机科学 2021-11-09 Sheshera Mysore , Tim O'Gorman , Andrew McCallum , Hamed Zamani

This paper introduces the Multi-Genre Natural Language Inference (MultiNLI) corpus, a dataset designed for use in the development and evaluation of machine learning models for sentence understanding. In addition to being one of the largest…

计算与语言 · 计算机科学 2018-02-21 Adina Williams , Nikita Nangia , Samuel R. Bowman

There are settings in which reproducibility of ranked lists is desirable, such as when extracting a subset of an evolving document corpus for downstream research tasks or in domains such as patent retrieval or in medical systematic reviews,…

信息检索 · 计算机科学 2024-11-07 Moritz Staudinger , Florina Piroi , Andreas Rauber