中文
相关论文

相关论文: Building Retrieval Systems for the ClueWeb22-B Cor…

200 篇论文

We present Lit2Vec, a reproducible workflow for constructing and validating a chemistry corpus from the Semantic Scholar Open Research Corpus using conservative, metadata-based license screening. Using this workflow, we assembled an…

数据库 · 计算机科学 2026-04-15 Mahmoud Amiri , Jamile Mohammad Jafari , Sara Mostafapour , Thomas Bocklitz

Recent multilingual named entity recognition (NER) work has shown that large language models (LLMs) can provide effective synthetic supervision, yet such datasets have mostly appeared as by-products of broader experiments rather than as…

计算与语言 · 计算机科学 2025-12-17 Jonas Golde , Patrick Haller , Alan Akbik

The number of biomedical research articles published has doubled in the past 20 years. Search engine based systems naturally center around searching, but researchers may not have a clear goal in mind, or the goal may be expressed in a query…

数字图书馆 · 计算机科学 2017-10-25 Jessica Perrie , Yanqi Hao , Zack Hayat , Recep Colak , Kelly Lyons , Shankar Vembu , Sam Molyneux

Retrieval-augmented question answering over heterogeneous corpora requires connected evidence across text, tables, and graph nodes. While entity-level knowledge graphs support structured access, they are costly to construct and maintain,…

信息检索 · 计算机科学 2026-02-20 Prasham Titiya , Rohit Khoja , Tomer Wolfson , Vivek Gupta , Dan Roth

The web contains vast repositories of unstructured text. We investigate the opportunity for building a knowledge graph from these text sources. We generate a set of triples which can be used in knowledge gathering and integration. We define…

计算与语言 · 计算机科学 2018-11-15 Shaun D'Souza

An important goal of online platforms is to enable content discovery, i.e. allow users to find a catalog entity they were not familiar with. A pre-requisite to discover an entity, e.g. a book, with a search engine is that the entity is…

信息检索 · 计算机科学 2023-03-22 Gustavo Penha , Enrico Palumbo , Maryam Aziz , Alice Wang , Hugues Bouchard

This paper studies fast fusion of dense retrieval and sparse lexical retrieval, and proposes a cluster-based selective dense retrieval method called CluSD guided by sparse lexical retrieval. CluSD takes a lightweight cluster-based approach…

信息检索 · 计算机科学 2025-02-18 Yingrui Yang , Parker Carlson , Yifan Qiao , Wentai Xie , Shanxiu He , Tao Yang

Retrieval-Augmented Generation (RAG) systems critically depend on retrieval quality, yet no systematic comparison of modern retrieval methods exists for heterogeneous documents containing both text and tabular data. We benchmark ten…

信息检索 · 计算机科学 2026-04-03 Meftun Akarsu , Recep Kaan Karaman , Christopher Mierbach

The main aim of an information retrieval system is to extract appropriate information from an enormous collection of data based on users need. The basic concept of the information retrieval system is that when a user sends out a query, the…

信息检索 · 计算机科学 2020-12-17 Abdulmalik Johar

Scientific knowledge discovery increasingly relies on large language models, yet many existing scholarly assistants depend on proprietary systems with tens or hundreds of billions of parameters. Such reliance limits reproducibility and…

信息检索 · 计算机科学 2026-04-23 Florian Kelber , Matthias Jobst , Yuni Susanti , Michael Färber

Document classification for text, images and other applicable entities has long been a focus of research in academia and also finds application in many industrial settings. Amidst a plethora of approaches to solve such problems,…

机器学习 · 计算机科学 2016-06-06 Faizan Javed , Matt McNair , Ferosh Jacob , Meng Zhao

Passage retrieval addresses the problem of locating relevant passages, usually from a large corpus, given a query. In practice, lexical term-matching algorithms like BM25 are popular choices for retrieval owing to their efficiency. However,…

Most of the current methods for mining parallel texts from the web assume that web pages of web sites share same structure across languages. We believe that there still exists a non-negligible amount of parallel data spread across sources…

计算与语言 · 计算机科学 2018-04-30 Jakub Kúdela , Irena Holubová , Ondřej Bojar

HC4 is a new suite of test collections for ad hoc Cross-Language Information Retrieval (CLIR), with Common Crawl News documents in Chinese, Persian, and Russian, topics in English and in the document languages, and graded relevance…

信息检索 · 计算机科学 2022-01-26 Dawn Lawrie , James Mayfield , Douglas Oard , Eugene Yang

Commercial web search engines employ near-duplicate detection to ensure that users see each relevant result only once, albeit the underlying web crawls typically include (near-)duplicates of many web pages. We revisit the risks and…

A publicly available dataset for federated search reflecting a real web environment has long been absent, making it difficult for researchers to test the validity of their federated search algorithms for the web setting. We present several…

信息检索 · 计算机科学 2016-09-16 Dong Nguyen , Thomas Demeester , Dolf Trieschnigg , Djoerd Hiemstra

Augmenting large language models (LLMs) with browsing tools substantially improves their potential as deep search agents to solve complex, real-world tasks. Yet, open LLMs still perform poorly in such settings due to limited long-horizon…

计算与语言 · 计算机科学 2025-10-15 Rui Lu , Zhenyu Hou , Zihan Wang , Hanchen Zhang , Xiao Liu , Yujiang Li , Shi Feng , Jie Tang , Yuxiao Dong

Retrieval-augmented generation over semi-structured sources such as HTML is constrained by a mismatch between document structure and the flat, sequence-based interfaces of today's embedding and generative models. Retrieval pipelines often…

信息检索 · 计算机科学 2026-04-24 Mike Rainey , Umut Acar , Muhammed Sezer

Detectives frequently engage in information detection and reasoning simultaneously when making decisions across various cases, especially when confronted with a vast amount of information. With the rapid development of large language…

计算与语言 · 计算机科学 2024-03-21 Zhouhong Gu , Lin Zhang , Jiangjie Chen , Haoning Ye , Xiaoxuan Zhu , Zihan Li , Zheyu Ye , Yan Gao , Yao Hu , Yanghua Xiao , Hongwei Feng

The search for research datasets is as important as laborious. Due to the importance of the choice of research data in further research, this decision must be made carefully. Additionally, because of the growing amounts of data in almost…

信息检索 · 计算机科学 2022-10-03 Jüri Keller , Leon Paul Mondrian Munz
‹ 上一页 1 8 9 10 下一页 ›