中文
相关论文

相关论文: BookReconciler: An Open-Source Tool for Metadata E…

200 篇论文

Identifying suitable datasets for a research question remains challenging because existing dataset search engines rely heavily on metadata quality and keyword overlap, which often fail to capture the semantic intent of scientific…

数字图书馆 · 计算机科学 2026-01-09 Zhiyin Tan , Changxu Duan

Digital libraries for research, such as the ACM Digital Library or Semantic Scholar, do not enable the machine-supported, efficient reuse of scientific knowledge (e.g., in synthesis research). This is because these libraries are based on…

信息检索 · 计算机科学 2025-11-12 Hadi Ghaemi , Lauren Snyder , Markus Stocker

One technique to improve the retrieval effectiveness of a search engine is to expand documents with terms that are related or representative of the documents' content.From the perspective of a question answering system, this might comprise…

信息检索 · 计算机科学 2019-09-26 Rodrigo Nogueira , Wei Yang , Jimmy Lin , Kyunghyun Cho

In this work we present a systematic empirical study focused on the suitability of the state-of-the-art multilingual encoders for cross-lingual document and sentence retrieval tasks across a number of diverse language pairs. We first treat…

计算与语言 · 计算机科学 2021-12-22 Robert Litschko , Ivan Vulić , Simone Paolo Ponzetto , Goran Glavaš

The OpenCitations organization is working on ingesting citation data and bibliographic metadata directly provided by the community (e.g., scholars and publishers). The aim is to improve the general coverage of open citations, which is still…

数字图书馆 · 计算机科学 2022-09-26 Arcangelo Massari , Ivan Heibi

The goal of local citation recommendation is to recommend a missing reference from the local citation context and optionally also from the global context. To balance the tradeoff between speed and accuracy of citation recommendation in the…

信息检索 · 计算机科学 2022-03-18 Nianlong Gu , Yingqiang Gao , Richard H. R. Hahnloser

The task of information retrieval is an important component of many natural language processing systems, such as open domain question answering. While traditional methods were based on hand-crafted features, continuous representations based…

计算与语言 · 计算机科学 2022-08-05 Gautier Izacard , Edouard Grave

Manual digitization of bibliographic metadata is time consuming and labor intensive, especially for historical and real-world archives with highly variable formatting across documents. Despite advances in machine learning, the absence of…

计算机视觉与模式识别 · 计算机科学 2025-03-26 Jan Kohút , Martin Dočekal , Michal Hradiš , Marek Vaško

Developed so far, multi-document summarization has reached its bottleneck due to the lack of sufficient training data and diverse categories of documents. Text classification just makes up for these deficiencies. In this paper, we propose a…

计算与语言 · 计算机科学 2016-11-29 Ziqiang Cao , Wenjie Li , Sujian Li , Furu Wei

In many government applications we often find that information about entities, such as persons, are available in disparate data sources such as passports, driving licences, bank accounts, and income tax records. Similar scenarios are…

数据库 · 计算机科学 2014-02-19 Pankaj Malhotra , Puneet Agarwal , Gautam Shroff

Due to the availability of references of research papers and the rich information contained in papers, various citation analysis approaches have been proposed to identify similar documents for scholar recommendation. Despite of the success…

信息检索 · 计算机科学 2017-03-21 Han Tian , Hankz Hankui Zhuo

Urban data support a wide range of applications across multiple disciplines. However, at the global scale, there is no unified platform for urban data discovery. As a result, researchers often have to manually search through websites or…

信息检索 · 计算机科学 2026-04-21 Runwen You , Tong Xia , Jingzhi Wang , Jiankun Zhang , Tengyao Tu , Jinghua Piao , Yi Chang , Yong Li

The task of expert finding has been getting increasing attention in information retrieval literature. However, the current state-of-the-art is still lacking in principled approaches for combining different sources of evidence. This paper…

信息检索 · 计算机科学 2020-10-28 Catarina Moreira , Bruno Martins , Pável Calado

The problem of reversing the compilation process, decompilation, is an important tool in reverse engineering of computer software. Recently, researchers have proposed using techniques from neural machine translation to automate the process…

密码学与安全 · 计算机科学 2022-12-20 Iman Hosseini , Brendan Dolan-Gavitt

Multilingual dense retrieval aims to retrieve relevant documents across different languages based on a unified retriever model. The challenge lies in aligning representations of different languages in a shared vector space. The common…

信息检索 · 计算机科学 2025-09-12 Chao Huang , Fengran Mo , Yufeng Chen , Changhao Guan , Zhenrui Yue , Xinyu Wang , Jinan Xu , Kaiyu Huang

Generalist visual captioning goes beyond a simple appearance description task, but requires integrating a series of visual cues into a caption and handling various visual domains. In this task, current open-source models present a large…

计算机视觉与模式识别 · 计算机科学 2025-10-17 Zhenxin Lei , Zhangwei Gao , Changyao Tian , Erfei Cui , Guanzhou Chen , Danni Yang , Yuchen Duan , Zhaokai Wang , Wenhao Li , Weiyun Wang , Xiangyu Zhao , Jiayi Ji , Yu Qiao , Wenhai Wang , Gen Luo

Large Language Models (LLMs) have shown remarkable prowess in text generation, yet producing long-form, factual documents grounded in extensive external knowledge bases remains a significant challenge. Existing "top-down" methods, which…

计算与语言 · 计算机科学 2025-09-17 Binquan Ji , Jiaqi Wang , Ruiting Li , Xingchen Han , Yiyang Qi , Shichao Wang , Yifei Lu , Yuantao Han , Feiliang Ren

Reference lists in scholarly manuscripts frequently contain errors, including incorrect identifiers, incomplete metadata, misattributed authors, and mismatches between preprint and published versions. These problems are tedious to repair…

数字图书馆 · 计算机科学 2026-03-19 Junhyeok Lee

A central challenge in science is to understand how systems behaviors emerge from complex networks. This often requires aggregating, reusing, and integrating heterogeneous information. Supplementary spreadsheets to articles are a key data…

Event Coreference Resolution (ECR) is the task of clustering event mentions that refer to the same real-world event. Despite significant advancements, ECR research faces two main challenges: limited generalizability across domains due to…

人工智能 · 计算机科学 2024-06-21 Yuncong Li , Tianhua Xu , Sheng-hua Zhong , Haiqin Yang