中文
相关论文

相关论文: DMDD: A Large-Scale Dataset for Dataset Mentions D…

200 篇论文

We focus on electronic theses and dissertations (ETDs), aiming to improve access and expand their utility, since more than 6 million are publicly available, and they constitute an important corpus to aid research and education across…

计算机视觉与模式识别 · 计算机科学 2021-06-30 Sampanna Yashwant Kahu , William A. Ingram , Edward A. Fox , Jian Wu

Electronic theses and dissertations (ETDs) have been proposed, advocated, and generated for more than 25 years. Although ETDs are hosted by commercial or institutional digital library repositories, they are still an understudied type of…

计算机视觉与模式识别 · 计算机科学 2023-11-09 Muntabir Hasan Choudhury , Lamia Salsabil , William A. Ingram , Edward A. Fox , Jian Wu

Funding acknowledgments in scholarly publications provide large-scale trace data on organizations that support scientific research. We present a dataset for linking global science funding organizations to research publications by…

数字图书馆 · 计算机科学 2026-03-26 Jacob Aarup Dalsgaard , Filipi Nascimento Silva , Jin AI

Dataset distillation (DD) condenses large datasets into compact yet informative substitutes, preserving performance comparable to the original dataset while reducing storage, transmission costs, and computational consumption. However,…

计算机视觉与模式识别 · 计算机科学 2025-07-01 Yawen Zou , Guang Li , Duo Su , Zi Wang , Jun Yu , Chao Zhang

In this paper, we propose a method to automatically classify AI-related documents from large-scale literature databases, leading to the creation of an AI-related literature dataset, named DeepDiveAI. The dataset construction approach…

人工智能 · 计算机科学 2025-04-23 Zhou Xiaochen , Liang Xingzhou , Zou Hui , Lu Yi , Qu Jingjing

The volume of academic paper submissions and publications is growing at an ever increasing rate. While this flood of research promises progress in various fields, the sheer volume of output inherently increases the amount of noise. We…

信息检索 · 计算机科学 2020-05-22 Marko Stamenovic , Jeibo Luo

Understanding how two datasets differ can help us determine whether one dataset under-represents certain sub-populations, and provides insights into how well models will generalize across datasets. Representative points selected by a…

统计方法学 · 统计学 2021-08-06 Sinead A. Williamson , Jette Henderson

Keeping track of all relevant recent publications and experimental results for a research area is a challenging task. Prior work has demonstrated the efficacy of information extraction models in various scientific areas. Recently, several…

计算与语言 · 计算机科学 2023-10-25 Timo Pierre Schrader , Matteo Finco , Stefan Grünewald , Felix Hildebrand , Annemarie Friedrich

Scientific article summarization is challenging: large, annotated corpora are not available, and the summary should ideally include the article's impacts on research community. This paper provides novel solutions to these two challenges. We…

计算与语言 · 计算机科学 2019-09-17 Michihiro Yasunaga , Jungo Kasai , Rui Zhang , Alexander R. Fabbri , Irene Li , Dan Friedman , Dragomir R. Radev

A main challenge of data-driven sciences is how to make maximal use of the progressively expanding databases of experimental datasets in order to keep research cumulative. We introduce the idea of a modeling-based dataset retrieval engine…

定量方法 · 定量生物学 2015-06-19 Ali Faisal , Jaakko Peltonen , Elisabeth Georgii , Johan Rung , Samuel Kaski

Citation recommendation is the task of finding appropriate citations based on a given piece of text. The proposed datasets for this task consist mainly of several scientific fields, lacking some core ones, such as law. Furthermore, citation…

信息检索 · 计算机科学 2023-11-13 Doğukan Arslan , Saadet Sena Erdoğan , Gülşen Eryiğit

In scientific research, the ability to effectively retrieve relevant documents based on complex, multifaceted queries is critical. Existing evaluation datasets for this task are limited, primarily due to the high cost and effort required to…

信息检索 · 计算机科学 2023-10-31 Jianyou Wang , Kaicheng Wang , Xiaoyue Wang , Prudhviraj Naidu , Leon Bergen , Ramamohan Paturi

Data quality is crucial for training accurate, unbiased, and trustworthy machine learning models as well as for their correct evaluation. Recent works, however, have shown that even popular datasets used to train and evaluate…

计算与语言 · 计算机科学 2024-03-12 Jan-Christoph Klie , Richard Eckart de Castilho , Iryna Gurevych

Deep learning technology has developed unprecedentedly in the last decade and has become the primary choice in many application domains. This progress is mainly attributed to a systematic collaboration in which rapidly growing computing…

机器学习 · 计算机科学 2023-12-27 Shiye Lei , Dacheng Tao

Tools to explore scientific literature are essential for scientists, especially in biomedicine, where about a million new papers are published every year. Many such tools provide users the ability to search for specific entities (e.g.…

计算与语言 · 计算机科学 2021-07-05 Sunil Mohan , Rico Angell , Nick Monath , Andrew McCallum

Scientific figure interpretation is a crucial capability for AI-driven scientific assistants built on advanced Large Vision Language Models. However, current datasets and benchmarks primarily focus on simple charts or other relatively…

Highly specific datasets of scientific literature are important for both research and education. However, it is difficult to build such datasets at scale. A common approach is to build these datasets reductively by applying topic modeling…

Knowledge about software used in scientific investigations is important for several reasons, for instance, to enable an understanding of provenance and methods involved in data handling. However, software is usually not formally cited, but…

信息检索 · 计算机科学 2021-08-23 David Schindler , Felix Bensmann , Stefan Dietze , Frank Krüger

In this paper, we ask the research question of whether all the datasets in the benchmark are necessary. We approach this by first characterizing the distinguishability of datasets when comparing different systems. Experiments on 9 datasets…

计算与语言 · 计算机科学 2022-05-05 Yang Xiao , Jinlan Fu , See-Kiong Ng , Pengfei Liu

We present a dataset of 833k paragraphs extracted from CC-BY licensed scientific publications, classified into four categories: acknowledgments, data mentions, software/code mentions, and clinical trial mentions. The paragraphs are…

计算与语言 · 计算机科学 2025-10-28 Eric Jeangirard