English
Related papers

Related papers: DMDD: A Large-Scale Dataset for Dataset Mentions D…

200 papers

We focus on electronic theses and dissertations (ETDs), aiming to improve access and expand their utility, since more than 6 million are publicly available, and they constitute an important corpus to aid research and education across…

Computer Vision and Pattern Recognition · Computer Science 2021-06-30 Sampanna Yashwant Kahu , William A. Ingram , Edward A. Fox , Jian Wu

Electronic theses and dissertations (ETDs) have been proposed, advocated, and generated for more than 25 years. Although ETDs are hosted by commercial or institutional digital library repositories, they are still an understudied type of…

Computer Vision and Pattern Recognition · Computer Science 2023-11-09 Muntabir Hasan Choudhury , Lamia Salsabil , William A. Ingram , Edward A. Fox , Jian Wu

Funding acknowledgments in scholarly publications provide large-scale trace data on organizations that support scientific research. We present a dataset for linking global science funding organizations to research publications by…

Digital Libraries · Computer Science 2026-03-26 Jacob Aarup Dalsgaard , Filipi Nascimento Silva , Jin AI

Dataset distillation (DD) condenses large datasets into compact yet informative substitutes, preserving performance comparable to the original dataset while reducing storage, transmission costs, and computational consumption. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Yawen Zou , Guang Li , Duo Su , Zi Wang , Jun Yu , Chao Zhang

In this paper, we propose a method to automatically classify AI-related documents from large-scale literature databases, leading to the creation of an AI-related literature dataset, named DeepDiveAI. The dataset construction approach…

Artificial Intelligence · Computer Science 2025-04-23 Zhou Xiaochen , Liang Xingzhou , Zou Hui , Lu Yi , Qu Jingjing

The volume of academic paper submissions and publications is growing at an ever increasing rate. While this flood of research promises progress in various fields, the sheer volume of output inherently increases the amount of noise. We…

Information Retrieval · Computer Science 2020-05-22 Marko Stamenovic , Jeibo Luo

Understanding how two datasets differ can help us determine whether one dataset under-represents certain sub-populations, and provides insights into how well models will generalize across datasets. Representative points selected by a…

Methodology · Statistics 2021-08-06 Sinead A. Williamson , Jette Henderson

Keeping track of all relevant recent publications and experimental results for a research area is a challenging task. Prior work has demonstrated the efficacy of information extraction models in various scientific areas. Recently, several…

Computation and Language · Computer Science 2023-10-25 Timo Pierre Schrader , Matteo Finco , Stefan Grünewald , Felix Hildebrand , Annemarie Friedrich

Scientific article summarization is challenging: large, annotated corpora are not available, and the summary should ideally include the article's impacts on research community. This paper provides novel solutions to these two challenges. We…

Computation and Language · Computer Science 2019-09-17 Michihiro Yasunaga , Jungo Kasai , Rui Zhang , Alexander R. Fabbri , Irene Li , Dan Friedman , Dragomir R. Radev

A main challenge of data-driven sciences is how to make maximal use of the progressively expanding databases of experimental datasets in order to keep research cumulative. We introduce the idea of a modeling-based dataset retrieval engine…

Quantitative Methods · Quantitative Biology 2015-06-19 Ali Faisal , Jaakko Peltonen , Elisabeth Georgii , Johan Rung , Samuel Kaski

Citation recommendation is the task of finding appropriate citations based on a given piece of text. The proposed datasets for this task consist mainly of several scientific fields, lacking some core ones, such as law. Furthermore, citation…

Information Retrieval · Computer Science 2023-11-13 Doğukan Arslan , Saadet Sena Erdoğan , Gülşen Eryiğit

In scientific research, the ability to effectively retrieve relevant documents based on complex, multifaceted queries is critical. Existing evaluation datasets for this task are limited, primarily due to the high cost and effort required to…

Information Retrieval · Computer Science 2023-10-31 Jianyou Wang , Kaicheng Wang , Xiaoyue Wang , Prudhviraj Naidu , Leon Bergen , Ramamohan Paturi

Data quality is crucial for training accurate, unbiased, and trustworthy machine learning models as well as for their correct evaluation. Recent works, however, have shown that even popular datasets used to train and evaluate…

Computation and Language · Computer Science 2024-03-12 Jan-Christoph Klie , Richard Eckart de Castilho , Iryna Gurevych

Deep learning technology has developed unprecedentedly in the last decade and has become the primary choice in many application domains. This progress is mainly attributed to a systematic collaboration in which rapidly growing computing…

Machine Learning · Computer Science 2023-12-27 Shiye Lei , Dacheng Tao

Tools to explore scientific literature are essential for scientists, especially in biomedicine, where about a million new papers are published every year. Many such tools provide users the ability to search for specific entities (e.g.…

Computation and Language · Computer Science 2021-07-05 Sunil Mohan , Rico Angell , Nick Monath , Andrew McCallum

Scientific figure interpretation is a crucial capability for AI-driven scientific assistants built on advanced Large Vision Language Models. However, current datasets and benchmarks primarily focus on simple charts or other relatively…

Highly specific datasets of scientific literature are important for both research and education. However, it is difficult to build such datasets at scale. A common approach is to build these datasets reductively by applying topic modeling…

Information Retrieval · Computer Science 2023-09-20 Nicholas Solovyev , Ryan Barron , Manish Bhattarai , Maksim E. Eren , Kim O. Rasmussen , Boian S. Alexandrov

Knowledge about software used in scientific investigations is important for several reasons, for instance, to enable an understanding of provenance and methods involved in data handling. However, software is usually not formally cited, but…

Information Retrieval · Computer Science 2021-08-23 David Schindler , Felix Bensmann , Stefan Dietze , Frank Krüger

In this paper, we ask the research question of whether all the datasets in the benchmark are necessary. We approach this by first characterizing the distinguishability of datasets when comparing different systems. Experiments on 9 datasets…

Computation and Language · Computer Science 2022-05-05 Yang Xiao , Jinlan Fu , See-Kiong Ng , Pengfei Liu

We present a dataset of 833k paragraphs extracted from CC-BY licensed scientific publications, classified into four categories: acknowledgments, data mentions, software/code mentions, and clinical trial mentions. The paragraphs are…

Computation and Language · Computer Science 2025-10-28 Eric Jeangirard
‹ Prev 1 3 4 5 6 7 10 Next ›