English
Related papers

Related papers: WithdrarXiv: A Large-Scale Dataset for Retraction …

200 papers

The automated extraction of structured questions from paper-based mathematics exams is fundamental to intelligent education, yet remains challenging in real-world settings due to severe visual noise. Existing benchmarks mainly focus on…

Computer Vision and Pattern Recognition · Computer Science 2026-01-16 Chenyue Zhou , Jiayi Tuo , Shitong Qin , Wei Dai , Mingxuan Wang , Ziwei Zhao , Duoyang Li , Shiyang Su , Yanxi Lu , Yanbiao Ma

Datasets are critical for scientific research, playing an important role in replication, reproducibility, and efficiency. Researchers have recently shown that datasets are becoming more important for science to function properly, even…

Computation and Language · Computer Science 2024-05-24 Tong Zeng , Daniel Acuna

Within the past few decades we have witnessed digital revolution, which moved scholarly communication to electronic media and also resulted in a substantial increase in its volume. Nowadays keeping track with the latest scientific…

Digital Libraries · Computer Science 2017-10-30 Dominika Tkaczyk

In this work, we study how URL extraction results depend on input format. We compiled a pilot dataset by extracting URLs from 10 arXiv papers and used the same heuristic method to extract URLs from four formats derived from the PDF files or…

Digital Libraries · Computer Science 2025-09-08 Rochana R. Obadage , Lamia Salsabil , Sawood Alam , Bipasha Banarjee , William A. Ingram , Edward A. Fox , Jian Wu

In October 2023, arXiv made HTML formatted papers available to readers. This was the exciting outcome of over a year of accessibility research and development with the scientific community. Currently, only 2.4% of research outputs meet…

Digital Libraries · Computer Science 2024-02-15 Charles Frankston , Jonathan Godfrey , Shamsi Brinn , Alison Hofer , Mark Nazzaro

Peer reviewing is a central component in the scientific publishing process. We present the first public dataset of scientific peer reviews available for research purposes (PeerRead v1) providing an opportunity to study this important…

Computation and Language · Computer Science 2018-04-26 Dongyeop Kang , Waleed Ammar , Bhavana Dalvi , Madeleine van Zuylen , Sebastian Kohlmeier , Eduard Hovy , Roy Schwartz

Keyphrase extraction is the task of extracting a small set of phrases that best describe a document. Most existing benchmark datasets for the task typically have limited numbers of annotated documents, making it challenging to train…

Computation and Language · Computer Science 2020-10-26 Tuan Manh Lai , Trung Bui , Doo Soon Kim , Quan Hung Tran

Scientific abstracts are increasingly used as primary data in computational metascience research, yet the quality of these abstracts in widely used bibliographic databases has not been systematically examined. We assess the integrity of…

Digital Libraries · Computer Science 2026-05-20 Seorin Kim , Vincent Holst , Vincent Ginis

Computational reproducibility is central to scientific credibility, yet verifying published results at scale remains costly. We develop an AI-assisted workflow for automated full-paper replication -- retrieving materials, reconstructing…

Econometrics · Economics 2026-03-27 Yiqing Xu , Leo Yang Yang

Retractions undermine the reliability of scientific literature and the foundation of future research. Analyzing collaboration networks in retracted papers can identify risk factors, such as recurring co-authors or institutions. This study…

Information Retrieval · Computer Science 2024-11-27 Kiran Sharma , Aanchal Sharma , Jazlyn Jose , Vansh Saini , Raghavraj Sobti , Ziya Uddin

As researchers use computational methods to study complex social behaviors at scale, the validity of this computational social science depends on the integrity of the data. On July 2, 2015, Jason Baumgartner published a dataset advertised…

Social and Information Networks · Computer Science 2018-09-05 Devin Gaffney , J. Nathan Matias

Writing a scientific article is a challenging task as it is a highly codified and specific genre, consequently proficiency in written communication is essential for effectively conveying research findings and ideas. In this article, we…

Computation and Language · Computer Science 2025-01-10 Leane Jourdan , Florian Boudin , Nicolas Hernandez , Richard Dufour

This paper presents a procedure to retrieve subsets of relevant documents from large text collections for Content Analysis, e.g. in social sciences. Document retrieval for this purpose needs to take account of the fact that analysts often…

Information Retrieval · Computer Science 2017-07-12 Gregor Wiedemann , Andreas Niekler

High-throughput screening techniques are commonly used to obtain large quantities of data in many fields of biology. It is well known that artifacts arising from variability in the technical execution of different experimental batches…

Non-textual components such as charts, diagrams and tables provide key information in many scientific documents, but the lack of large labeled datasets has impeded the development of data-driven methods for scientific figure extraction. In…

Digital Libraries · Computer Science 2018-06-01 Noah Siegel , Nicholas Lourie , Russell Power , Waleed Ammar

Extracting information from academic PDF documents is crucial for numerous indexing, retrieval, and analysis use cases. Choosing the best tool to extract specific content elements is difficult because many, technically diverse tools are…

Information Retrieval · Computer Science 2023-03-20 Norman Meuschke , Apurva Jagdale , Timo Spinde , Jelena Mitrović , Bela Gipp

Comprehensively retrieving diverse documents is crucial to address queries that admit a wide range of valid answers. We introduce retrieve-verify-retrieve (RVR), a multi-round retrieval framework designed to maximize answer coverage.…

Computation and Language · Computer Science 2026-02-23 Deniz Qian , Hung-Ting Chen , Eunsol Choi

The LongEval lab focuses on the evaluation of information retrieval systems over time. Two datasets are provided that capture evolving search scenarios with changing documents, queries, and relevance assessments. Systems are assessed from a…

We respond to Holst et al.'s critique that the decline in scientific disruptiveness documented in Park et al. (Nature, 2023) is an artifact of including works with zero backward citations. Using their advocated dataset, metric, and…

Social and Information Networks · Computer Science 2026-03-23 Michael Park , Erin Leahey , Russell J. Funk

Extracting hypotheses and their supporting statistical evidence from full-text scientific articles is central to the synthesis of empirical findings, but remains difficult due to document length and the distribution of scientific arguments…

Computation and Language · Computer Science 2026-03-24 Sai Koneru , Jian Wu , Sarah Rajtmajer