English
Related papers

Related papers: WithdrarXiv: A Large-Scale Dataset for Retraction …

200 papers

The availability of a vast array of research papers in any area of study, necessitates the need of automated summarisation systems that can present the key research conducted and their corresponding findings. Scientific paper summarisation…

Computation and Language · Computer Science 2024-07-30 Grishma Sharma , Aditi Paretkar , Deepak Sharma

This paper is withdrawn due to some errors, which are corrected in arXiv:0912.0071v4 [cs.LG].

Cryptography and Security · Computer Science 2011-06-22 Kamalika Chaudhuri , Anand D. Sarwate

We report on the ongoing development of arXiv's HTML Papers offering, available on every new TeX/LaTeX submission since its initial release in 2023. The main highlights from 2025 and early 2026 are: (i) community-driven improvements to HTML…

Computation and Language · Computer Science 2026-05-26 Deyan Ginev , Brian Caruso , Bruce Miller , Jeff Sank , Jacob Weiskoff

Researchers produce thousands of scholarly documents containing valuable technical knowledge. The community faces the laborious task of reading these documents to identify, extract, and synthesize information. To automate information…

Computation and Language · Computer Science 2023-12-13 Tavish McDonald , Brian Tsan , Amar Saini , Juanita Ordonez , Luis Gutierrez , Phan Nguyen , Blake Mason , Brenda Ng

In this paper we introduce a new publicly available dataset for verification against textual sources, FEVER: Fact Extraction and VERification. It consists of 185,445 claims generated by altering sentences extracted from Wikipedia and…

Computation and Language · Computer Science 2018-12-19 James Thorne , Andreas Vlachos , Christos Christodoulopoulos , Arpit Mittal

In many fields of experimental science, papers that failed to replicate continue to be cited as a result of the poor discoverability of replication studies. As a first step to creating a system that automatically finds replication studies…

Digital Libraries · Computer Science 2023-11-28 Bob de Ruiter

This paper quantifies to which extent preprints in arXiv accelerate scholarly communication. The following subject fields were investigated up to the year 2012: High Energy Physics (HEP), Mathematics, Astrophysics, Quantitative Biology, and…

Digital Libraries · Computer Science 2013-06-21 Valeria Aman

Scientific literature has itself been the subject of much scientific study, for a variety of reasons: understanding how results are communicated, how ideas spread, and assessing the influence of areas or individuals. However, most prior…

Digital Libraries · Computer Science 2013-08-01 Graham Cormode , S. Muthukrishnan , Jinyun Yan

Scientists involved in scientific misconduct may face social stigmatization, leading to isolation and limited opportunities for collaboration. The reputation of every individual is reflected on the team, as the fraud attempted by any member…

Digital Libraries · Computer Science 2023-10-11 Kiran Sharma , Satyam Mukherjee

Since its creation in 1991, arXiv has become central to the diffusion of research in a number of fields. Combining data from the entirety of arXiv and the Web of Science (WoS), this paper investigates (a) the proportion of papers across all…

Digital Libraries · Computer Science 2018-12-19 Vincent Lariviere , Cassidy R. Sugimoto , Benoit Macaluso , Stasa Milojevic , Blaise Cronin , Mike Thelwall

Effective data management and sharing are critical success factors in industry-academia collaboration. This paper explores the motivations and lessons learned from publishing open data sets in such collaborations. Through a survey of…

Software Engineering · Computer Science 2025-01-28 Per Erik Strandberg , Philipp Peterseil , Julian Karoliny , Johanna Kallio , Johannes Peltola

This study presents OpenExtract, an open-source pipeline for automated data extraction in large-scale systematic literature reviews. The pipeline queries large language models (LLMs) to predict data entries based on relevant sections of…

We study the data deletion problem for convex models. By leveraging techniques from convex optimization and reservoir sampling, we give the first data deletion algorithms that are able to handle an arbitrarily long sequence of adversarial…

Machine Learning · Statistics 2020-07-07 Seth Neel , Aaron Roth , Saeed Sharifi-Malvajerdi

Large language models (LLMs) demonstrate strong capabilities in reasoning and question answering, yet their tendency to generate factually incorrect content remains a critical challenge. This study evaluates proprietary and open-source LLMs…

Information Retrieval · Computer Science 2025-08-08 Ning Li , Jingran Zhang , Justin Cui

This paper presents a hierarchical classification system that automatically categorizes a scholarly publication using its abstract into a three-tier hierarchical label set (discipline, field, subfield) in a multi-class setting. This system…

Digital Libraries · Computer Science 2024-07-26 Susie Xi Rao , Peter H. Egger , Ce Zhang

This article statistically analyses how the citation impact of articles deposited in the Condensed Matter section of the preprint server ArXiv (hosted by Cornell University), and subsequently published in a scientific journal, compares to…

Digital Libraries · Computer Science 2007-05-23 Henk F. Moed

Background. Reproducibility is essential to the scientific method, but reproduction is often a laborious task. Recent works have attempted to automate this process and relieve researchers of this workload. However, due to varying…

Computers and Society · Computer Science 2026-01-09 Thijs Snelleman , Peter Lundestad Lawrence , Holger H. Hoos , Odd Erik Gundersen

An abstractive snippet is an originally created piece of text to summarize a web page on a search engine results page. Compared to the conventional extractive snippets, which are generated by extracting phrases and sentences verbatim from a…

Information Retrieval · Computer Science 2020-03-17 Wei-Fan Chen , Shahbaz Syed , Benno Stein , Matthias Hagen , Martin Potthast

The goal is to study the prevalence of citajacked papers: papers in authentic scientific journals citing hijacked journals, in academic literature. A Citejacked detector was designed as a part of the Problematic Paper Screener…

Digital Libraries · Computer Science 2022-09-13 Anna Abalkina , Guillaume Cabanac , Cyril Labbé , Alexander Magazinov

The growing impact of preprint servers enables the rapid sharing of time-sensitive research. Likewise, it is becoming increasingly difficult to distinguish high-quality, peer-reviewed research from preprints. Although preprints are often…

Digital Libraries · Computer Science 2023-10-05 Louise Bloch , Johannes Rückert , Christoph M. Friedrich