English
Related papers

Related papers: STEREO: Scientific Text Reuse in Open Access Publi…

200 papers

Urban data support a wide range of applications across multiple disciplines. However, at the global scale, there is no unified platform for urban data discovery. As a result, researchers often have to manually search through websites or…

Information Retrieval · Computer Science 2026-04-21 Runwen You , Tong Xia , Jingzhi Wang , Jiankun Zhang , Tengyao Tu , Jinghua Piao , Yi Chang , Yong Li

Systematic literature reviews (SLRs) play an essential role in summarising, synthesising and validating scientific evidence. In recent years, there has been a growing interest in using machine learning techniques to automate the…

Computation and Language · Computer Science 2023-11-22 Wojciech Kusa , Oscar E. Mendoza , Matthias Samwald , Petr Knoth , Allan Hanbury

The ability to find data is central to the FAIR principles underlying research data stewardship. As with the ability to reuse data, efforts to ensure and enhance findability have historically focused on discoverability of data by other…

Digital Libraries · Computer Science 2026-02-02 Bryan M. Gee

This paper is devoted to the study of methods for information extraction (entity recognition and relation classification) from scientific texts on information technology. Scientific publications provide valuable information into…

Computation and Language · Computer Science 2020-12-29 Elena Bruches , Alexey Pauls , Tatiana Batura , Vladimir Isachenko

Dockerfiles are one of the most prevalent kinds of DevOps artifacts used in industry. Despite their prevalence, there is a lack of sophisticated semantics-aware static analysis of Dockerfiles. In this paper, we introduce a dataset of…

Software Engineering · Computer Science 2020-03-31 Jordan Henkel , Christian Bird , Shuvendu K. Lahiri , Thomas Reps

Detecting salient parts in text using natural language processing has been widely used to mitigate the effects of information overflow. Nevertheless, most of the datasets available for this task are derived mainly from academic…

Computation and Language · Computer Science 2024-03-26 Andrés García-Silva , Cristian Berrío , José Manuel Gómez-Pérez

Computational reproducibility is central to scientific credibility, yet verifying published results at scale remains costly. We develop an AI-assisted workflow for automated full-paper replication -- retrieving materials, reconstructing…

Econometrics · Economics 2026-03-27 Yiqing Xu , Leo Yang Yang

Considerable scientific work involves locating, analyzing, systematizing, and synthesizing other publications. Its results end up in a paper's "background" section or in standalone articles, which include meta-analyses and systematic…

Digital Libraries · Computer Science 2023-12-05 Diomidis Spinellis

The TERRE-ISTEX project aims to identify scientific research dealing with specific geographical territories areas based on heterogeneous digital content available in scientific papers. The project is divided into three main work packages:…

We address the task of explaining relationships between two scientific documents using natural language text. This task requires modeling the complex content of long technical documents, deducing a relationship between these documents, and…

Computation and Language · Computer Science 2021-08-16 Kelvin Luu , Xinyi Wu , Rik Koncel-Kedziorski , Kyle Lo , Isabel Cachola , Noah A. Smith

Scientific information expresses human understanding of nature. This knowledge is largely disseminated in different forms of text, including scientific papers, news articles, and discourse among people on social media. While important for…

Computation and Language · Computer Science 2025-07-01 Dustin Wright

Analysis of acknowledgments is particularly interesting as acknowledgments may give information not only about funding, but they are also able to reveal hidden contributions to authorship and the researcher's collaboration patterns, context…

Digital Libraries · Computer Science 2023-05-26 Nina Smirnova , Philipp Mayr

Large multimodal models trained on natural documents, which interleave images and text, outperform models trained on image-text pairs on various multimodal benchmarks. However, the datasets used to train these models have not been released,…

Since the beginning of COVID pandemic, there have been around 700000 scientific papers published on the subject. A human researcher cannot possibly get acquainted with such a huge text corpus -- and therefore developing AI-based tools to…

Computation and Language · Computer Science 2021-11-01 Dmitry Soshnikov , Vickie Soshnikova

Concerns about reproducibility in artificial intelligence (AI) have emerged, as researchers have reported unsuccessful attempts to directly reproduce published findings in the field. Replicability, the ability to affirm a finding using the…

Computer Vision and Pattern Recognition · Computer Science 2023-04-21 Kehinde Ajayi , Muntabir Hasan Choudhury , Sarah Rajtmajer , Jian Wu

Nowadays, science has been coming into a new paradigm, called data-intensive science. While current studies of the new phenomenon focused on building up infrastructure for this new paradigm, yet a few studies concern users of scientific…

Instrumentation and Methods for Astrophysics · Physics 2011-02-08 Jian Zhang , Chaomei Chen , Michael S. Vogeley

Knowledge about software used in scientific investigations is important for several reasons, for instance, to enable an understanding of provenance and methods involved in data handling. However, software is usually not formally cited, but…

Information Retrieval · Computer Science 2021-08-23 David Schindler , Felix Bensmann , Stefan Dietze , Frank Krüger

The World Wide Web is not only one of the most important platforms of communication and information at present, but also an area of growing interest for scientific research. This motivates a lot of work and projects that require large…

Computer Vision and Pattern Recognition · Computer Science 2021-05-18 Christian Mejia-Escobar , Miguel Cazorla , Ester Martinez-Martin

This article presents CRED-1, an open, reproducible domain-level credibility dataset combining two openly-licensed source lists (OpenSources.co and Iffy.news) with four computed enrichment signals: domain age (WHOIS/RDAP), web popularity…

Information Retrieval · Computer Science 2026-04-24 Alexander Loth , Martin Kappes , Marc-Oliver Pahl

Sequence-to-sequence models have recently gained the state of the art performance in summarization. However, not too many large-scale high-quality datasets are available and almost all the available ones are mainly news articles with…

Computation and Language · Computer Science 2018-10-23 Mahnaz Koupaee , William Yang Wang
‹ Prev 1 3 4 5 6 7 10 Next ›