English
Related papers

Related papers: STEREO: Scientific Text Reuse in Open Access Publi…

200 papers

In this paper, we investigate the retrievability of datasets and publications in a real-life Digital Library (DL). The measure of retrievability was originally developed to quantify the influence that a retrieval system has on the access to…

Information Retrieval · Computer Science 2022-07-22 Dwaipayan Roy , Zeljko Carevic , Philipp Mayr

We present MESA-Web, a cloud resource with an online interface to the Modules for Experiments in Stellar Astrophysics (MESA) software instrument. MESA-Web allows learners to evolve stellar models without the need to download and install…

Solar and Stellar Astrophysics · Physics 2023-09-29 Carl E. Fields , Richard H. D. Townsend , A. L. Dotter , Michael Zingale , F. X. Timmes

We present Lit2Vec, a reproducible workflow for constructing and validating a chemistry corpus from the Semantic Scholar Open Research Corpus using conservative, metadata-based license screening. Using this workflow, we assembled an…

Databases · Computer Science 2026-04-15 Mahmoud Amiri , Jamile Mohammad Jafari , Sara Mostafapour , Thomas Bocklitz

Scientific data collected at ESO's observatories are freely and openly accessible online through the ESO Science Archive Facility. In addition to the raw data straight out of the instruments, the ESO Science Archive also contains four…

Accurately assigning the document type of review articles in citation index databases like Web of Science(WoS) and Scopus is important. This study aims to investigate the document type assignation of review articles in web of Science,…

Digital Libraries · Computer Science 2023-12-01 Manman Zhu , Xinyue Lu , Fuyou Chen , Liying Yang , Zhesi Shen

We report results on benchmarking Open Information Extraction (OIE) systems using RelVis, a toolkit for benchmarking Open Information Extraction systems. Our comprehensive benchmark contains three data sets from the news domain and one data…

Computation and Language · Computer Science 2017-07-25 Rudolf Schneider , Tom Oberhauser , Tobias Klatt , Felix A. Gers , Alexander Löser

The growth rate of the number of scientific publications is constantly increasing, creating important challenges in the identification of valuable research and in various scholarly data management applications, in general. In this context,…

A major challenge in fine-tuning deep learning models for automatic summarization is the need for large domain specific datasets. One of the barriers to curating such data from resources like online publications is navigating the license…

Information Retrieval · Computer Science 2020-12-01 Amanpreet Singh , Niranjan Balasubramanian

We develop and evaluate multilingual scientific documents similarity measurement models in this work. Such models can be used to find related works in different languages, which can help multilingual researchers find and explore papers more…

Computation and Language · Computer Science 2023-09-20 Yang Gao , Ji Ma , Ivan Korotkov , Keith Hall , Dana Alon , Don Metzler

Many commercial and open-source models claim to detect machine-generated text with extremely high accuracy (99% or more). However, very few of these detectors are evaluated on shared benchmark datasets and even when they are, the datasets…

Computation and Language · Computer Science 2024-06-11 Liam Dugan , Alyssa Hwang , Filip Trhlik , Josh Magnus Ludan , Andrew Zhu , Hainiu Xu , Daphne Ippolito , Chris Callison-Burch

Effective data management and sharing are critical success factors in industry-academia collaboration. This paper explores the motivations and lessons learned from publishing open data sets in such collaborations. Through a survey of…

Software Engineering · Computer Science 2025-01-28 Per Erik Strandberg , Philipp Peterseil , Julian Karoliny , Johanna Kallio , Johannes Peltola

Open-source scientific software is effectively closed to modification by its complexity. With recent advances in technology, an agentic AI team led by a single human can now rapidly and robustly modify large codebases and re-open science to…

Software Engineering · Computer Science 2025-12-16 Ling-Hong Hung , Ka Yee Yeung

The accelerating pace of research on autoregressive generative models has produced thousands of papers, making manual literature surveys and reproduction studies increasingly impractical. We present a fully open-source, reproducible…

Information Retrieval · Computer Science 2025-08-07 Faruk Alpay , Bugra Kilictas , Hamdi Alakkad

Using three years of the Journal Citation Reports (2011, 2012, and 2013), indicators of transitions in 2012 (between 2011 and 2013) are studied using methodologies based on entropy statistics. Changes can be indicated at the level of…

Digital Libraries · Computer Science 2015-08-18 Loet Leydesdorff , Wouter de Nooy

We introduce a large-scale dataset of the complete texts of free/open source software (FOSS) license variants. To assemble it we have collected from the Software Heritage archive-the largest publicly available archive of FOSS source code…

Software Engineering · Computer Science 2022-04-04 Stefano Zacchiroli

Recent advances in methods and techniques enable us to develop an interactive overlay to the global map of science based on aggregated citation relations among the 9,162 journals contained in the Science Citation Index and Social Science…

Digital Libraries · Computer Science 2011-05-17 Loet Leydesdorff , Ismael Rafols

With social media datasets being increasingly shared by researchers, it also presents the caveat that those datasets are not always completely replicable. Having to adhere to requirements of platforms like Twitter, researchers cannot…

Digital Libraries · Computer Science 2018-03-08 Arkaitz Zubiaga

An increasing number of researchers support reproducibility by including pointers to and descriptions of datasets, software and methods in their publications. However, scientific articles may be ambiguous, incomplete and difficult to…

Communicating complex system designs or scientific processes through text alone is inefficient and prone to ambiguity. A system that automatically generates scientific architecture diagrams from text with high semantic fidelity can be…

Computation and Language · Computer Science 2026-04-17 Shivank Garg , Sankalp Mittal , Manish Gupta

Independent of established data centers, and partly for my own research, since 1989 I have been collecting the tabular data from over 2600 articles concerned with radio sources and extragalactic objects in general. Optical character…

Instrumentation and Methods for Astrophysics · Physics 2009-11-13 Heinz Andernach
‹ Prev 1 4 5 6 7 8 10 Next ›