English
Related papers

Related papers: STEREO: Scientific Text Reuse in Open Access Publi…

200 papers

We present a dataset of 833k paragraphs extracted from CC-BY licensed scientific publications, classified into four categories: acknowledgments, data mentions, software/code mentions, and clinical trial mentions. The paragraphs are…

Computation and Language · Computer Science 2025-10-28 Eric Jeangirard

With the Open Science approach becoming important for research, the evolution towards open scientific-paper reviews is making an impact on the scientific community. However, there is a lack of publicly available resources for conducting…

Digital Libraries · Computer Science 2023-12-11 Jaroslaw Szumega , Lamine Bougueroua , Blerina Gkotse , Pierre Jouvelot , Federico Ravotti

We present SemOpenAlex, an extensive RDF knowledge graph that contains over 26 billion triples about scientific publications and their associated entities, such as authors, institutions, journals, and concepts. SemOpenAlex is licensed under…

Digital Libraries · Computer Science 2023-08-08 Michael Färber , David Lamprecht , Johan Krause , Linn Aung , Peter Haase

Medical imaging papers often focus on methodology, but the quality of the algorithms and the validity of the conclusions are highly dependent on the datasets used. As creating datasets requires a lot of effort, researchers often use…

In the implementation and use of research information systems (RIS) in scientific institutions, text data mining and semantic technologies are a key technology for the meaningful use of large amounts of data. It is not the collection of…

Digital Libraries · Computer Science 2018-12-12 Otmane Azeroual , Gunter Saake , Mohammad Abuosba , Joachim Schöpfel

At the foundation of scientific evaluation is the labor-intensive process of peer review. This critical task requires participants to consume vast amounts of highly technical text. Prior work has annotated different aspects of review…

We report evidence of a new set of sneaked references discovered in the scientific literature. Sneaked references are references registered in the metadata of publications without being listed in reference section or in the full text of the…

The Reception Reader is a web tool for studying text reuse in the Early English Books Online (EEBO-TCP) and Eighteenth Century Collections Online (ECCO) data. Users can: 1) explore a visual overview of the reception of a work, or its…

Digital Libraries · Computer Science 2023-04-20 David Rosson , Eetu Mäkelä , Ville Vaara , Ananth Mahadevan , Yann Ryan , Mikko Tolonen

Open Information Extraction (OIE) is the task of the unsupervised creation of structured information from text. OIE is often used as a starting point for a number of downstream tasks including knowledge base construction, relation…

Computation and Language · Computer Science 2018-08-23 Paul Groth , Michael Lauruhn , Antony Scerri , Ron Daniel

"Open access" has become a central theme of journal reform in academic publishing. In this article, I examine the relationship between open access publishing and an important infrastructural element of a modern research enterprise,…

Computers and Society · Computer Science 2018-04-25 Gopal P. Sarma

Publishing research data aims to improve the transparency of research results and facilitate the reuse of datasets. In both cases, referencing the datasets that were used is recommended. Research data repositories can support data…

Digital Libraries · Computer Science 2025-05-14 Dorothea Strecker , Kerstin Soltau , Felix Bach

Without sufficient information about research data practices occurring in a particular research organisation, there is a risk of mismatching research data service efforts with the needs of its researchers. This study describes how data…

Digital Libraries · Computer Science 2022-05-12 Antti Mikael Rousi

Systematic reviews are the standard method for synthesizing scientific evidence, but their creation requires substantial manual effort, particularly during retrieval and screening. While recent work has explored automating these steps,…

Information Retrieval · Computer Science 2026-04-21 Pierre Achkar , Tim Gollub amd Martin Potthast

Complex scientific codes and the datasets they generate are in need of a sophisticated categorization environment that allows the community to store, search, and enhance metadata in an open, dynamic system. Currently, data is often…

Digital Libraries · Computer Science 2012-03-20 Eric L. Seidel

With the advent of open source software, a veritable treasure trove of previously proprietary software development data was made available. This opened the field of empirical software engineering research to anyone in academia. Data that is…

Software Engineering · Computer Science 2022-04-19 Adam Tutko , Austin Z. Henley , Audris Mockus

This paper proposes OCR++, an open-source framework designed for a variety of information extraction tasks from scholarly articles including metadata (title, author names, affiliation and e-mail), structure (section headings and body text,…

Extracting structured and grounded fact triples from raw text is a fundamental task in Information Extraction (IE). Existing IE datasets are typically collected from Wikipedia articles, using hyperlinks to link entities to the Wikidata…

Computation and Language · Computer Science 2023-06-16 Chenxi Whitehouse , Clara Vania , Alham Fikri Aji , Christos Christodoulopoulos , Andrea Pierleoni

Information Extraction is a well-researched area of Natural Language Processing with applications in web search and question answering concerned with identifying entities and relationships between them as expressed in a given context,…

Information Retrieval · Computer Science 2020-11-17 Erin Macdonald , Denilson Barbosa

Proper citation is of great importance in academic writing for it enables knowledge accumulation and maintains academic integrity. However, citing properly is not an easy task. For published scientific entities, the ever-growing academic…

Digital Libraries · Computer Science 2022-10-20 Jialiang Lin , Yao Yu , Jiaxin Song , Xiaodong Shi

Materials science literature contains millions of materials synthesis procedures described in unstructured natural language text. Large-scale analysis of these synthesis procedures would facilitate deeper scientific understanding of…

Computation and Language · Computer Science 2019-07-16 Sheshera Mysore , Zach Jensen , Edward Kim , Kevin Huang , Haw-Shiuan Chang , Emma Strubell , Jeffrey Flanigan , Andrew McCallum , Elsa Olivetti