English
Related papers

Related papers: The CALBC RDF Triple Store: retrieval over large l…

200 papers

We present SemOpenAlex, an extensive RDF knowledge graph that contains over 26 billion triples about scientific publications and their associated entities, such as authors, institutions, journals, and concepts. SemOpenAlex is licensed under…

Digital Libraries · Computer Science 2023-08-08 Michael Färber , David Lamprecht , Johan Krause , Linn Aung , Peter Haase

We present SemRepo, an RDF knowledge graph comprising over 81 million triples describing nearly 200,000 GitHub repositories associated with scientific research. SemRepo captures repository-level metadata, such as contributors, issues, and…

Digital Libraries · Computer Science 2026-05-14 Abdul Rafay , Yuni Susanti , David Lamprecht , Michael Färber

Solutions to the classic problems of dealing with heterogeneous data and making entire collections interoperable while ensuring that any annotation, which includes the recognition-and-reward system of scientific publishing, need to fit into…

Digital Libraries · Computer Science 2015-03-17 Paul Boekschoten , Kees Burger , Barend Mons , Christine Chichester

In recent years, the increased need to house and process large volumes of data has prompted the need for distributed storage and querying systems. The growth of machine-readable RDF triples has prompted both industry and academia to develop…

Databases · Computer Science 2016-01-11 Albert Haque

To date, there are no effective treatments for most neurodegenerative diseases. Knowledge graphs can provide comprehensive and semantic representation for heterogeneous data, and have been successfully leveraged in many biomedical…

Artificial Intelligence · Computer Science 2022-11-30 Yi Nian , Xinyue Hu , Rui Zhang , Jingna Feng , Jingcheng Du , Fang Li , Yong Chen , Cui Tao

Existing scholarly information extraction (SIE) datasets focus on scientific papers and overlook implementation-level details in code repositories. README files describe datasets, source code, and other implementation-level artifacts,…

Computation and Language · Computer Science 2026-03-09 Genet Asefa Gesese , Zongxiong Chen , Shufan Jiang , Mary Ann Tan , Zhaotai Liu , Sonja Schimmler , Harald Sack

From 2012 to 2015 together with other Linked Data community members and experts from the social, behavioral, and economic sciences (SBE), we developed diverse vocabularies to represent SBE metadata and tabular data in RDF. The DDI-RDF…

Digital Libraries · Computer Science 2015-09-16 Thomas Hartmann , Benjamin Zapilko , Joachim Wackerow , Kai Eckert

Scholarly data are largely fragmented across siloed databases with divergent metadata and missing linkages among them. We present the Science Data Lake, a locally-deployable infrastructure built on DuckDB and simple Parquet files that…

Digital Libraries · Computer Science 2026-03-04 Jonas Wilinski

Following the global COVID-19 pandemic, the number of scientific papers studying the virus has grown massively, leading to increased interest in automated literate review. We present a clinical text mining system that improves on previous…

Computation and Language · Computer Science 2020-12-09 Veysel Kocaman , David Talby

We present Belebele, a multiple-choice machine reading comprehension (MRC) dataset spanning 122 language variants. Significantly expanding the language coverage of natural language understanding (NLU) benchmarks, this dataset enables the…

Objective: To develop a corpus annotated for diet-microbiome associations from the biomedical literature and train natural language processing (NLP) models to identify these associations, thereby improving the understanding of their role in…

Computation and Language · Computer Science 2025-04-01 Gibong Hong , Veronica Hindle , Nadine M. Veasley , Hannah D. Holscher , Halil Kilicoglu

The Russian Drug Reaction Corpus (RuDReC) is a new partially annotated corpus of consumer reviews in Russian about pharmaceutical products for the detection of health-related named entities and the effectiveness of pharmaceutical products.…

Computation and Language · Computer Science 2023-11-21 Elena Tutubalina , Ilseyar Alimova , Zulfat Miftahutdinov , Andrey Sakhovskiy , Valentin Malykh , Sergey Nikolenko

Objective: To discover candidate drugs to repurpose for COVID-19 using literature-derived knowledge and knowledge graph completion methods. Methods: We propose a novel, integrative, and neural network-based literature-based discovery (LBD)…

Computation and Language · Computer Science 2021-02-10 Rui Zhang , Dimitar Hristovski , Dalton Schutte , Andrej Kastrin , Marcelo Fiszman , Halil Kilicoglu

The Resource Description Framework (RDF) is continuing to grow outside the bounds of its initial function as a metadata framework and into the domain of general-purpose data modeling. This expansion has been facilitated by the continued…

Artificial Intelligence · Computer Science 2008-07-25 Marko A. Rodriguez

Proper citation is of great importance in academic writing for it enables knowledge accumulation and maintains academic integrity. However, citing properly is not an easy task. For published scientific entities, the ever-growing academic…

Digital Libraries · Computer Science 2022-10-20 Jialiang Lin , Yao Yu , Jiaxin Song , Xiaodong Shi

Capturing the semantics of related biological concepts, such as genes and mutations, is of significant importance to many research tasks in computational biology such as protein-protein interaction detection, gene-drug association…

Computation and Language · Computer Science 2020-07-01 Qingyu Chen , Kyubum Lee , Shankai Yan , Sun Kim , Chih-Hsuan Wei , Zhiyong Lu

Research resources (RRs) such as data, software, and tools are essential pillars of scientific research. The field of biomedicine, a critical scientific discipline, is witnessing a surge in research publications resulting in the…

Digital Libraries · Computer Science 2024-09-24 Li Zhang , Mengting Sun , Chong Jiang , Haihua Chen

The processing of entities in natural language is essential to many medical NLP systems. Unfortunately, existing datasets vastly under-represent the entities required to model public health relevant texts such as health advice often found…

Computation and Language · Computer Science 2022-10-10 Joseph Gatto , Parker Seegmiller , Garrett Johnston , Sarah M. Preum

We are faced with an unprecedented production in scholarly publications worldwide. Stakeholders in the digital libraries posit that the document-based publishing paradigm has reached the limits of adequacy. Instead, structured,…

Computation and Language · Computer Science 2022-05-25 Jennifer D'Souza

High throughput extraction and structured labeling of data from academic articles is critical to enable downstream machine learning applications and secondary analyses. We have embedded multimodal data curation into the academic publishing…

Computation and Language · Computer Science 2024-09-26 Jorge Abreu-Vicente , Hannah Sonntag , Thomas Eidens , Cassie S. Mitchell , Thomas Lemberger
‹ Prev 1 2 3 10 Next ›