English

A Semantically Enriched Dataset based on Biomedical NER for the COVID19 Open Research Dataset Challenge

Digital Libraries 2020-05-19 v1

Abstract

Research into COVID-19 is a big challenge and highly relevant at the moment. New tools are required to assist medical experts in their research with relevant and valuable information. The COVID-19 Open Research Dataset Challenge (CORD-19) is a "call to action" for computer scientists to develop these innovative tools. Many of these applications are empowered by entity information, i. e. knowing which entities are used within a sentence. For this paper, we have developed a pipeline upon the latest Named Entity Recognition tools for Chemicals, Diseases, Genes and Species. We apply our pipeline to the COVID-19 research challenge and share the resulting entity mentions with the community.

Keywords

Cite

@article{arxiv.2005.08823,
  title  = {A Semantically Enriched Dataset based on Biomedical NER for the COVID19 Open Research Dataset Challenge},
  author = {Hermann Kroll and Jan Pirklbauer and Johannes Ruthmann and Wolf-Tilo Balke},
  journal= {arXiv preprint arXiv:2005.08823},
  year   = {2020}
}