English
Related papers

Related papers: MedMentions: A Large Biomedical Corpus Annotated w…

200 papers

The number of biomedical literature on new biomedical concepts is rapidly increasing, which necessitates a reliable biomedical named entity recognition (BioNER) model for identifying new and unseen entity mentions. However, it is…

Computation and Language · Computer Science 2022-03-15 Hyunjae Kim , Jaewoo Kang

The growth rate in the amount of biomedical documents is staggering. Unlocking information trapped in these documents can enable researchers and practitioners to operate confidently in the information world. Biomedical NER, the task of…

Computation and Language · Computer Science 2021-06-24 Xiang Dai

Causal relation extraction (CRE) is central to biomedical text mining, but current resources often conflate causal relations with broader associations, restrict annotation to sentence-level examples, or focus mainly on explicit causal cues.…

Computation and Language · Computer Science 2026-05-28 Ifeoluwa Kunle-John , Josiah Paul , Oluwatosin Agbaakin , Peter Aina , Ikenna Odezuligbo , Sydney Anuyah

Biomedical concept normalization links concept mentions in texts to a semantically equivalent concept in a biomedical knowledge base. This task is challenging as concepts can have different expressions in natural languages, e.g.…

Computation and Language · Computer Science 2018-07-10 Roland Roller , Madeleine Kittner , Dirk Weissenborn , Ulf Leser

Large language models (LLMs) like ChatGPT can generate and revise text with human-level performance. These models come with clear limitations: they can produce inaccurate information, reinforce existing biases, and be easily misused. Yet,…

Computation and Language · Computer Science 2025-07-04 Dmitry Kobak , Rita González-Márquez , Emőke-Ágnes Horvát , Jan Lause

Named Entity Recognition (NER) or the extraction of concepts from clinical text is the task of identifying entities in text and slotting them into categories such as problems, treatments, tests, clinical departments, occurrences (such as…

Computation and Language · Computer Science 2022-08-31 Namrata Nath , Sang-Heon Lee , Ivan Lee

Objective: Social media-based public health research is crucial for epidemic surveillance, but most studies identify relevant corpora with keyword-matching. This study develops a system to streamline the process of curating colloquial…

Computation and Language · Computer Science 2024-03-19 Yining Hua , Jiageng Wu , Shixu Lin , Minghui Li , Yujie Zhang , Dinah Foer , Siwen Wang , Peilin Zhou , Jie Yang , Li Zhou

Distributed representations of medical concepts have been used to support downstream clinical tasks recently. Electronic Health Records (EHR) capture different aspects of patients' hospital encounters and serve as a rich source for…

Computation and Language · Computer Science 2020-01-07 Shaika Chowdhury , Chenwei Zhang , Philip S. Yu , Yuan Luo

We present MedConceptsQA, a dedicated open source benchmark for medical concepts question answering. The benchmark comprises of questions of various medical concepts across different vocabularies: diagnoses, procedures, and drugs. The…

Computation and Language · Computer Science 2024-05-15 Ofir Ben Shoham , Nadav Rappoport

Clinical information extraction, which involves structuring clinical concepts from unstructured medical text, remains a challenging problem that could benefit from the inclusion of tabular background information available in electronic…

Artificial Intelligence · Computer Science 2025-12-10 Paloma Rabaey , Stefan Heytens , Thomas Demeester

We present a system that constructs and maintains an up-to-date co-occurrence network of medical concepts based on continuously mining the latest biomedical literature. Users can explore this network visually via a concise online interface…

Information Retrieval · Computer Science 2015-03-20 Alexei Yavlinsky

The number of clinical citations received from clinical guidelines or clinical trials has been considered as one of the most appropriate indicators for quantifying the clinical impact of biomedical papers. Therefore, the early prediction of…

Computation and Language · Computer Science 2022-10-24 Xin Li , Xuli Tang , Qikai Cheng

Current medical language model (LM) benchmarks often over-simplify the complexities of day-to-day clinical practice tasks and instead rely on evaluating LMs on multiple-choice board exam questions. In psychiatry especially, these challenges…

We describe the CZ Software Mentions dataset, a new dataset of software mentions in biomedical papers. Plain-text software mentions are extracted with a trained SciBERT model from several sources: the NIH PubMed Central collection and from…

Digital Libraries · Computer Science 2022-09-29 Ana-Maria Istrate , Donghui Li , Dario Taraborelli , Michaela Torkar , Boris Veytsman , Ivana Williams

We report characteristics of in-text citations in over five million full text articles from two large databases - the PubMed Central Open Access subset and Elsevier journals - as functions of time, textual progression, and scientific field.…

Digital Libraries · Computer Science 2017-10-10 Kevin W. Boyack , Nees Jan van Eck , Giovanni Colavizza , Ludo Waltman

Disease name recognition and normalization, which is generally called biomedical entity linking, is a fundamental process in biomedical text mining. Recently, neural joint learning of both tasks has been proposed to utilize the mutual…

Computation and Language · Computer Science 2021-04-22 Shogo Ujiie , Hayate Iso , Shuntaro Yada , Shoko Wakamiya , Eiji Aramaki

The Medical Subject Headings (MeSH), one of the main knowledge organization systems in the biomedical domain, continuously evolves to reflect the latest scientific discoveries in health and life sciences. Previous research has focused on…

Social and Information Networks · Computer Science 2025-10-09 Jenny Copara , Nona Naderi , Gilles Falquet , Douglas Teodoro

This document, based on feedback from UMR TETIS members and the scientific literature, provides a generic methodology for creating annotation guidelines and annotated textual datasets (corpora). It covers methodological aspects, as well as…

Information Retrieval · Computer Science 2026-01-21 Bahdja Boudoua , Nadia Guiffant , Mathieu Roche , Maguelonne Teisseire , Annelise Tran

The RareDis corpus contains more than 5,000 rare diseases and almost 6,000 clinical manifestations are annotated. Moreover, the Inter Annotator Agreement evaluation shows a relatively high agreement (F1-measure equal to 83.5% under exact…

Computation and Language · Computer Science 2021-12-10 Claudia Martínez-deMiguel , Isabel Segura-Bedmar , Esteban Chacón-Solano , Sara Guerrero-Aspizua

Medical texts are notoriously challenging to read. Properly measuring their readability is the first step towards making them more accessible. In this paper, we present a systematic study on fine-grained readability measurements in the…

Computation and Language · Computer Science 2024-10-29 Chao Jiang , Wei Xu