中文
相关论文

相关论文: MedMentions: A Large Biomedical Corpus Annotated w…

200 篇论文

The number of biomedical literature on new biomedical concepts is rapidly increasing, which necessitates a reliable biomedical named entity recognition (BioNER) model for identifying new and unseen entity mentions. However, it is…

计算与语言 · 计算机科学 2022-03-15 Hyunjae Kim , Jaewoo Kang

The growth rate in the amount of biomedical documents is staggering. Unlocking information trapped in these documents can enable researchers and practitioners to operate confidently in the information world. Biomedical NER, the task of…

计算与语言 · 计算机科学 2021-06-24 Xiang Dai

Causal relation extraction (CRE) is central to biomedical text mining, but current resources often conflate causal relations with broader associations, restrict annotation to sentence-level examples, or focus mainly on explicit causal cues.…

计算与语言 · 计算机科学 2026-05-28 Ifeoluwa Kunle-John , Josiah Paul , Oluwatosin Agbaakin , Peter Aina , Ikenna Odezuligbo , Sydney Anuyah

Biomedical concept normalization links concept mentions in texts to a semantically equivalent concept in a biomedical knowledge base. This task is challenging as concepts can have different expressions in natural languages, e.g.…

计算与语言 · 计算机科学 2018-07-10 Roland Roller , Madeleine Kittner , Dirk Weissenborn , Ulf Leser

Large language models (LLMs) like ChatGPT can generate and revise text with human-level performance. These models come with clear limitations: they can produce inaccurate information, reinforce existing biases, and be easily misused. Yet,…

计算与语言 · 计算机科学 2025-07-04 Dmitry Kobak , Rita González-Márquez , Emőke-Ágnes Horvát , Jan Lause

Named Entity Recognition (NER) or the extraction of concepts from clinical text is the task of identifying entities in text and slotting them into categories such as problems, treatments, tests, clinical departments, occurrences (such as…

计算与语言 · 计算机科学 2022-08-31 Namrata Nath , Sang-Heon Lee , Ivan Lee

Objective: Social media-based public health research is crucial for epidemic surveillance, but most studies identify relevant corpora with keyword-matching. This study develops a system to streamline the process of curating colloquial…

计算与语言 · 计算机科学 2024-03-19 Yining Hua , Jiageng Wu , Shixu Lin , Minghui Li , Yujie Zhang , Dinah Foer , Siwen Wang , Peilin Zhou , Jie Yang , Li Zhou

Distributed representations of medical concepts have been used to support downstream clinical tasks recently. Electronic Health Records (EHR) capture different aspects of patients' hospital encounters and serve as a rich source for…

计算与语言 · 计算机科学 2020-01-07 Shaika Chowdhury , Chenwei Zhang , Philip S. Yu , Yuan Luo

We present MedConceptsQA, a dedicated open source benchmark for medical concepts question answering. The benchmark comprises of questions of various medical concepts across different vocabularies: diagnoses, procedures, and drugs. The…

计算与语言 · 计算机科学 2024-05-15 Ofir Ben Shoham , Nadav Rappoport

Clinical information extraction, which involves structuring clinical concepts from unstructured medical text, remains a challenging problem that could benefit from the inclusion of tabular background information available in electronic…

人工智能 · 计算机科学 2025-12-10 Paloma Rabaey , Stefan Heytens , Thomas Demeester

We present a system that constructs and maintains an up-to-date co-occurrence network of medical concepts based on continuously mining the latest biomedical literature. Users can explore this network visually via a concise online interface…

信息检索 · 计算机科学 2015-03-20 Alexei Yavlinsky

The number of clinical citations received from clinical guidelines or clinical trials has been considered as one of the most appropriate indicators for quantifying the clinical impact of biomedical papers. Therefore, the early prediction of…

计算与语言 · 计算机科学 2022-10-24 Xin Li , Xuli Tang , Qikai Cheng

Current medical language model (LM) benchmarks often over-simplify the complexities of day-to-day clinical practice tasks and instead rely on evaluating LMs on multiple-choice board exam questions. In psychiatry especially, these challenges…

We describe the CZ Software Mentions dataset, a new dataset of software mentions in biomedical papers. Plain-text software mentions are extracted with a trained SciBERT model from several sources: the NIH PubMed Central collection and from…

数字图书馆 · 计算机科学 2022-09-29 Ana-Maria Istrate , Donghui Li , Dario Taraborelli , Michaela Torkar , Boris Veytsman , Ivana Williams

We report characteristics of in-text citations in over five million full text articles from two large databases - the PubMed Central Open Access subset and Elsevier journals - as functions of time, textual progression, and scientific field.…

数字图书馆 · 计算机科学 2017-10-10 Kevin W. Boyack , Nees Jan van Eck , Giovanni Colavizza , Ludo Waltman

Disease name recognition and normalization, which is generally called biomedical entity linking, is a fundamental process in biomedical text mining. Recently, neural joint learning of both tasks has been proposed to utilize the mutual…

计算与语言 · 计算机科学 2021-04-22 Shogo Ujiie , Hayate Iso , Shuntaro Yada , Shoko Wakamiya , Eiji Aramaki

The Medical Subject Headings (MeSH), one of the main knowledge organization systems in the biomedical domain, continuously evolves to reflect the latest scientific discoveries in health and life sciences. Previous research has focused on…

社会与信息网络 · 计算机科学 2025-10-09 Jenny Copara , Nona Naderi , Gilles Falquet , Douglas Teodoro

This document, based on feedback from UMR TETIS members and the scientific literature, provides a generic methodology for creating annotation guidelines and annotated textual datasets (corpora). It covers methodological aspects, as well as…

信息检索 · 计算机科学 2026-01-21 Bahdja Boudoua , Nadia Guiffant , Mathieu Roche , Maguelonne Teisseire , Annelise Tran

The RareDis corpus contains more than 5,000 rare diseases and almost 6,000 clinical manifestations are annotated. Moreover, the Inter Annotator Agreement evaluation shows a relatively high agreement (F1-measure equal to 83.5% under exact…

Medical texts are notoriously challenging to read. Properly measuring their readability is the first step towards making them more accessible. In this paper, we present a systematic study on fine-grained readability measurements in the…

计算与语言 · 计算机科学 2024-10-29 Chao Jiang , Wei Xu