English
Related papers

Related papers: MeSHup: A Corpus for Full Text Biomedical Document…

200 papers

The quest for seeking health information has swamped the web with consumers' health-related questions. Generally, consumers use overly descriptive and peripheral information to express their medical condition or other healthcare needs,…

Computation and Language · Computer Science 2022-06-16 Shweta Yadav , Deepak Gupta , Dina Demner-Fushman

Information retrieval (IR) for precision medicine (PM) often involves looking for multiple pieces of evidence that characterize a patient case. This typically includes at least the name of a condition and a genetic variation that applies to…

Computation and Language · Computer Science 2020-12-18 Jiho Noh , Ramakanth Kavuluru

Automatic text summarization tools help users in biomedical domain to acquire their intended information from various textual resources more efficiently. Some of the biomedical text summarization systems put the basis of their sentence…

Computation and Language · Computer Science 2017-05-31 Milad Moradi , Nasser Ghadiri

In recent years, the trend of deploying digital systems in numerous industries has hiked. The health sector has observed an extensive adoption of digital systems and services that generate significant medical records. Electronic health…

Computation and Language · Computer Science 2022-03-01 Neel Kanwal , Giuseppe Rizzo

The pandemic has accelerated the pace at which COVID-19 scientific papers are published. In addition, the process of manually assigning semantic indexes to these papers by experts is even more time-consuming and overwhelming in the current…

Information Retrieval · Computer Science 2020-10-08 Nima Ebadi , Peyman Najafirad

Given the dominance of dense retrievers that do not generalize well beyond their training dataset distributions, domain-specific test sets are essential in evaluating retrieval. There are few test datasets for retrieval systems intended for…

Information Retrieval · Computer Science 2025-07-01 Nadia Athar Sheikh , Daniel Buades Marcos , Anne-Laure Jousse , Akintunde Oladipo , Olivier Rousseau , Jimmy Lin

Text clustering serves as a fundamental technique for organizing and interpreting unstructured textual data, particularly in contexts where manual annotation is prohibitively costly. With the rapid advancement of Large Language Models…

Computation and Language · Computer Science 2025-10-08 Chen Huang , Guoxiu He

Named entities in text documents are the names of people, organization, location or other types of objects in the documents that exist in the real world. A persisting research challenge is to use computational techniques to identify such…

Computation and Language · Computer Science 2019-07-09 Abdulkareem Alsudais , Hovig Tchalian

Keyphrase generation is the task consisting in generating a set of words or phrases that highlight the main topics of a document. There are few datasets for keyphrase generation in the biomedical domain and they do not meet the expectations…

Computation and Language · Computer Science 2022-11-23 Mael Houbre , Florian Boudin , Beatrice Daille

Whilst there has been growing progress in Entity Linking (EL) for general language, existing datasets fail to address the complex nature of health terminology in layman's language. Meanwhile, there is a growing need for applications that…

Computation and Language · Computer Science 2020-10-09 Marco Basaldella , Fangyu Liu , Ehsan Shareghi , Nigel Collier

Biomedical named entity recognition (NER) and entity linking (EL) strongly depend on annotated corpora, but the utility of these resources for benchmarking is often assumed rather than characterized. We present a corpus-centric framework…

Computation and Language · Computer Science 2026-05-21 Robert Leaman , Rezarta Islamaj , Zhiyong Lu

We present Reddit Health Online Talk (RedHOT), a corpus of 22,000 richly annotated social media posts from Reddit spanning 24 health conditions. Annotations include demarcations of spans corresponding to medical claims, personal…

Computation and Language · Computer Science 2023-02-09 Somin Wadhwa , Vivek Khetan , Silvio Amir , Byron Wallace

Electronic Health Records (EHRs) play an important role in the healthcare system. However, their complexity and vast volume pose significant challenges to data interpretation and analysis. Recent advancements in Artificial Intelligence…

The abstract of a scientific paper distills the contents of the paper into a short paragraph. In the biomedical literature, it is customary to structure an abstract into discourse categories like BACKGROUND, OBJECTIVE, METHOD, RESULT, and…

Computation and Language · Computer Science 2020-05-28 Soumya Banerjee , Debarshi Kumar Sanyal , Samiran Chattopadhyay , Plaban Kumar Bhowmick , Parthapratim Das

The USA Food and Drug Administration has accorded increasing importance to patient-reported problems in clinical and research settings. In this paper, we explore one of the largest online datasets comprising 170,141 open-ended self-reported…

Computation and Language · Computer Science 2023-05-09 Lakshmi Arbatti , Abhishek Hosamath , Vikram Ramanarayanan , Ira Shoulson

Keeping track of all relevant recent publications and experimental results for a research area is a challenging task. Prior work has demonstrated the efficacy of information extraction models in various scientific areas. Recently, several…

Computation and Language · Computer Science 2023-10-25 Timo Pierre Schrader , Matteo Finco , Stefan Grünewald , Felix Hildebrand , Annemarie Friedrich

Motivation: Accurate data analysis and quality control is critical for metagenomic studies. Though many tools exist to analyze metagenomic data there is no consistent framework to integrate and run these tools across projects. Currently,…

Quantitative Methods · Quantitative Biology 2020-09-28 David C Danko , Chris Mason

We present a novel system providing summaries for Computer Science publications. Through a qualitative user study, we identified the most valuable scenarios for discovery, exploration and understanding of scientific documents. Based on…

Multi-document summarization is a challenging task for which there exists little large-scale datasets. We propose Multi-XScience, a large-scale multi-document summarization dataset created from scientific articles. Multi-XScience introduces…

Computation and Language · Computer Science 2020-10-28 Yao Lu , Yue Dong , Laurent Charlin

In recent years, many methods have been developed to identify important portions of text documents. Summarization tools can utilize these methods to extract summaries from large volumes of textual information. However, to identify concepts…

Computation and Language · Computer Science 2019-03-08 Milad Moradi