English
Related papers

Related papers: Self-Driving Datasets: From 20 Million Papers to N…

200 papers

We present PubMed 200k RCT, a new dataset based on PubMed for sequential sentence classification. The dataset consists of approximately 200,000 abstracts of randomized controlled trials, totaling 2.3 million sentences. Each sentence of each…

Computation and Language · Computer Science 2017-10-18 Franck Dernoncourt , Ji Young Lee

Biological relation networks contain rich information for understanding the biological mechanisms behind the relationship of entities such as genes, proteins, diseases, and chemicals. The vast growth of biomedical literature poses…

Computation and Language · Computer Science 2025-01-27 Po-Ting Lai , Chih-Hsuan Wei , Shubo Tian , Robert Leaman , Zhiyong Lu

Extracting structured information from scientific literature is critical for accelerating discovery, yet Large Language Models (LLMs) often struggle in specialized domains that require expert knowledge and generalize poorly across tasks. We…

Computation and Language · Computer Science 2026-05-22 Tek Raj Chhetri , Yibei Chen , Puja Trivedi , Dorota Jarecka , Saif Haobsh , Patrick Ray , Lydia Ng , Satrajit S. Ghosh

Subject classification schemes are foundational to the organization, evaluation, and navigation of scientific knowledge. While expert-curated systems like Scopus provide widely used taxonomies, they often suffer from coarse granularity,…

Digital Libraries · Computer Science 2025-12-30 Zhuoqi Lyu , Qing Ke

Entity linking faces significant challenges such as prolific variations and prevalent ambiguities, especially in high-value domains with myriad entities. Standard classification approaches suffer from the annotation bottleneck and cannot…

Computation and Language · Computer Science 2022-05-24 Sheng Zhang , Hao Cheng , Shikhar Vashishth , Cliff Wong , Jinfeng Xiao , Xiaodong Liu , Tristan Naumann , Jianfeng Gao , Hoifung Poon

Effective data-driven biomedical discovery requires data curation: a time-consuming process of finding, organizing, distilling, integrating, interpreting, annotating, and validating diverse information into a structured form suitable for…

The amount of scientific papers published every day is daunting and constantly increasing. Keeping up with literature represents a challenge. If one wants to start exploring new topics it is hard to have a big picture without reading lots…

Information Retrieval · Computer Science 2020-11-10 Alberto Calderone

Recognizing the layout of unstructured digital documents is an important step when parsing the documents into structured machine-readable format for downstream applications. Deep neural networks that are developed for computer vision have…

Computation and Language · Computer Science 2019-08-22 Xu Zhong , Jianbin Tang , Antonio Jimeno Yepes

Large Language Models (LLMs) are increasingly adopted for applications in healthcare, reaching the performance of domain experts on tasks such as question answering and document summarisation. Despite their success on these tasks, it is…

Computation and Language · Computer Science 2025-05-20 Aishik Nagar , Viktor Schlegel , Thanh-Tung Nguyen , Hao Li , Yuping Wu , Kuluhan Binici , Stefan Winkler

The biomedical literature contains a vast collection of omics studies, yet most published data remain functionally inaccessible for computational reuse. When raw data are deposited in public repositories, essential information for…

Genomics · Quantitative Biology 2026-03-16 Alexandre Hutton , Jesse G. Meyer

Progress in biomedical Named Entity Recognition (NER) and Entity Linking (EL) is currently hindered by a fragmented data landscape, a lack of resources for building explainable models, and the limitations of semantically-blind evaluation…

Computation and Language · Computer Science 2025-11-17 Nishant Mishra , Wilker Aziz , Iacer Calixto

Biomedical semantic question answering rooted in information retrieval can play a crucial role in keeping up to date with vast, rapidly evolving and ever-growing biomedical literature. A robust system can help researchers, healthcare…

Information Retrieval · Computer Science 2025-07-09 Shashank Verma , Fengyi Jiang , Xiangning Xue

Though exponentially growing health-related literature has been made available to a broad audience online, the language of scientific articles can be difficult for the general public to understand. Therefore, adapting this expert-level…

Computation and Language · Computer Science 2022-10-25 Kush Attal , Brian Ondov , Dina Demner-Fushman

Named-entity recognition (NER) is fundamental to extracting structured information from the >80% of healthcare data that resides in unstructured clinical notes and biomedical literature. Despite recent advances with large language models,…

Computation and Language · Computer Science 2025-08-05 Maziyar Panahi

Current medical language model (LM) benchmarks often over-simplify the complexities of day-to-day clinical practice tasks and instead rely on evaluating LMs on multiple-choice board exam questions. In psychiatry especially, these challenges…

There has been rapid growth in biomedical literature, yet capturing the heterogeneity of the bibliographic information of these articles remains relatively understudied. Although graph mining research via heterogeneous graph neural networks…

Machine Learning · Computer Science 2023-08-28 Eric W Lee , Joyce C Ho

Entity recognition is a critical first step to a number of clinical NLP applications, such as entity linking and relation extraction. We present the first attempt to apply state-of-the-art entity recognition approaches on a newly released…

Computation and Language · Computer Science 2019-10-04 Kathleen C. Fraser , Isar Nejadgholi , Berry De Bruijn , Muqun Li , Astha LaPlante , Khaldoun Zine El Abidine

The effectiveness of artificial intelligence (AI) in healthcare is significantly hindered by unstructured clinical documentation, which results in noisy, inconsistent, and logically fragmented training data. To address this challenge, we…

Machine Learning · Computer Science 2025-10-21 Dun Liu , Qin Pang , Guangai Liu , Hongyu Mou , Jipeng Fan , Yiming Miao , Pin-Han Ho , Limei Peng

Large language models (LLMs) are increasingly recognized as valuable tools across the medical environment, supporting clinical, research, and administrative workflows. However, strict privacy and network security regulations in hospital…

Computation and Language · Computer Science 2026-01-09 Seokhwan Ko , Donghyeon Lee , Jaewoo Chun , Hyungsoo Han , Junghwan Cho

The fundamental process of evidence extraction and synthesis in evidence-based medicine involves extracting PICO (Population, Intervention, Comparison, and Outcome) elements from biomedical literature. However, Outcomes, being the most…

Computation and Language · Computer Science 2025-06-09 Yiliang Zhou , Abigail M. Newbury , Gongbo Zhang , Betina Ross Idnay , Hao Liu , Chunhua Weng , Yifan Peng