English
Related papers

Related papers: PMC text mining subset in BioC: 2.3 million full t…

200 papers

Continuity of care is crucial to ensuring positive health outcomes for patients discharged from an inpatient hospital setting, and improved information sharing can help. To share information, caregivers write discharge notes containing…

Computation and Language · Computer Science 2021-06-07 James Mullenbach , Yada Pruksachatkun , Sean Adler , Jennifer Seale , Jordan Swartz , T. Greg McKelvey , Hui Dai , Yi Yang , David Sontag

There is an ongoing need for scalable tools to aid researchers in both retrospective and prospective standardization of discrete entity types -- such as disease names, cell types or chemicals -- that are used in metadata associated with…

Databases · Computer Science 2024-07-04 Rafael S. Gonçalves , Jason Payne , Amelia Tan , Carmen Benitez , Jamie Haddock , Robert Gentleman

We introduce S2ORC, a large corpus of 81.1M English-language academic papers spanning many academic disciplines. The corpus consists of rich metadata, paper abstracts, resolved bibliographic references, as well as structured full text for…

Computation and Language · Computer Science 2020-07-08 Kyle Lo , Lucy Lu Wang , Mark Neumann , Rodney Kinney , Dan S. Weld

We present a simple text mining method that is easy to implement, requires minimal data collection and preparation, and is easy to use for proposing ranked associations between a list of target terms and a key phrase. We call this method…

Information Retrieval · Computer Science 2019-06-13 Finn Kuusisto , John Steill , Zhaobin Kuang , James Thomson , David Page , Ron Stewart

The development of AI-assisted chemical synthesis tools requires comprehensive datasets covering diverse reaction types, yet current high-throughput experimental (HTE) approaches are expensive and limited in scope. Chemical literature…

Information Retrieval · Computer Science 2025-07-01 Kexin Chen , Yuyang Du , Junyou Li , Hanqun Cao , Menghao Guo , Xilin Dang , Lanqing Li , Jiezhong Qiu , Pheng Ann Heng , Guangyong Chen

Medical Subject Heading (MeSH) indexing refers to the problem of assigning a given biomedical document with the most relevant labels from an extremely large set of MeSH terms. Currently, the vast number of biomedical articles in the PubMed…

Computation and Language · Computer Science 2022-04-29 Xindi Wang , Robert E. Mercer , Frank Rudzicz

Creating scientific publications is a complex process, typically composed of a number of different activities, such as designing the experiments, data preparation, programming software and writing and editing the manuscript. The information…

Digital Libraries · Computer Science 2018-02-06 Dominika Tkaczyk , Andrew Collins , Joeran Beel

Keyphrase generation is the task consisting in generating a set of words or phrases that highlight the main topics of a document. There are few datasets for keyphrase generation in the biomedical domain and they do not meet the expectations…

Computation and Language · Computer Science 2022-11-23 Mael Houbre , Florian Boudin , Beatrice Daille

Background: Biomedical research projects deal with data management requirements from multiple sources like funding agencies' guidelines, publisher policies, discipline best practices, and their own users' needs. We describe functional and…

Citing comprehensively and appropriately has become a challenging task with the explosive growth of scientific publications. Current citation recommendation systems aim to recommend a list of scientific papers for a given text context or a…

Information Retrieval · Computer Science 2024-03-05 Kehan Long , Shasha Li , Pancheng Wang , Chenlong Bao , Jintao Tang , Ting Wang

We present a system that constructs and maintains an up-to-date co-occurrence network of medical concepts based on continuously mining the latest biomedical literature. Users can explore this network visually via a concise online interface…

Information Retrieval · Computer Science 2015-03-20 Alexei Yavlinsky

Nanopublications are a Linked Data format for scholarly data publishing that has received considerable uptake in the last few years. In contrast to the common Linked Data publishing practice, nanopublications work at the granular level of…

Scientific action graphs extraction from materials synthesis procedures is important for reproducible research, machine automation, and material prediction. But the lack of annotated data has hindered progress in this field. We demonstrate…

Computation and Language · Computer Science 2022-10-25 Xianjun Yang , Ya Zhuo , Julia Zuo , Xinlu Zhang , Stephen Wilson , Linda Petzold

This study presents OpenExtract, an open-source pipeline for automated data extraction in large-scale systematic literature reviews. The pipeline queries large language models (LLMs) to predict data entries based on relevant sections of…

In the implementation and use of research information systems (RIS) in scientific institutions, text data mining and semantic technologies are a key technology for the meaningful use of large amounts of data. It is not the collection of…

Digital Libraries · Computer Science 2018-12-12 Otmane Azeroual , Gunter Saake , Mohammad Abuosba , Joachim Schöpfel

The exponential growth of biomedical texts such as biomedical literature and electronic health records (EHRs), poses a significant challenge for clinicians and researchers to access clinical information efficiently. To tackle this…

Computation and Language · Computer Science 2023-07-17 Qianqian Xie , Zheheng Luo , Benyou Wang , Sophia Ananiadou

Earlier techniques of text mining included algorithms like k-means, Naive Bayes, SVM which classify and cluster the text document for mining relevant information about the documents. The need for improving the mining techniques has us…

Information Retrieval · Computer Science 2016-05-10 Jinju Joby , Jyothi Korra

We present a novel approach to automating the identification of risk factors for diseases from medical literature, leveraging pre-trained models in the bio-medical domain, while tuning them for the specific task. Faced with the challenges…

Computation and Language · Computer Science 2024-07-11 Maxim Rubchinsky , Ella Rabinovich , Adi Shraibman , Netanel Golan , Tali Sahar , Dorit Shweiki

The use of social media data, like Twitter, for biomedical research has been gradually increasing over the years. With the COVID-19 pandemic, researchers have turned to more nontraditional sources of clinical data to characterize the…

Information Retrieval · Computer Science 2021-07-28 Luis Alberto Robles Hernandez , Tiffany J. Callahan , Juan M. Banda

We present SciClaims, an interactive web-based system for end-to-end scientific claim analysis in the biomedical domain. Designed for high-stakes use cases such as systematic literature reviews and patent validation, SciClaims extracts…

Computation and Language · Computer Science 2026-01-09 Raúl Ortega , José Manuel Gómez-Pérez