English
Related papers

Related papers: Self-Driving Datasets: From 20 Million Papers to N…

200 papers

The exponential growth of scientific knowledge has created significant barriers to cross-disciplinary knowledge discovery, synthesis and research collaboration. In response to this challenge, we present BioSage, a novel compound AI…

A main challenge of data-driven sciences is how to make maximal use of the progressively expanding databases of experimental datasets in order to keep research cumulative. We introduce the idea of a modeling-based dataset retrieval engine…

Quantitative Methods · Quantitative Biology 2015-06-19 Ali Faisal , Jaakko Peltonen , Elisabeth Georgii , Johan Rung , Samuel Kaski

Mentions of new concepts appear regularly in texts and require automated approaches to harvest and place them into Knowledge Bases (KB), e.g., ontologies and taxonomies. Existing datasets suffer from three issues, (i) mostly assuming that a…

Computation and Language · Computer Science 2023-09-04 Hang Dong , Jiaoyan Chen , Yuan He , Ian Horrocks

Understanding scientific papers requires more than answering isolated questions or summarizing content. It involves an integrated reasoning process that grounds textual and visual information, interprets experimental evidence, synthesizes…

Information Retrieval · Computer Science 2026-04-29 Yanjun Zhao , Tianxin Wei , Jiaru Zou , Xuying Ning , Yuanchen Bei , Lingjie Chen , Simmi Rana , Wendy H. Yang , Hanghang Tong , Jingrui He

Now a day's, search engines are been most widely used for extracting information's from various resources throughout the world. Where, majority of searches lies in the field of biomedical for retrieving related documents from various…

Information Retrieval · Computer Science 2009-12-14 Jayanthi Manicassamy , P. Dhavachelvan

As an important biomedical database, PubMed provides users with free access to abstracts of its documents. However, citations between these documents need to be collected from external data sources. Although previous studies have…

Digital Libraries · Computer Science 2021-11-02 Zhentao Liang , Jin Mao , Kun Lu , Gang Li

In the evolving landscape of clinical informatics, the integration and utilization of software tools developed through governmental funding represent a pivotal advancement in research and application. However, the dispersion of these tools…

Digital Libraries · Computer Science 2024-03-28 Jeremy R. Harper

Predictive models in biomedicine depend on structured assay data locked in the text, tables, and supplements of primary publications. This bottleneck is especially acute in targeted protein degradation (TPD), where each assay record must…

Quantitative Methods · Quantitative Biology 2026-05-13 Yaochen Rao , Farzaneh Jalalypour , N. M. Anoop Krishnan , Rocío Mercado

Highly specific datasets of scientific literature are important for both research and education. However, it is difficult to build such datasets at scale. A common approach is to build these datasets reductively by applying topic modeling…

Information Retrieval · Computer Science 2023-09-20 Nicholas Solovyev , Ryan Barron , Manish Bhattarai , Maksim E. Eren , Kim O. Rasmussen , Boian S. Alexandrov

Recently, significant progress has been made applying machine learning to the problem of table structure inference and extraction from unstructured documents. However, one of the greatest challenges remains the creation of datasets with…

Machine Learning · Computer Science 2021-11-22 Brandon Smock , Rohith Pesala , Robin Abraham

Public health reasoning requires population level inference grounded in scientific evidence, expert consensus, and safety constraints. However, it remains underexplored as a structured machine learning problem with limited supervised…

Computation and Language · Computer Science 2026-05-27 Zhaokun Yan , Shan Xu , Wuzheng Dong , Zhaohan Liu , Lijie Feng , Chengxiao Dai , Chen Tianqi , Binfan Liu , Yunpu Ma , Wenting Wei , Yingting Li , Yi Zhang , Tongning Wu

With the exponential increase in online scientific literature, identifying reliable domain-specific data has become increasingly important but also very challenging. Manual data collection and filtering for domain-specific scientific…

Information Retrieval · Computer Science 2026-03-10 Nikita Gautam , Doina Caragea , Ignacio Ciampitti , Federico Gomez

Unsupervised clustering has broad applications in data stratification, pattern investigation and new discovery beyond existing knowledge. In particular, clustering of bioactive molecules facilitates chemical space mapping,…

In recent years, the research landscape of machine learning in medical imaging has changed drastically from supervised to semi-, weakly- or unsupervised methods. This is mainly due to the fact that ground-truth labels are time-consuming and…

Image and Video Processing · Electrical Eng. & Systems 2021-10-04 Turkay Kart , Wenjia Bai , Ben Glocker , Daniel Rueckert

Biomedical literature often uses complex language and inaccessible professional terminologies. That is why simplification plays an important role in improving public health literacy. Applying Natural Language Processing (NLP) models to…

Computation and Language · Computer Science 2024-03-19 Zihao Li , Samuel Belkadi , Nicolo Micheletti , Lifeng Han , Matthew Shardlow , Goran Nenadic

AI-enabled precision medicine promises a transformational improvement in healthcare outcomes by enabling data-driven personalized diagnosis, prognosis, and treatment. However, the well-known "curse of dimensionality" and the clustered…

Machine Learning · Computer Science 2023-05-19 Amanda M. Buch , Conor Liston , Logan Grosenick

The extraction of structured knowledge from scientific literature remains a major bottleneck in nutraceutical research, particularly when identifying microbial strains involved in compound biosynthesis. This study presents a domain-adapted…

Quantitative Methods · Quantitative Biology 2025-12-30 Xinyang Sun , Nipon Sarmah , Miao Guo

Automatically locating named entities in natural language text - named entity recognition - is an important task in the biomedical domain. Many named entity mentions are ambiguous between several bioconcept types, however, causing text…

Computation and Language · Computer Science 2019-09-24 Chih-Hsuan Wei , Kyubum Lee , Robert Leaman , Zhiyong Lu

Manual chart review remains an extremely time-consuming and resource-intensive component of clinical research, requiring experts to extract often complex information from unstructured electronic health record (EHR) narratives. We present a…

Duplication, whether exact or partial, is a common issue in many datasets. In clinical notes data, duplication (and near duplication) can arise for many reasons, such as the pervasive use of templates, copy-pasting, or notes being generated…

Databases · Computer Science 2017-04-20 Sanjeev Shenoy , Tsung-Ting Kuo , Rodney Gabriel , Julian McAuley , Chun-Nan Hsu