English
Related papers

Related papers: Rethinking the generalization of drug target affin…

200 papers

Positive-Unlabeled (PU) learning aims to learn a model with rare positive samples and abundant unlabeled samples. Compared with classical binary classification, the task of PU learning is much more challenging due to the existence of many…

Computer Vision and Pattern Recognition · Computer Science 2022-12-01 Chengming Xu , Chen Liu , Siqian Yang , Yabiao Wang , Shijie Zhang , Lijie Jia , Yanwei Fu

Methods for split conformal prediction leverage calibration samples to transform any prediction rule into a set-prediction rule that complies with a target coverage probability. Existing methods provide remarkably strong performance…

Machine Learning · Statistics 2025-10-15 Santiago Mazuelas

Current pharmaceutical formulation development still strongly relies on the traditional trial-and-error approach by individual experiences of pharmaceutical scientists, which is laborious, time-consuming and costly. Recently, deep learning…

Machine Learning · Computer Science 2018-12-05 Yilong Yang , Zhuyifan Ye , Yan Su , Qianqian Zhao , Xiaoshan Li , Defang Ouyang

Accurate prediction of Drug-Target Affinity (DTA) is of vital importance in early-stage drug discovery, facilitating the identification of drugs that can effectively interact with specific targets and regulate their activities. While wet…

Biomolecules · Quantitative Biology 2023-10-18 Qizhi Pei , Lijun Wu , Jinhua Zhu , Yingce Xia , Shufang Xie , Tao Qin , Haiguang Liu , Tie-Yan Liu , Rui Yan

Batched synthesis and testing of molecular designs is the key bottleneck of drug development. There has been great interest in leveraging biomolecular foundation models as surrogates to accelerate this process. In this work, we show how to…

Deep learning models are being adopted and applied on various critical decision-making tasks, yet they are trained to provide point predictions without providing degrees of confidence. The trustworthiness of deep learning models can be…

Machine Learning · Computer Science 2024-10-28 Daniel Nolte , Souparno Ghosh , Ranadip Pal

A growing body of research has demonstrated the inability of NLP models to generalize compositionally and has tried to alleviate it through specialized architectures, training schemes, and data augmentation, among other approaches. In this…

Computation and Language · Computer Science 2022-11-03 Shivanshu Gupta , Sameer Singh , Matt Gardner

Virtual Screening (VS) of vast compound libraries guided by Artificial Intelligence (AI) models is a highly productive approach to early drug discovery. Data splitting is crucial for better benchmarking of such AI models. Traditional random…

Quantitative Methods · Quantitative Biology 2024-07-02 Qianrong Guo , Saiveth Hernandez-Hernandez , Pedro J Ballester

Over the past few years, deep learning methods have been applied for a wide range of Software Engineering (SE) tasks, including in particular for the important task of automatically predicting and localizing faults in software. With the…

Software Engineering · Computer Science 2024-02-09 Adil Mukhtar , Dietmar Jannach , Franz Wotawa

Failure of machine learning models to generalize to new data is a core problem limiting the reliability of AI systems, partly due to the lack of simple and robust methods for comparing new data to the original training dataset. We propose a…

Machine Learning · Computer Science 2025-02-26 W. Max Schreyer , Christopher Anderson , Reid F. Thompson

A widely recognized limitation of molecular prediction models is their reliance on structures observed in the training data, resulting in poor generalization to out-of-distribution compounds. Yet in drug discovery, the compounds most…

Machine Learning · Computer Science 2026-01-05 Jina Kim , Jeffrey Willette , Bruno Andreis , Sung Ju Hwang

Artificial intelligence, trained via machine learning or computational statistics algorithms, holds much promise for the improvement of small molecule drug discovery. However, structure-activity data are high dimensional with low…

Applications · Statistics 2018-07-25 Oliver Watson , Isidro Cortes-Ciriano , Aimee Taylor , James A Watson

Motivation: Drug discovery demands rapid quantification of compound-protein interaction (CPI). However, there is a lack of methods that can predict compound-protein affinity from sequences alone with high applicability, accuracy, and…

Biomolecules · Quantitative Biology 2020-12-17 Mostafa Karimi , Di Wu , Zhangyang Wang , Yang Shen

Accurate prediction of protein-ligand binding affinities is crucial for drug development. Recent advances in machine learning show promising results on this task. However, these methods typically rely heavily on labeled data, which can be…

Machine Learning · Computer Science 2024-06-13 Meng Liu , Saee Gopal Paliwal

Drug similarity has been studied to support downstream clinical tasks such as inferring novel properties of drugs (e.g. side effects, indications, interactions) from known properties. The growing availability of new types of drug features…

Machine Learning · Computer Science 2018-05-01 Tengfei Ma , Cao Xiao , Jiayu Zhou , Fei Wang

Recently, Saeb et al (2017) showed that, in diagnostic machine learning applications, having data of each subject randomly assigned to both training and test sets (record-wise data split) can lead to massive underestimation of the…

Recent work to enhance data partitioning strategies for more realistic model evaluation face challenges in providing a clear optimal choice. This study addresses these challenges, focusing on morphological segmentation and synthesizing…

Computation and Language · Computer Science 2024-04-16 Zoey Liu , Bonnie J. Dorr

Spurious correlations in training data significantly hinder the generalization capability of machine learning models when faced with distribution shifts, leading to the proposition of numberous debiasing methods. However, it remains to be…

Machine Learning · Computer Science 2025-05-22 Peng Kuang , Zhibo Wang , Zhixuan Chu , Jingyi Wang , Kui Ren

Matching is a commonly used causal inference study design in observational studies. Through matching on measured confounders between different treatment groups, valid randomization inferences can be conducted under the no unmeasured…

Methodology · Statistics 2024-09-20 Jeffrey Zhang , Siyu Heng

In this paper, we investigate potential biases in datasets used to make drug binding predictions using machine learning. We investigate a recently published metric called the Asymmetric Validation Embedding (AVE) bias which is used to…

Biomolecules · Quantitative Biology 2020-01-13 Brian Davis , Kevin Mcloughlin , Jonathan Allen , Sally Ellingson