English
Related papers

Related papers: Hermes: Large DEL Datasets Train Generalizable Pro…

200 papers

Protein-ligand binding affinity prediction is essential for drug discovery and toxicity assessment. While machine learning (ML) promises fast and accurate predictions, its progress is constrained by the availability of reliable data. In…

Providing accurate and reliable predictions about the future of an epidemic is an important problem for enabling informed public health decisions. Recent works have shown that leveraging data-driven solutions that utilize advances in deep…

Machine Learning · Computer Science 2023-11-21 Harshavardhan Kamarthi , B. Aditya Prakash

Confidence-based pseudo-label selection usually generates overly confident yet incorrect predictions, due to the early misleadingness of model and overfitting inaccurate pseudo-labels in the learning process, which heavily degrades the…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Peng Zhang , Zhihui Lai , Heng Kong

Pre-trained Transformer language models (LM) have become go-to text representation encoders. Prior research fine-tunes deep LMs to encode text sequences such as sentences and passages into single dense vector representations for efficient…

Computation and Language · Computer Science 2021-09-22 Luyu Gao , Jamie Callan

Proteins, essential to biological systems, perform functions intricately linked to their three-dimensional structures. Understanding the relationship between protein structures and their amino acid sequences remains a core challenge in…

Quantitative Methods · Quantitative Biology 2024-11-04 Liang He , Peiran Jin , Yaosen Min , Shufang Xie , Lijun Wu , Tao Qin , Xiaozhuan Liang , Kaiyuan Gao , Yuliang Jiang , Tie-Yan Liu

Predicting the structure of interacting chains is crucial for understanding biological systems and developing new drugs. Large-scale pre-trained Protein Language Models (PLMs), such as ESM2, have shown impressive abilities in extracting…

Biomolecules · Quantitative Biology 2023-12-05 Shuxian Zou , Hui Li , Shentong Mo , Xingyi Cheng , Eric Xing , Le Song

Density Estimation Trees (DETs) are decision trees trained on a multivariate dataset to estimate its probability density function. While not competitive with kernel techniques in terms of accuracy, they are incredibly fast, embarrassingly…

Applications · Statistics 2016-12-21 Lucio Anderlini

Drug discovery projects entail cycles of design, synthesis, and testing that yield a series of chemically related small molecules whose properties, such as binding affinity to a given target protein, are progressively tailored to a…

Machine Learning · Computer Science 2020-02-10 Paul Maragakis , Hunter Nisonoff , Brian Cole , David E. Shaw

Discovering mutations enhancing protein-protein interactions (PPIs) is critical for advancing biomedical research and developing improved therapeutics. While machine learning approaches have substantially advanced the field, they often…

Deep learning models that leverage large datasets are often the state of the art for modelling molecular properties. When the datasets are smaller (< 2000 molecules), it is not clear that deep learning approaches are the right modelling…

Computational Engineering, Finance, and Science · Computer Science 2022-12-07 Gary Tom , Riley J. Hickman , Aniket Zinzuwadia , Afshan Mohajeri , Benjamin Sanchez-Lengeling , Alan Aspuru-Guzik

In the domain of computer vision, Parameter-Efficient Tuning (PET) is increasingly replacing the traditional paradigm of pre-training followed by full fine-tuning. PET is particularly favored for its effectiveness in large foundation…

Computer Vision and Pattern Recognition · Computer Science 2025-01-16 Jiaqi Huang , Zunnan Xu , Ting Liu , Yong Liu , Haonan Han , Kehong Yuan , Xiu Li

A protein's function depends critically on its conformational ensemble, a collection of energy weighted structures whose balance depends on temperature and environment. Though recent deep learning (DL) methods have substantially advanced…

Biomolecules · Quantitative Biology 2026-01-09 Myeongsang Lee , Lauren L. Porter

In structure-based drug design, accurately estimating the binding affinity between a candidate ligand and its protein receptor is a central challenge. Recent advances in artificial intelligence, particularly deep learning, have demonstrated…

Biomolecules · Quantitative Biology 2025-09-18 Md Masud Rana , Farjana Tasnim Mukta , Duc D. Nguyen

Large Artificial Neural Network (ANN) models have demonstrated success in various domains, including general text and image generation, drug discovery, and protein-RNA (ribonucleic acid) binding tasks. However, these models typically demand…

Biomolecules · Quantitative Biology 2025-11-13 Stanislav Selitskiy

The field of drug discovery hinges on the accurate prediction of binding affinity between prospective drug molecules and target proteins, especially when such proteins directly influence disease progression. However, estimating binding…

GNNs and chemical fingerprints are the predominant approaches to representing molecules for property prediction. However, in NLP, transformers have become the de-facto standard for representation learning thanks to their strong downstream…

Machine Learning · Computer Science 2020-10-26 Seyone Chithrananda , Gabriel Grand , Bharath Ramsundar

The majority of machine learning scoring functions used in drug discovery for predicting protein-ligand binding poses and affinities have been trained on the PDBBind dataset. However, it is unclear whether these new scoring functions are…

Biological Physics · Physics 2026-01-13 Jie Li , Xingyi Guan , Oufan Zhang , Kunyang Sun , Yingze Wang , Dorian Bagni , Teresa Head-Gordon

In the competitive landscape of sponsored search, balancing retrieval quality with production latency is a critical challenge. While large retrieval models based on Small Language Models (SLMs) such as Qwen3-Embedding-4B/8B set strong upper…

Information Retrieval · Computer Science 2026-05-25 Vipul Gupta , Shikhar Mohan , Lakshya Kumar , Pranjal Chitale , Nikit Begwani , Amit Singh , Manik Varma

Many sensitive domains -- such as the clinical domain -- lack widely available datasets due to privacy risks. The increasing generative capabilities of large language models (LLMs) have made synthetic datasets a viable path forward. In this…

Computation and Language · Computer Science 2025-06-03 Thomas Vakili , Aron Henriksson , Hercules Dalianis

Empirical scoring functions based on either molecular force fields or cheminformatics descriptors are widely used, in conjunction with molecular docking, during the early stages of drug discovery to predict potency and binding affinity of a…

Machine Learning · Computer Science 2017-03-31 Joseph Gomes , Bharath Ramsundar , Evan N. Feinberg , Vijay S. Pande
‹ Prev 1 3 4 5 6 7 10 Next ›