中文
相关论文

相关论文: ChemBERTa-2: Towards Chemical Foundation Models

200 篇论文

GNNs and chemical fingerprints are the predominant approaches to representing molecules for property prediction. However, in NLP, transformers have become the de-facto standard for representation learning thanks to their strong downstream…

机器学习 · 计算机科学 2020-10-26 Seyone Chithrananda , Gabriel Grand , Bharath Ramsundar

With the emergence of Transformer architectures and their powerful understanding of textual data, a new horizon has opened up to predict the molecular properties based on text description. While SMILES are the most common form of…

化学物理 · 物理学 2023-10-11 Suryanarayanan Balaji , Rishikesh Magar , Yayati Jadhav , Amir Barati Farimani

Recent advances in large language models (LLMs) have demonstrated transformative potential across diverse fields. While LLMs have been applied to molecular simplified molecular input line entry system (SMILES) in computer-aided synthesis…

机器学习 · 计算机科学 2026-01-07 Kenan Li , Yijian Zhang , Jin Wang , Haipeng Gan , Zeying Sun , Xiaoguang Lei , Hao Dong

Molecular property prediction is an increasingly critical task within drug discovery and development. Typically, neural networks can learn molecular properties using graph-based, language-based or feature-based methods. Recent advances in…

机器学习 · 计算机科学 2025-07-31 Philip Spence , Brooks Paige , Anne Osbourn

Recent advancements in computational chemistry have leveraged the power of trans-former-based language models, such as MoLFormer, pre-trained using a vast amount of simplified molecular-input line-entry system (SMILES) sequences, to…

生物大分子 · 定量生物学 2024-11-05 Tianhao Peng , Yuchen Li , Xuhong Li , Jiang Bian , Zeke Xie , Ning Sui , Shahid Mumtaz , Yanwu Xu , Linghe Kong , Haoyi Xiong

We apply a Transformer architecture, specifically BERT, to learn flexible and high quality molecular representations for drug discovery problems. We study the impact of using different combinations of self-supervised tasks for pre-training,…

机器学习 · 计算机科学 2020-11-30 Benedek Fabian , Thomas Edlich , Héléna Gaspar , Marwin Segler , Joshua Meyers , Marco Fiscato , Mohamed Ahmed

In the computational prediction of chemical compound properties, molecular descriptors and fingerprints encoded to low dimensional vectors are used. The selection of proper molecular descriptors and fingerprints is both important and…

机器学习 · 计算机科学 2020-10-23 Sangrak Lim , Yong Oh Lee

Chemical representation learning has gained increasing interest due to the limited availability of supervised data in fields such as drug and materials design. This interest particularly extends to chemical language representation learning,…

化学物理 · 物理学 2024-08-06 Jun-Hyung Park , Yeachan Kim , Mingyu Lee , Hyuntae Park , SangKeun Lee

Molecular property prediction is essential in chemistry, especially for drug discovery applications. However, available molecular property data is often limited, encouraging the transfer of information from related data. Transfer learning…

机器学习 · 计算机科学 2022-07-07 Johan Broberg , Maria Bånkestad , Erik Ylipää

In drug-discovery-related tasks such as virtual screening, machine learning is emerging as a promising way to predict molecular properties. Conventionally, molecular fingerprints (numerical representations of molecules) are calculated…

机器学习 · 计算机科学 2019-11-13 Shion Honda , Shoi Shi , Hiroki R. Ueda

Models based on machine learning can enable accurate and fast molecular property predictions, which is of interest in drug discovery and material design. Various supervised machine learning models have demonstrated promising performance,…

机器学习 · 计算机科学 2022-12-15 Jerret Ross , Brian Belgodere , Vijil Chenthamarakshan , Inkit Padhi , Youssef Mroueh , Payel Das

Models that accurately predict properties based on chemical structure are valuable tools in drug discovery. However, for many properties, public and private training sets are typically small, and it is difficult for the models to generalize…

定量方法 · 定量生物学 2022-11-08 Oscar Méndez-Lucio , Christos Nicolaou , Berton Earnshaw

Machine learning is becoming a preferred method for the virtual screening of organic materials due to its cost-effectiveness over traditional computationally demanding techniques. However, the scarcity of labeled data for organic materials…

化学物理 · 物理学 2024-03-06 Chengwei Zhang , Yushuang Zhai , Ziyang Gong , Hongliang Duan , Yuan-Bin She , Yun-Fang Yang , An Su

Large-scale pre-training methodologies for chemical language models represent a breakthrough in cheminformatics. These methods excel in tasks such as property prediction and molecule generation by learning contextualized representations of…

In drug discovery, predicting the absorption, distribution, metabolism, excretion, and toxicity (ADMET) properties of small-molecule drugs is critical for ensuring safety and efficacy. However, the process of accurately predicting these…

机器学习 · 计算机科学 2026-03-27 Bohao Xu , Yingzhou Lu , Chenhao Li , Ling Yue , Xiao Wang , Tianfan Fu , Minjie Shen , Lulu Chen

Modern SMILES-based chemical language models obtain strong MoleculeNet performance by treating SMILES as generic text and compensating with multi-million-molecule self-supervised pretraining. We ask: when a domain carries structural priors…

机器学习 · 计算机科学 2026-05-14 Deepak Warrier , Raja Sekhar Pappala

Purpose: Large Language Models (LLMs) like GPT (Generative Pre-trained Transformer) from OpenAI and LLaMA (Large Language Model Meta AI) from Meta AI are increasingly recognized for their potential in the field of cheminformatics,…

生物大分子 · 定量生物学 2024-05-22 Shaghayegh Sadeghi , Alan Bui , Ali Forooghi , Jianguo Lu , Alioune Ngom

Chemical Language Models (CLMs) pre-trained on large scale molecular data are widely used for molecular property prediction. However, the common belief that increasing training resources such as model size, dataset size, and training…

机器学习 · 计算机科学 2026-05-14 Tatsuya Sagawa , Ryosuke Kojima

One reason pretraining on self-supervised linguistic tasks is effective is that it teaches models features that are helpful for language understanding. However, we want pretrained models to learn not only to represent linguistic features,…

计算与语言 · 计算机科学 2020-10-13 Alex Warstadt , Yian Zhang , Haau-Sing Li , Haokun Liu , Samuel R. Bowman

Predicting the inhibitory potency of small molecules against Tyrosyl-DNA Phosphodiesterase 1 (TDP1)-a key target in overcoming cancer chemoresistance-remains a critical challenge in early drug discovery. We present a deep learning framework…

机器学习 · 计算机科学 2025-12-05 Baichuan Zeng
‹ 上一页 1 2 3 10 下一页 ›