English
Related papers

Related papers: ChemBERTa-2: Towards Chemical Foundation Models

200 papers

Automatic medication mining from clinical and biomedical text has become a popular topic due to its real impact on healthcare applications and the recent development of powerful language models (LMs). However, fully-automatic extraction…

Computation and Language · Computer Science 2023-08-09 Haifa Alrdahi , Lifeng Han , Hendrik Šuvalov , Goran Nenadic

Molecular property prediction is crucial for drug discovery and materials science, yet existing approaches suffer from limited interpretability, poor cross-task generalization, and lack of chemical reasoning capabilities. Traditional…

Machine Learning · Computer Science 2025-10-20 Jiaxi Zhuang , Yaorui Shi , Jue Hou , Yunong He , Mingwei Ye , Mingjun Xu , Yuming Su , Linfeng Zhang , Ying Qian , Linfeng Zhang , Guolin Ke , Hengxing Cai

The simplified molecular-input line-entry system (SMILES) is the most popular representation of chemical compounds. Therefore, many SMILES-based molecular property prediction models have been developed. In particular, transformer-based…

Quantitative Methods · Quantitative Biology 2022-05-03 Ingoo Lee , Hojung Nam

Transformer-based large language models have remarkable potential to accelerate design optimization for applications such as drug development and materials discovery. Self-supervised pretraining of transformer models requires large-scale…

Machine Learning · Computer Science 2023-10-27 Pei Zhang , Logan Kearney , Debsindhu Bhowmik , Zachary Fox , Amit K. Naskar , John Gounley

Chemical pretrained models, sometimes referred to as foundation models, are receiving considerable interest for drug discovery applications. The general chemical knowledge extracted from self-supervised training has the potential to improve…

Machine Learning · Computer Science 2025-10-15 Matthew Adrian , Yunsie Chung , Kevin Boyd , Saee Paliwal , Srimukh Prasad Veccham , Alan C. Cheng

Generative pre-trained Transformer (GPT) has demonstrates its great success in natural language processing and related techniques have been adapted into molecular modeling. Considering that text is the most important record for scientific…

Computation and Language · Computer Science 2023-05-29 Zequn Liu , Wei Zhang , Yingce Xia , Lijun Wu , Shufang Xie , Tao Qin , Ming Zhang , Tie-Yan Liu

Chemistry plays a crucial role in many domains, such as drug discovery and material science. While large language models (LLMs) such as GPT-4 exhibit remarkable capabilities on natural language processing tasks, existing research indicates…

Artificial Intelligence · Computer Science 2024-08-13 Botao Yu , Frazier N. Baker , Ziqi Chen , Xia Ning , Huan Sun

In recent years, researchers tend to pre-train ever-larger language models to explore the upper limit of deep models. However, large language model pre-training costs intensive computational resources and most of the models are trained from…

Computation and Language · Computer Science 2021-10-15 Cheng Chen , Yichun Yin , Lifeng Shang , Xin Jiang , Yujia Qin , Fengyu Wang , Zhi Wang , Xiao Chen , Zhiyuan Liu , Qun Liu

Artificial intelligence and machine learning have shown great promise in their ability to accelerate novel materials discovery. As researchers and domain scientists seek to unify and consolidate chemical knowledge, the case for models with…

Representing molecular structures effectively in chemistry remains a challenging task. Language models and graph-based models are extensively utilized within this domain, consistently achieving state-of-the-art results across an array of…

Machine Learning · Computer Science 2025-05-27 Nikolai Rekut , Alexey Orlov , Klea Ziu , Elizaveta Starykh , Martin Takac , Aleksandr Beznosikov

Automated computational analysis of the vast chemical space is critical for numerous fields of research such as drug discovery and material science. Representation learning techniques have recently been employed with the primary objective…

Quantitative Methods · Quantitative Biology 2023-05-26 Atakan Yüksel , Erva Ulusoy , Atabey Ünlü , Tunca Doğan

We discover a robust self-supervised strategy tailored towards molecular representations for generative masked language models through a series of tailored, in-depth ablations. Using this pre-training strategy, we train BARTSmiles, a…

Natural language understanding has recently seen a surge of progress with the use of sentence encoders like ELMo (Peters et al., 2018a) and BERT (Devlin et al., 2019) which are pretrained on variants of language modeling. We conduct the…

Multimodal large language models (MLLMs) have made impressive progress in many applications in recent years. However, chemical MLLMs that can handle cross-modal understanding and generation remain underexplored. To fill this gap, we propose…

Machine Learning · Computer Science 2025-08-05 Qian Tan , Dongzhan Zhou , Peng Xia , Wanhao Liu , Wanli Ouyang , Lei Bai , Yuqiang Li , Tianfan Fu

We demonstrate the ability of large language models (LLMs) to perform material and molecular property regression tasks, a significant deviation from the conventional LLM use case. We benchmark the Large Language Model Meta AI (LLaMA) 3 on…

Materials Science · Physics 2026-04-22 Ryan Jacobs , Maciej P. Polak , Lane E. Schultz , Hamed Mahdavi , Vasant Honavar , Dane Morgan

Transformer-based masked language models such as BERT, trained on general corpora, have shown impressive performance on downstream tasks. It has also been demonstrated that the downstream task performance of such models can be improved by…

Computation and Language · Computer Science 2023-05-04 Zhi Hong , Aswathy Ajith , Gregory Pauloski , Eamon Duede , Kyle Chard , Ian Foster

Machine learning has transformed material discovery for inorganic compounds and small molecules, yet polymers remain largely inaccessible to these methods. While data scarcity is often cited as the primary bottleneck, we demonstrate that…

Machine Learning · Computer Science 2025-12-09 Jihun Ahn , Gabriella Pasya Irianti , Vikram Thapar , Su-Mi Hur

Identification of high affinity drug-target interactions is a major research question in drug discovery. Proteins are generally represented by their structures or sequences. However, structures are available only for a small subset of…

Machine Learning · Computer Science 2020-12-22 Rıza Özçelik , Hakime Öztürk , Arzucan Özgür , Elif Ozkirimli

Pre-trained Language Models (PLMs) have been successful for a wide range of natural language processing (NLP) tasks. The state-of-the-art of PLMs, however, are extremely large to be used on edge devices. As a result, the topic of model…

Large protein language models are adept at capturing the underlying evolutionary information in primary structures, offering significant practical value for protein engineering. Compared to natural language models, protein amino acid…

Computation and Language · Computer Science 2023-10-27 Yang Tan , Mingchen Li , Pan Tan , Ziyi Zhou , Huiqun Yu , Guisheng Fan , Liang Hong