English
Related papers

Related papers: Exploring Data-Driven Chemical SMILES Tokenization…

200 papers

Identification of high affinity drug-target interactions is a major research question in drug discovery. Proteins are generally represented by their structures or sequences. However, structures are available only for a small subset of…

Machine Learning · Computer Science 2020-12-22 Rıza Özçelik , Hakime Öztürk , Arzucan Özgür , Elif Ozkirimli

Tokenization is a crucial step in processing protein sequences for machine learning models, as proteins are complex sequences of amino acids that require meaningful segmentation to capture their functional and structural properties.…

Computation and Language · Computer Science 2024-11-27 Burak Suyunu , Enes Taylan , Arzucan Özgür

Complex chemical structures, like drugs, are usually defined by SMILES strings as a sequence of molecules and bonds. These SMILES strings are used in different complex machine learning-based drug-related research and representation works.…

Biomolecules · Quantitative Biology 2024-03-29 Azmine Toushik Wasi , Šerbetar Karlo , Raima Islam , Taki Hasan Rafi , Dong-Kyu Chae

Deep neural-network-based language models (LMs) are increasingly applied to large-scale protein sequence data to predict protein function. However, being largely black-box models and thus challenging to interpret, current protein LM…

Quantitative Methods · Quantitative Biology 2024-08-06 Mai Ha Vu , Rahmad Akbar , Philippe A. Robert , Bartlomiej Swiatczak , Victor Greiff , Geir Kjetil Sandve , Dag Trygve Truslew Haug

Text-based foundation models have become an important part of scientific discovery, with molecular foundation models accelerating advancements in material science and molecular design.However, existing models are constrained by…

Machine Learning · Computer Science 2026-01-29 Alexius Wadell , Anoushka Bhutani , Venkatasubramanian Viswanathan

Representing molecular structures effectively in chemistry remains a challenging task. Language models and graph-based models are extensively utilized within this domain, consistently achieving state-of-the-art results across an array of…

Machine Learning · Computer Science 2025-05-27 Nikolai Rekut , Alexey Orlov , Klea Ziu , Elizaveta Starykh , Martin Takac , Aleksandr Beznosikov

Models based on machine learning can enable accurate and fast molecular property predictions, which is of interest in drug discovery and material design. Various supervised machine learning models have demonstrated promising performance,…

Machine Learning · Computer Science 2022-12-15 Jerret Ross , Brian Belgodere , Vijil Chenthamarakshan , Inkit Padhi , Youssef Mroueh , Payel Das

AI for drug discovery has been a research hotspot in recent years, and SMILES-based language models has been increasingly applied in drug molecular design. However, no work has explored whether and how language models understand the…

Machine Learning · Computer Science 2024-01-17 Xiuyuan Hu , Guoqing Liu , Yang Zhao , Hao Zhang

Molecular property prediction is an increasingly critical task within drug discovery and development. Typically, neural networks can learn molecular properties using graph-based, language-based or feature-based methods. Recent advances in…

Machine Learning · Computer Science 2025-07-31 Philip Spence , Brooks Paige , Anne Osbourn

Recent advancements in computational chemistry have leveraged the power of trans-former-based language models, such as MoLFormer, pre-trained using a vast amount of simplified molecular-input line-entry system (SMILES) sequences, to…

Biomolecules · Quantitative Biology 2024-11-05 Tianhao Peng , Yuchen Li , Xuhong Li , Jiang Bian , Zeke Xie , Ning Sui , Shahid Mumtaz , Yanwu Xu , Linghe Kong , Haoyi Xiong

SMILES is a linear representation of chemical structures which encodes the connection table, and the stereochemistry of a molecule as a line of text with a grammar structure denoting atoms, bonds, rings and chains, and this information can…

Machine Learning · Computer Science 2018-12-03 Arindam Paul , Dipendra Jha , Reda Al-Bahrani , Wei-keng Liao , Alok Choudhary , Ankit Agrawal

The application of large language models (LLMs) to chemistry is frequently hampered by a "tokenization bottleneck", where tokenizers tuned on general-domain text tend to fragment chemical representations such as SMILES into semantically…

Computation and Language · Computer Science 2025-11-19 Prathamesh Kalamkar , Ned Letcher , Meissane Chami , Sahger Lad , Shayan Mohanty , Prasanna Pendse

The success of language models, especially transformer-based architectures, has trickled into other domains giving rise to "scientific language models" that operate on small molecules, proteins or polymers. In chemistry, language models…

Chemical Physics · Physics 2024-10-22 Nikita Janakarajan , Tim Erdmann , Sarath Swaminathan , Teodoro Laino , Jannis Born

The detailed analysis of molecular structures and properties holds great potential for drug development discovery through machine learning. Developing an emergent property in the model to understand molecules would broaden the horizons for…

Subword tokenization has become the prevailing standard in the field of natural language processing (NLP) over recent years, primarily due to the widespread utilization of pre-trained language models. This shift began with Byte-Pair…

Computation and Language · Computer Science 2024-06-11 Yanis Labrak , Adrien Bazoge , Beatrice Daille , Mickael Rouvier , Richard Dufour

Large-scale pre-training methodologies for chemical language models represent a breakthrough in cheminformatics. These methods excel in tasks such as property prediction and molecule generation by learning contextualized representations of…

Machine Learning · Computer Science 2025-07-18 Eduardo Soares , Victor Shirasuna , Emilio Vital Brazil , Renato Cerqueira , Dmitry Zubarev , Kristin Schmidt

Despite the high accuracy of 'black box' deep learning models, drug discovery still relies on protein-ligand interaction principles and heuristics. To improve interpretability of protein-small molecule binding predictions, we developed the…

Machine Learning · Computer Science 2026-04-21 Jingke Chen , Jingrui Zhong , Tazneen Hossain Tani , Zidong Su , Xiaochun Zhang , Boxue Tian

Automated computational analysis of the vast chemical space is critical for numerous fields of research such as drug discovery and material science. Representation learning techniques have recently been employed with the primary objective…

Quantitative Methods · Quantitative Biology 2023-05-26 Atakan Yüksel , Erva Ulusoy , Atabey Ünlü , Tunca Doğan

We introduce a protein language model for determining the complete sequence of a peptide based on measurement of a limited set of amino acids. To date, protein sequencing relies on mass spectrometry, with some novel edman degregation based…

Since the advent of machine learning, interpretability has remained a persistent challenge, becoming increasingly urgent as generative models support high-stakes applications in drug and material discovery. Recent advances in large language…

Machine Learning · Computer Science 2025-12-10 Jaron Cohen , Alexander G. Hasson , Sara Tanovic
‹ Prev 1 2 3 10 Next ›