English
Related papers

Related papers: Exploring Data-Driven Chemical SMILES Tokenization…

200 papers

Tokenization is the act of breaking down text into smaller parts, or tokens, that are easier for machines to process. This is a key phase in machine translation (MT) models. Subword tokenization enhances this process by breaking down words…

Computation and Language · Computer Science 2025-05-23 Sudhansu Bala Das , Samujjal Choudhury , Tapas Kumar Mishra , Bidyut Kr. Patra

Traditional molecular string representations, such as SMILES, often pose challenges for AI-driven molecular design due to their non-sequential depiction of molecular substructures. To address this issue, we introduce Sequential…

Machine Learning · Computer Science 2023-12-13 Emmanuel Noutahi , Cristian Gabellini , Michael Craig , Jonathan S. C Lim , Prudencio Tossou

Despite their ability to understand chemical knowledge, large language models (LLMs) remain limited in their capacity to propose novel molecules with desired functions (e.g., drug-like properties). In addition, the molecules that LLMs…

Chemical databases store information in text representations, and the SMILES format is a universal standard used in many cheminformatics software. Encoded in each SMILES string is structural information that can be used to predict complex…

Machine Learning · Statistics 2018-08-16 Garrett B. Goh , Nathan O. Hodas , Charles Siegel , Abhinav Vishnu

This paper shows how one can directly apply natural language processing (NLP) methods to classification problems in cheminformatics. Connection between these seemingly separate fields is shown by considering standard textual representation…

Computation and Language · Computer Science 2018-03-09 Stanisław Jastrzębski , Damian Leśniak , Wojciech Marian Czarnecki

The simplified molecular-input line-entry system (SMILES) is the most popular representation of chemical compounds. Therefore, many SMILES-based molecular property prediction models have been developed. In particular, transformer-based…

Quantitative Methods · Quantitative Biology 2022-05-03 Ingoo Lee , Hojung Nam

Artificial intelligence (AI) and machine learning (ML) are expanding in popularity for broad applications to challenging tasks in chemistry and materials science. Examples include the prediction of properties, the discovery of new reaction…

Motivation: Prediction of the interaction affinity between proteins and compounds is a major challenge in the drug discovery process. WideDTA is a deep-learning based prediction model that employs chemical and biological textual sequence…

Quantitative Methods · Quantitative Biology 2019-02-13 Hakime Öztürk , Elif Ozkirimli , Arzucan Özgür

Chemical language models (CLMs) are prominent for their effectiveness in exploring chemical space and enabling molecular engineering. However, while exploring chemical-linguistic space, CLMs suffer from the gap between natural language and…

Computational Engineering, Finance, and Science · Computer Science 2025-01-07 Liuzhenghao Lv , Hao Li , Yu Wang , Zhiyuan Yan , Zijun Chen , Zongying Lin , Li Yuan , Yonghong Tian

Protein structure tokenization converts 3D structures into discrete or vectorized representations, enabling the integration of structural and sequence data. Despite many recent works on structure tokenization, the properties of the…

Machine Learning · Computer Science 2025-11-14 Zijing Liu , Bin Feng , He Cao , Yu Li

Recent dynamic tokenisation methods operate directly on bytes and pool their latent representations into patches. This bears similarities to computational models of word segmentation that determine lexical boundaries using spikes in an…

Computation and Language · Computer Science 2025-06-24 Zébulon Goriely , Suchir Salhan , Pietro Lesci , Julius Cheng , Paula Buttery

As a cornerstone in language modeling, tokenization involves segmenting text inputs into pre-defined atomic units. Conventional statistical tokenizers often disrupt constituent boundaries within words, thereby corrupting semantic…

Computation and Language · Computer Science 2025-07-11 Qingyang Zhu , Xiang Hu , Pengyu Ji , Wei Wu , Kewei Tu

What are the units of text that we want to model? From bytes to multi-word expressions, text can be analyzed and generated at many granularities. Until recently, most natural language processing (NLP) models operated over words, treating…

Subword tokenization methods, such as Byte-Pair Encoding (BPE), significantly impact the performance and efficiency of large language models (LLMs). The standard approach involves training a general-purpose tokenizer that uniformly…

Computation and Language · Computer Science 2026-01-30 Vijini Liyanage , François Yvon

Language models are powerful tools for molecular design. Currently, the dominant paradigm is to parse molecular graphs into linear string representations that can easily be trained on. This approach has been very successful, however, it is…

Machine Learning · Computer Science 2023-05-11 Daniel Flam-Shepherd , Alán Aspuru-Guzik

In drug-discovery-related tasks such as virtual screening, machine learning is emerging as a promising way to predict molecular properties. Conventionally, molecular fingerprints (numerical representations of molecules) are calculated…

Machine Learning · Computer Science 2019-11-13 Shion Honda , Shoi Shi , Hiroki R. Ueda

Signaling proteins are an important topic in drug development due to the increased importance of finding fast, accurate and cheap methods to evaluate new molecular targets involved in specific diseases. The complexity of the protein…

Quantitative Methods · Quantitative Biology 2019-04-11 Carlos Fernandez-Lozano , Ruben F. Cuinas , Jose A. Seoane , Enrique Fernandez-Blanco , Julian Dorado , Cristian R. Munteanu

In recent years machine learning (ML) took bio- and cheminformatics fields by storm, providing new solutions for a vast repertoire of problems related to protein sequence, structure, and interactions analysis. ML techniques, deep neural…

Biomolecules · Quantitative Biology 2020-03-31 Marta M. Stepniewska-Dziubinska , Piotr Zielenkiewicz , Pawel Siedlecki

Understanding protein sequences is vital and urgent for biology, healthcare, and medicine. Labeling approaches are expensive yet time-consuming, while the amount of unlabeled data is increasing quite faster than that of the labeled data due…

Computation and Language · Computer Science 2021-11-01 Liang He , Shizhuo Zhang , Lijun Wu , Huanhuan Xia , Fusong Ju , He Zhang , Siyuan Liu , Yingce Xia , Jianwei Zhu , Pan Deng , Bin Shao , Tao Qin , Tie-Yan Liu

In recent years, machine learning has been proposed as a promising strategy to build accurate scoring functions for computational docking finalized to numerically empowered drug discovery. However, the latest studies have suggested that…

Quantitative Methods · Quantitative Biology 2023-02-17 F. Pellicani , D. Dal Ben , A. Perali , S. Pilati