English
Related papers

Related papers: SMILES Enumeration as Data Augmentation for Neural…

200 papers

A statistical model is a mathematical representation of an often simplified or idealised data-generating process. In this paper, we focus on a particular type of statistical model, called linear mixed models (LMMs), that is widely used in…

Methodology · Statistics 2020-01-23 Emi Tanaka , Francis K. C. Hui

The application of large language models (LLMs) to chemistry is frequently hampered by a "tokenization bottleneck", where tokenizers tuned on general-domain text tend to fragment chemical representations such as SMILES into semantically…

Computation and Language · Computer Science 2025-11-19 Prathamesh Kalamkar , Ned Letcher , Meissane Chami , Sahger Lad , Shayan Mohanty , Prasanna Pendse

Like many scientific fields, new chemistry literature has grown at a staggering pace, with thousands of papers released every month. A large portion of chemistry literature focuses on new molecules and reactions between molecules. Most…

Quantitative Methods · Quantitative Biology 2021-09-10 Daniel Campos , Heng Ji

Large Language Models (LLMs) are increasingly being used to support scientific discovery. In chemistry, tasks such as reaction prediction and structure elucidation require reasoning about the structures of molecules. As such, LLM-based…

Machine Learning · Computer Science 2026-05-05 Nicholas T. Runcie , Fergus Imrie , Charlotte M. Deane

Molecular language modeling tasks such as molecule captioning have been recognized for their potential to further understand molecular properties that can aid drug discovery or material synthesis based on chemical reactions. Unlike the…

Machine Learning · Computer Science 2025-03-12 Sangyeup Kim , Nayeon Kim , Yinhua Piao , Sun Kim

Model ensembling is a well-established technique for improving the performance of machine learning models. Conventionally, this involves averaging the output distributions of multiple models and selecting the most probable label. This idea…

Machine Learning · Computer Science 2026-05-26 Jiale Fu , Yuchu Jiang , Peijun Wu , Chonghan Liu , Joey Tianyi Zhou , Xu Yang

The Saliency Model Implementation Library for Experimental Research (SMILER) is a new software package which provides an open, standardized, and extensible framework for maintaining and executing computational saliency models. This work…

Computer Vision and Pattern Recognition · Computer Science 2018-12-24 Calden Wloka , Toni Kunić , Iuliia Kotseruba , Ramin Fahimi , Nicholas Frosst , Neil D. B. Bruce , John K. Tsotsos

The development of large language models and multi-modal models has enabled the appealing idea of generating novel molecules from text descriptions. Generative modeling would shift the paradigm from relying on large-scale chemical screening…

Machine Learning · Computer Science 2025-08-25 Yifan Deng , Spencer S. Ericksen , Anthony Gitter

Molecular representation is a critical element in our understanding of the physical world and the foundation for modern molecular machine learning. Previous molecular machine learning models have employed strings, fingerprints, global…

Machine Learning · Computer Science 2025-05-28 Daniil A. Boiko , Thiago Reschützegger , Benjamin Sanchez-Lengeling , Samuel M. Blau , Gabe Gomes

Smishing, which aims to illicitly obtain personal information from unsuspecting victims, holds significance due to its negative impacts on our society. In prior studies, as a tool to counteract smishing, machine learning (ML) has been…

Social and Information Networks · Computer Science 2024-11-07 Ho Sung Shim , Hyoungjun Park , Kyuhan Lee , Jang-Sun Park , Seonhye Kang

Biological multimodal large language models (MLLMs) have emerged as powerful foundation models for scientific discovery. However, existing models are specialized to a single modality, limiting their ability to solve inherently cross-modal…

Machine Learning · Computer Science 2026-03-17 Wonbin Lee , Dongki Kim , Sung Ju Hwang

Mammalian cells have about 30,000-fold more protein molecules than mRNA molecules. This larger number of molecules and the associated larger dynamic range have major implications in the development of proteomics technologies. We examine…

Quantitative Methods · Quantitative Biology 2023-03-14 Michael J. MacCoss , Javier Alfaro , Meni Wanunu , Danielle A. Faivre , Nikolai Slavov

In nature, the behaviors of many complex systems can be described by parsimonious math equations. Automatically distilling these equations from limited data is cast as a symbolic regression process which hitherto remains a grand challenge.…

Machine Learning · Computer Science 2023-05-25 Yilong Xu , Yang Liu , Hao Sun

The potential number of drug like small molecules is estimated to be between 10^23 and 10^60 while current databases of known compounds are orders of magnitude smaller with approximately 10^8 compounds. This discrepancy has led to an…

Machine Learning · Computer Science 2017-05-18 Esben Jannik Bjerrum , Richard Threlfall

Generative pre-trained Transformer (GPT) has demonstrates its great success in natural language processing and related techniques have been adapted into molecular modeling. Considering that text is the most important record for scientific…

Computation and Language · Computer Science 2023-05-29 Zequn Liu , Wei Zhang , Yingce Xia , Lijun Wu , Shufang Xie , Tao Qin , Ming Zhang , Tie-Yan Liu

The quantum Hamiltonian is a fundamental property that governs a molecule's electronic structure and behavior, and its calculation and prediction are paramount in computational chemistry and materials science. Accurate prediction is highly…

Computational Engineering, Finance, and Science · Computer Science 2026-01-23 Zhenzhong Wang , Yongjie Hou , Chenggong Huang , Yuxuan Du , Dacheng Tao , Min Jiang

Cross-modal retrieval between food images and recipe texts is an important task with applications in nutritional management, dietary logging, and cooking assistance. Existing methods predominantly rely on dual-encoder architectures with…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Keisuke Gomi , Keiji Yanai

Machine learning has transformed material discovery for inorganic compounds and small molecules, yet polymers remain largely inaccessible to these methods. While data scarcity is often cited as the primary bottleneck, we demonstrate that…

Machine Learning · Computer Science 2025-12-09 Jihun Ahn , Gabriella Pasya Irianti , Vikram Thapar , Su-Mi Hur

Before entering the neural network, a token is generally converted to the corresponding one-hot representation, which is a discrete distribution of the vocabulary. Smoothed representation is the probability of candidate tokens obtained from…

Computation and Language · Computer Science 2022-03-01 Xing Wu , Chaochen Gao , Meng Lin , Liangjun Zang , Zhongyuan Wang , Songlin Hu

Identification of high affinity drug-target interactions is a major research question in drug discovery. Proteins are generally represented by their structures or sequences. However, structures are available only for a small subset of…

Machine Learning · Computer Science 2020-12-22 Rıza Özçelik , Hakime Öztürk , Arzucan Özgür , Elif Ozkirimli