English

Translation between Molecules and Natural Language

Computation and Language 2022-11-07 v3 Artificial Intelligence

Abstract

We present MolT5\textbf{MolT5} - a self-supervised learning framework for pretraining models on a vast amount of unlabeled natural language text and molecule strings. MolT5\textbf{MolT5} allows for new, useful, and challenging analogs of traditional vision-language tasks, such as molecule captioning and text-based de novo molecule generation (altogether: translation between molecules and language), which we explore for the first time. Since MolT5\textbf{MolT5} pretrains models on single-modal data, it helps overcome the chemistry domain shortcoming of data scarcity. Furthermore, we consider several metrics, including a new cross-modal embedding-based metric, to evaluate the tasks of molecule captioning and text-based molecule generation. Our results show that MolT5\textbf{MolT5}-based models are able to generate outputs, both molecules and captions, which in many cases are high quality.

Keywords

Cite

@article{arxiv.2204.11817,
  title  = {Translation between Molecules and Natural Language},
  author = {Carl Edwards and Tuan Lai and Kevin Ros and Garrett Honke and Kyunghyun Cho and Heng Ji},
  journal= {arXiv preprint arXiv:2204.11817},
  year   = {2022}
}

Comments

Accepted at EMNLP 2022. Data and code can be found on [Github](https://github.com/blender-nlp/MolT5)

R2 v1 2026-06-24T10:58:05.164Z