English

Local-Global Multimodal Contrastive Learning for Molecular Property Prediction

Machine Learning 2026-02-02 v1 Artificial Intelligence

Abstract

Accurate molecular property prediction requires integrating complementary information from molecular structure and chemical semantics. In this work, we propose LGM-CL, a local-global multimodal contrastive learning framework that jointly models molecular graphs and textual representations derived from SMILES and chemistry-aware augmented texts. Local functional group information and global molecular topology are captured using AttentiveFP and Graph Transformer encoders, respectively, and aligned through self-supervised contrastive learning. In addition, chemically enriched textual descriptions are contrasted with original SMILES to incorporate physicochemical semantics in a task-agnostic manner. During fine-tuning, molecular fingerprints are further integrated via Dual Cross-attention multimodal fusion. Extensive experiments on MoleculeNet benchmarks demonstrate that LGM-CL achieves consistent and competitive performance across both classification and regression tasks, validating the effectiveness of unified local-global and multimodal representation learning.

Keywords

Cite

@article{arxiv.2601.22610,
  title  = {Local-Global Multimodal Contrastive Learning for Molecular Property Prediction},
  author = {Xiayu Liu and Zhengyi Lu and Yunhong Liao and Chan Fan and Hou-biao Li},
  journal= {arXiv preprint arXiv:2601.22610},
  year   = {2026}
}

Comments

16 pages, 9 figures. Submitted to Briefings in Bioinformatics

R2 v1 2026-07-01T09:27:12.505Z