English

TRIDENT: Tri-Modal Molecular Representation Learning with Taxonomic Annotations and Local Correspondence

Machine Learning 2026-03-03 v2

Abstract

Molecular property prediction aims to learn representations that map chemical structures to functional properties. While multimodal learning has emerged as a powerful paradigm to learn molecular representations, prior works have largely overlooked textual and taxonomic information of molecules for representation learning. We introduce TRIDENT, a novel framework that integrates molecular SMILES, textual descriptions, and taxonomic functional annotations to learn rich molecular representations. To achieve this, we curate a comprehensive dataset of molecule-text pairs with structured, multi-level functional annotations. Instead of relying on conventional contrastive loss, TRIDENT employs a volume-based alignment objective to jointly align tri-modal features at the global level, enabling soft, geometry-aware alignment across modalities. Additionally, TRIDENT introduces a novel local alignment objective that captures detailed relationships between molecular substructures and their corresponding sub-textual descriptions. A momentum-based mechanism dynamically balances global and local alignment, enabling the model to learn both broad functional semantics and fine-grained structure-function mappings. TRIDENT achieves state-of-the-art performance on 11 downstream tasks, demonstrating the value of combining SMILES, textual, and taxonomic functional annotations for molecular property prediction.

Keywords

Cite

@article{arxiv.2506.21028,
  title  = {TRIDENT: Tri-Modal Molecular Representation Learning with Taxonomic Annotations and Local Correspondence},
  author = {Feng Jiang and Mangal Prakash and Hehuan Ma and Jianyuan Deng and Yuzhi Guo and Amina Mollaysa and Tommaso Mansi and Rui Liao and Junzhou Huang},
  journal= {arXiv preprint arXiv:2506.21028},
  year   = {2026}
}

Comments

Accepted to NeurIPS 2025

R2 v1 2026-07-01T03:34:05.313Z