English

AROMMA: Unifying Olfactory Embeddings for Single Molecules and Mixtures

Machine Learning 2026-01-28 v1 Artificial Intelligence

Abstract

Public olfaction datasets are small and fragmented across single molecules and mixtures, limiting learning of generalizable odor representations. Recent works either learn single-molecule embeddings or address mixtures via similarity or pairwise label prediction, leaving representations separate and unaligned. In this work, we propose AROMMA, a framework that learns a unified embedding space for single molecules and two-molecule mixtures. Each molecule is encoded by a chemical foundation model and the mixtures are composed by an attention-based aggregator, ensuring both permutation invariance and asymmetric molecular interactions. We further align odor descriptor sets using knowledge distillation and class-aware pseudo-labeling to enrich missing mixture annotations. AROMMA achieves state-of-the-art performance in both single-molecule and molecule-pair datasets, with up to 19.1% AUROC improvement, demonstrating a robust generalization in two domains.

Keywords

Cite

@article{arxiv.2601.19561,
  title  = {AROMMA: Unifying Olfactory Embeddings for Single Molecules and Mixtures},
  author = {Dayoung Kang and JongWon Kim and Jiho Park and Keonseock Lee and Ji-Woong Choi and Jinhyun So},
  journal= {arXiv preprint arXiv:2601.19561},
  year   = {2026}
}
R2 v1 2026-07-01T09:22:13.331Z