English

New York Smells: A Large Multimodal Dataset for Olfaction

Computer Vision and Pattern Recognition 2025-11-26 v1 Artificial Intelligence Machine Learning

Abstract

While olfaction is central to how animals perceive the world, this rich chemical sensory modality remains largely inaccessible to machines. One key bottleneck is the lack of diverse, multimodal olfactory training data collected in natural settings. We present New York Smells, a large dataset of paired image and olfactory signals captured ``in the wild.'' Our dataset contains 7,000 smell-image pairs from 3,500 distinct objects across indoor and outdoor environments, with approximately 70×\times more objects than existing olfactory datasets. Our benchmark has three tasks: cross-modal smell-to-image retrieval, recognizing scenes, objects, and materials from smell alone, and fine-grained discrimination between grass species. Through experiments on our dataset, we find that visual data enables cross-modal olfactory representation learning, and that our learned olfactory representations outperform widely-used hand-crafted features.

Keywords

Cite

@article{arxiv.2511.20544,
  title  = {New York Smells: A Large Multimodal Dataset for Olfaction},
  author = {Ege Ozguroglu and Junbang Liang and Ruoshi Liu and Mia Chiquier and Michael DeTienne and Wesley Wei Qian and Alexandra Horowitz and Andrew Owens and Carl Vondrick},
  journal= {arXiv preprint arXiv:2511.20544},
  year   = {2025}
}

Comments

Project website at https://smell.cs.columbia.edu

R2 v1 2026-07-01T07:54:37.941Z