English

SPROUT: A Scalable Diffusion Foundation Model for Agricultural Vision

Computer Vision and Pattern Recognition 2026-03-31 v1

Abstract

Vision Foundation Models (VFM) pre-trained on large-scale unlabeled data have achieved remarkable success on general computer vision tasks, yet typically suffer from significant domain gaps when applied to agriculture. In this context, we introduce SPROUTSPROUT (SScalable PPlant RRepresentation model via OOpen-field UUnsupervised TTraining), a multi-crop, multi-task agricultural foundation model trained via diffusion denoising. SPROUT leverages a VAE-free Pixel-space Diffusion Transformer to learn rich, structure-aware representations through denoising and enabling efficient end-to-end training. We pre-train SPROUT on a curated dataset of 2.6 million high-quality agricultural images spanning diverse crops, growth stages, and environments. Extensive experiments demonstrate that SPROUT consistently outperforms state-of-the-art web-pretrained and agricultural foundation models across a wide range of downstream tasks, while requiring substantially lower pre-training cost. The code and model are available at https://github.com/UTokyo-FieldPhenomics-Lab/SPROUT.

Keywords

Cite

@article{arxiv.2603.27519,
  title  = {SPROUT: A Scalable Diffusion Foundation Model for Agricultural Vision},
  author = {Shuai Xiang and Wei Guo and James Burridge and Shouyang Liu and Hao Lu and Tokihiro Fukatsu},
  journal= {arXiv preprint arXiv:2603.27519},
  year   = {2026}
}
R2 v1 2026-07-01T11:42:39.540Z