English

FungalZSL: Zero-Shot Fungal Classification with Image Captioning Using a Synthetic Data Approach

Computer Vision and Pattern Recognition 2025-02-27 v1

Abstract

The effectiveness of zero-shot classification in large vision-language models (VLMs), such as Contrastive Language-Image Pre-training (CLIP), depends on access to extensive, well-aligned text-image datasets. In this work, we introduce two complementary data sources, one generated by large language models (LLMs) to describe the stages of fungal growth and another comprising a diverse set of synthetic fungi images. These datasets are designed to enhance CLIPs zero-shot classification capabilities for fungi-related tasks. To ensure effective alignment between text and image data, we project them into CLIPs shared representation space, focusing on different fungal growth stages. We generate text using LLaMA3.2 to bridge modality gaps and synthetically create fungi images. Furthermore, we investigate knowledge transfer by comparing text outputs from different LLM techniques to refine classification across growth stages.

Keywords

Cite

@article{arxiv.2502.19038,
  title  = {FungalZSL: Zero-Shot Fungal Classification with Image Captioning Using a Synthetic Data Approach},
  author = {Anju Rani and Daniel O. Arroyo and Petar Durdevic},
  journal= {arXiv preprint arXiv:2502.19038},
  year   = {2025}
}

Comments

11 pages, 5 Figures, 1 Table

R2 v1 2026-06-28T21:58:32.715Z