English

AgroTools: A Benchmark for Tool-Augmented Multimodal Agents in Agriculture

Computer Vision and Pattern Recognition 2026-05-22 v1

Abstract

Agricultural decision-making increasingly requires multimodal systems that can transform visual observations into reliable, executable actions. However, existing agricultural multimodal benchmarks mainly evaluate final-answer correctness and provide limited support for assessing whether models can use external tools to complete precision-sensitive workflows. In this paper, we introduce AgroTools, a benchmark for evaluating tool-augmented multimodal agents in agriculture. AgroTools contains 539 question-answer instances paired with 1,097 heterogeneous agricultural images, spanning five task families and an executable environment of 14 agricultural tools. Each query is annotated with structured tool-use traces, enabling a dual-view evaluation of both process-level execution quality and outcome-level task success. We benchmark 9 open-source and 4 closed-source multimodal large language models on AgroTools. Results show that current models remain far from reliable in agricultural tool-use settings, with clear bottlenecks in tool planning, argument generation, execution recovery, and final-answer synthesis. We hope AgroTools will support future research on multimodal agents for high-precision agricultural applications. The benchmark and evaluation are available at https://huggingface.co/datasets/AgroTools/AgroTools.

Keywords

Cite

@article{arxiv.2605.22366,
  title  = {AgroTools: A Benchmark for Tool-Augmented Multimodal Agents in Agriculture},
  author = {Zi Ye and Yibin Wen and Xiaoya Fan and Xinyu Zhang and Jing Wu and Kun Zeng and Zurong Mai and Jiarui Zhang and Bohan Shi and Juepeng Zheng and Jianxi Huang and Yutong Lu and Haohuan Fu},
  journal= {arXiv preprint arXiv:2605.22366},
  year   = {2026}
}
R2 v1 2026-07-22T07:26:04.188Z