English

FieldWorkArena: Agentic AI Benchmark for Real Field Work Tasks

Artificial Intelligence 2026-04-16 v3 Computer Vision and Pattern Recognition

Abstract

This paper introduces FieldWorkArena, a benchmark for agentic AI targeting real-world field work. With the recent increase in demand for agentic AI, they are built to detect and document safety hazards, procedural violations, and other critical incidents across real-world manufacturing and retail environments. Whereas most agentic AI benchmarks focus on performance in simulated or digital environments, our work addresses the fundamental challenge of evaluating agents in the real-world. In this paper, we improve the evaluation function from previous methods to assess the performance of agentic AI in diverse real-world tasks. Our dataset comprises on-site captured images/videos in factories, warehouses and retails. Tasks were meticulously developed through interviews with site workers and managers. Evaluation results confirmed that performance evaluation considering the characteristics of Multimodal LLM (MLLM) such as GPT-4o is feasible. Furthermore, this study identifies both the effectiveness and limitations of the proposed new evaluation methodology. The complete dataset and evaluation program are publicly accessible on the website (https://en-documents.research.global.fujitsu.com/fieldworkarena/)

Keywords

Cite

@article{arxiv.2505.19662,
  title  = {FieldWorkArena: Agentic AI Benchmark for Real Field Work Tasks},
  author = {Jun Takahashi and Atsunori Moteki and Akiyoshi Uchida and Shoichi Masui and Fan Yang and Kanji Uchino and Yueqi Song and Yonatan Bisk and Graham Neubig and Ikuo Kusajima and Yasuto Watanabe and Hiroyuki Ishida and Koki Nakagawa and Shan Jiang},
  journal= {arXiv preprint arXiv:2505.19662},
  year   = {2026}
}

Comments

15 pages, 2 figures, 5 tables [ICPR 2026 Accepted] Changes from Version 2: 1) Added retail domain as third scenario; dataset scaled from 455 to 886 tasks, 2) Task taxonomy restructured (Planning/Perception/Action -> Perception/Decision Making/Combination), 3) Experiments updated: GPT-5.1/5.2, Gemini 2.5 Flash/Pro (replaced Claude); added human baseline and video chunking/Qwen3-VL experiments

R2 v1 2026-07-01T02:38:43.402Z