English

FactoryBench: Evaluating Industrial Machine Understanding

Artificial Intelligence 2026-05-11 v1 Machine Learning

Abstract

We introduce FactoryBench, a benchmark for evaluating time-series models and LLMs on machine understanding over industrial robotic telemetry. Q&A pairs are organized along four causal levels (state, intervention, counterfactual, decision) instantiating Pearl's ladder of causation, and span five answer formats: four structured formats are scored deterministically and free-form answers are scored by an LLM-as-judge voting protocol. We propose a scalable Q&A generation framework built around structured question templates, present FactoryWave (a dense, multitask, multivariate sensor dataset collected from a UR3 cobot and a KUKA KR10 industrial arm), and construct FactoryBench as a large-scale benchmark of over 70k Q&A items grounded in roughly 15k normalized episodes from FactoryWave, AURSAD, and voraus-AD. Zero-shot evaluation of six frontier LLMs shows that no model exceeds 50% on structured levels or 18% on decision-making, revealing a wide gap between current models and operational machine understanding.

Keywords

Cite

@article{arxiv.2605.07675,
  title  = {FactoryBench: Evaluating Industrial Machine Understanding},
  author = {Yanis Merzouki and Coral Izquierdo and Matei Ignuta-Ciuncanu and Marcos Gomez-Bracamonte and Riccardo Maggioni and Alessandro Lombardi and Camilla Mazzoleni and Federico Martelli and Balazs Gunther and Jonas Petersen and Philipp Petersen},
  journal= {arXiv preprint arXiv:2605.07675},
  year   = {2026}
}

Comments

9 pages, 4 figures, 14 tables; appendix with 24 pages

R2 v1 2026-07-01T12:57:39.826Z