English

FactoryNet: A Large-Scale Dataset toward Industrial Time-Series Foundation Models

Machine Learning 2026-05-26 v3 Artificial Intelligence

Abstract

We introduce the first universal pretraining corpus for industrial time-series data: FactoryNet. 51M datapoints across 23k end-to-end task executions (13.3k real, 9.8k synthetic) on six embodiments, unified by a shared schema that enables robust zero-shot cross-embodiment transfer and highly parameter-efficient anomaly detection. We introduce a novel schema: Setpoint, Effort, Feedback, Context (S-E-F-C) underlying the whole pipeline that maps any actuated system into a common representational frame. The corpus spans 27 annotated anomaly types alongside healthy baselines and counterfactual pairs across robotic manipulation and machining domains. Cross-embodiment transfer experiments yield positive results: under bias-aware metrics our model demonstrates fair cross-embodiment transfer capabilities on the evaluated source-target pair, while 24 schema-aligned signals achieves competitive anomaly detection performance compared to high-dimensional baselines. We release FactoryNet as a growing, multi-embodiment dataset to drive progress toward industrial foundation models.

Keywords

Cite

@article{arxiv.2605.09081,
  title  = {FactoryNet: A Large-Scale Dataset toward Industrial Time-Series Foundation Models},
  author = {Karim Othman and Jonas Petersen and Matei Ignuta-Ciuncanu and Camilla Mazzoleni and Federico Martelli and Alessandro Lombardi and Riccardo Maggioni and Philipp Petersen},
  journal= {arXiv preprint arXiv:2605.09081},
  year   = {2026}
}

Comments

8 pages, 4 figures, 5 tables

R2 v1 2026-07-01T13:00:15.236Z