English

Beyond the All-in-One Agent: Benchmarking Role-Specialized Multi-Agent Collaboration in Enterprise Workflows

Multiagent Systems 2026-05-12 v1 Machine Learning

Abstract

Large language model (LLM) agents are increasingly expected to operate in enterprise environments, where work is distributed across specialized roles, permission-controlled systems, and cross-departmental procedures. However, existing enterprise benchmarks largely evaluate single agents with broad tool access, while existing multi-agent benchmarks rarely capture realistic enterprise constraints such as role specialization, access control, stateful business systems, and policy-based approvals. We introduce \textsc{EntCollabBench}, a benchmark for evaluating enterprise multi-agent collaboration. \textsc{EntCollabBench} simulates a permission-isolated organization with 11 role-specialized agents across six departments and contains two evaluation subsets: a Workflow subset, where agents collaboratively modify enterprise system states, and an Approval subset, where agents make policy-grounded decisions. Evaluation is based on execution traces, database state verification, and deterministic policy adjudication rather than natural-language response judging. Experiments with representative LLM agents show that current models still struggle with end-to-end enterprise collaboration, especially in delegation, context transfer, parameter grounding, workflow closure, and decision commitment. \textsc{EntCollabBench} provides a reproducible testbed for measuring and improving agent systems intended for realistic organizational environments.

Keywords

Cite

@article{arxiv.2605.08761,
  title  = {Beyond the All-in-One Agent: Benchmarking Role-Specialized Multi-Agent Collaboration in Enterprise Workflows},
  author = {Tao Yu and Hao Wang and Changyu Li and Shenghua Chai and Minghui Zhang and Zhongtian Luo and Yuxuan Zhou and Haopeng Jin and Zhaolu Kang and Jiabing Yang and YiFan Zhang and Xinming Wang and Hongzhu Yi and Zheqi He and Jing-Shu Zheng and Xi Yang and Yan Huang and Liang Wang},
  journal= {arXiv preprint arXiv:2605.08761},
  year   = {2026}
}

Comments

45 pages

R2 v1 2026-07-01T12:59:38.417Z