HAS-Bench: Evaluating LLM-Based Human-Agent Systems under Configurable Human Participation
Abstract
Large language models increasingly operate in settings where humans are active collaborators rather than passive task providers. We introduce HAS-Framework, a graph-based framework that represents humans and LLM-powered agents as first-class participants with explicit roles, permissions, communication paths, and action authority. Building on this framework, HAS-Bench evaluates Human-Agent Systems under configurable human participation across agency levels, interaction channels, and persona policies. The benchmark measures both task outcomes and process-level collaboration behavior, including clarification quality, feedback utilization, control calibration, safety, initiative, and interaction cost. Experiments across six domains show that human participation can substantially improve task completion and failure recovery, but the gains depend on when, how, and by whom human input is exercised.
Cite
@article{arxiv.2607.04329,
title = {HAS-Bench: Evaluating LLM-Based Human-Agent Systems under Configurable Human Participation},
author = {Yaozu Wu and Wei-Chieh Huang and Jizhou Guo and Dongyuan Li and Renhe Jiang and Henry Peng Zou and Chunyu Miao and Shanghao Li and Weizhi Zhang and WeiWei Ye and Yankai Chen and Meng Zhang and Xue Liu and Philip S. Yu},
journal= {arXiv preprint arXiv:2607.04329},
year = {2026}
}