NeuroState-Bench:用于 LLM 智能体轮廓承诺完整性的人类校准基准
摘要
仅凭结果评估无法确定评估的智能体轮廓是否保留了解决多轮任务所需的承诺。NeuroState-Bench 是一个人类校准的基准,通过基准定义的侧查询探针来操作化承诺完整性,而非推断隐藏激活。所发布的库存包含 144 个确定性任务和 306 个基准定义的侧查询探针,涵盖八个认知动机下的失败家族,配对干净和干扰变体,以及三个难度带。主要的 32 个轮廓评估包含固定的 16 个轮廓本地子集和匹配的 16 个轮廓托管大型模型子集,通过相同的基准管道进行评估。人类校准使用最终合并的报告范围:104 个抽样任务单元、216 原始注释和 108 裁决任务行,with weighted kappa = 0.977 and ICC(2,1) = 0.977。Empirically, task success and commitment integrity diverge across this expanded grid: the success leader is not the integrity leader, 31 of 32 profiles change rank when integrity replaces task success, and integrity rankings are more stable under distractor perturbation. The primary confidence-free score HCCIS-CORE reaches 0.8469 AUC and 0.6992 PR-AUC for post-probe diagnostic discrimination of terminal task failure; the legacy full heuristic variant HCCIS-FULL reaches 0.7997 AUC and 0.6410 PR-AUC. Probe accuracy and state drift achieve slightly higher ROC-AUC, 0.8587, and better Brier/ECE, while HCCIS-CORE has substantially higher point-estimate PR-AUC and remains more closely tied to the benchmark's intended construct. The exploratory neural-augmented variant HCCIS+N is weaker overall, and a randomized subspace control approaches chance. NeuroState-Bench therefore contributes a calibrated evaluation axis for exposing commitment failures over a broader model grid than the original local-only subset.
引用
@article{arxiv.2605.01847,
title = {NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles},
author = {Xiao Jia},
journal= {arXiv preprint arXiv:2605.01847},
year = {2026}
}
备注
30 pages, 11 figures