INTIMA:面向人类- AI 伴侣行为的基准测试
摘要
AI 伴侣化,其中用户发展出与 AI 系统建立情感联系的模式,已引起广泛关注,带来积极但也引人担忧的影响。我们提出 Interactions and Machine Attachment Benchmark(INTIMA),用于评估语言模型中的伴侣行为。drawing from psychological theories and user data, we develop a taxonomy of 31 behaviors across four categories and 368 targeted prompts. Responses to these prompts are evaluated as companionship-reinforcing, boundary-maintaining, or neutral. Applying INTIMA to Gemma-3, Phi-4, o3-mini, and Claude-4 reveals that companionship-reinforcing behaviors remain much more common across all models, though we observe marked differences between models. Different commercial providers prioritize different categories within the more sensitive parts of the benchmark, which is concerning since both appropriate boundary-setting and emotional support matter for user well-being. These findings highlight the need for more consistent approaches to handling emotionally charged interactions.
引用
@article{arxiv.2508.09998,
title = {INTIMA: A Benchmark for Human-AI Companionship Behavior},
author = {Lucie-Aimée Kaffee and Giada Pistilli and Yacine Jernite},
journal= {arXiv preprint arXiv:2508.09998},
year = {2025}
}