中文

PersonaTeaming: 揭示在 AI 红队测试中引入 persona 的作用

人工智能 2025-10-28 v3 人机交互

摘要

近期的 AI 治理和 safety research 的发展呼吁采用 red-teaming 方法以有效揭示 AI 模型潜在风险。许多呼声都强调 red-teamers 的 identities 和 backgrounds 会影响 their red-teaming strategies,从而决定其 likely 能揭示的风险类型。虽然 automated red-teaming 方法 promise to complement human red-teaming by enabling larger-scale exploration of model behavior,但 current approaches 不考虑 identity 的 role。在将 people's background 和 identities 纳入 automated red-teaming 的 initial step 方面,我们开发并评估了一种 novel method,PersonaTeaming,通过在 adversarial prompt generation process 中引入 personas 来 explore 更广 spectrum 的 adversarial strategies。具体而言,我们首先引入一种 methodology for mutating prompts based on either "red-teaming expert" personas 或 "regular AI user" personas。随后,我们开发了 dynamic persona-generating algorithm,以自动生成 various persona types adaptive to different seed prompts。此外,我们开发了一组 new metrics 来 explicitly measure "mutation distance" 以补充 existing diversity measurements of adversarial prompts。我们的实验显示,与 RainbowPlus(一种 state-of-the-art automated red-teaming method)相比,通过 persona mutation 在 attack success rates 上实现了 promising improvements (up to 144.1%),同时保持 prompt diversity。我们讨论了不同 persona types 和 mutation methods 的 strengths and limitations,为探索 automated 与 human red-teaming approaches 之间的 complementarities 提供了启发。

关键词

引用

@article{arxiv.2509.03728,
  title  = {PersonaTeaming: Exploring How Introducing Personas Can Improve Automated AI Red-Teaming},
  author = {Wesley Hanwen Deng and Sunnie S. Y. Kim and Akshita Jha and Ken Holstein and Motahhare Eslami and Lauren Wilcox and Leon A Gatys},
  journal= {arXiv preprint arXiv:2509.03728},
  year   = {2025}
}