POLAR:一种嵌入空间中的 per-user 关联性测试
计算与语言
2026-03-18 v1 计算机与社会
社会与信息网络
摘要
大多数内在关联探针 operate at 词、句或语料库水平,掩盖了作者层面的变异性。我们提出 POLAR(Per-user On-axis Lexical Association Re-port),即一种在嵌入空间中运行的 per-user 词汇关联测试。作者由私有确定性 to-kens 表示;POLAR 将这些向量投影到精选的词汇轴上,并报告标准化效应及置换 p 值和 Benjamini--Hochberg 控制。在一个平衡的 bot--human Twitter 基准测试中,POLAR 能清晰区分 LLM 生成的机器人与自然账户;在极端主义论坛上,它量化了与侮辱性词汇词典的强烈关联,并揭示了随时间右移趋势。该方法可模块化到新的属性集,为计算社会科学提供简洁的 per-author 诊断。所有代码均在 https://github.com/pedroaugtb/POLAR-A-Per-User-Association-Test-in-Embedding-Space 上公开。
引用
@article{arxiv.2603.15950,
title = {POLAR:A Per-User Association Test in Embedding Space},
author = {Pedro Bento and Arthur Buzelin and Arthur Chagas and Yan Aquino and Victoria Estanislau and Samira Malaquias and Pedro Robles Dutenhefner and Gisele L. Pappa and Virgilio Almeida and Wagner MeiraJr},
journal= {arXiv preprint arXiv:2603.15950},
year = {2026}
}
备注
Accepted paper at ICWSM 2026