中文

A Systematic Evaluation of Large Language Models for PTSD Severity Estimation: The Role of Contextual Knowledge and Modeling Strategies

计算与语言 2026-02-06 v1

摘要

大型语言模型 (LLM) 越来越多地以零-shot 方法用于评估精神健康状况,但我们对其准确性影响因素知之甚少。在本研究中,我们利用来自 1,437 名个体的临床数据集,包含自然语言叙述和自报告的 PTSD 严重程度评分,对 11 个最新技术 LLM 进行全面评估。为了解影响准确性的因素,我们系统性地变异(i)如子尺度定义、分布摘要和访谈问题等的上下文知识,以及(ii)包括零-shot 与 few shot、推理 effort 量、模型规模、结构子尺度 vs 直接标量预测、输出重新缩放和九种集成方法的建模策略.我们的发现表明:(a)在提供详细的 construct 定义和叙述上下文时,LLM 最精确;(b)增加推理 effort 导致更好的估计准确性;(c)开源权重模型 (Llama、Deepseek) 在超过 70B 参数后性能平台化,而封闭权重 (o3-mini、gpt-5) 模型随新一代而提升;(d)通过将受监督模型与零-shot LLM 集成可实现最佳性能。总体而言,结果表明上下文知识和建模策略的选择对于将 LLM 部署到准确评估精神健康方面至关重要。

关键词

引用

@article{arxiv.2602.06015,
  title  = {A Systematic Evaluation of Large Language Models for PTSD Severity Estimation: The Role of Contextual Knowledge and Modeling Strategies},
  author = {Panagiotis Kaliosis and Adithya V Ganesan and Oscar N. E. Kjell and Whitney Ringwald and Scott Feltman and Melissa A. Carr and Dimitris Samaras and Camilo Ruggero and Benjamin J. Luft and Roman Kotov and Andrew H. Schwartz},
  journal= {arXiv preprint arXiv:2602.06015},
  year   = {2026}
}

备注

18 pages, 3 figures, 5 tables