面向软件工程社区中基于LLM的心理安全定性编码的提示工程策略:受控实证研究
摘要
定性分析在理解软件工程的人类和社会方面发挥着关键作用,但仍是一个受主观解释影响且对方法论选择(如提示设计)敏感的艰巨过程。近期大型语言模型(LLM)的发展为支持这类分析提供了有前景的机会,尽管其在不同提示条件下能否可靠地再现人类定性推理尚未广泛测试。本研究对三种 LLM——Claude Haiku、DeepSeek-Chat 和 Gemini 2.5 Flash——在两种提示工程策略(零shot 和多shot 闭合编码)下的表现进行受控实证评估。以 Cohen's kappa 作为主要一致性指标,对每种配置进行十次独立运行。结果表明,多shot 提示显著提高了 Claude Haiku 的一致性(Delta kappa = +0.034, Wilcoxon p = 0.004),但对 DeepSeek-Chat 或 Gemini 2.5 Flash 并无显著效果。不同模型内部稳定性差异显著——DeepSeek-Chat 和 Claude Haiku 的方差最小(SD 约 0.017),而 Gemini 2.5 Flash 最不稳定(SD = 0.038)。我们识别出所有模型在预测“分享负面反馈”时存在系统性过度预测(偏差比最高可达 5.25 倍),同时“表达关切” consistently 被低估。总体而言,这些发现为软件工程研究中 LLM 辅助定性编码的提示工程指南提供了实证依据。
引用
@article{arxiv.2605.07422,
title = {Prompt Engineering Strategies for LLM-based Qualitative Coding of Psychological Safety in Software Engineering Communities: A Controlled Empirical Study},
author = {Moaath Alshaikh and Tasneem Alshaher and Ricardo Vieira and Beatriz Santana and Clelio Xavier and Jose Amancio and Glauco Carneiro and Julio Leite and Savio Freire and Manoel Mendonca},
journal= {arXiv preprint arXiv:2605.07422},
year = {2026}
}
备注
9 pages, 5 figures. Accepted at the 1st International Workshop on Prompt Engineering for Software Engineering (PROMPT-SE 2026), co-located with the 30th International Conference on Evaluation and Assessment in Software Engineering (EASE 2026), Glasgow, Scotland, United Kingdom, June 9--12, 2026