English

Simulated Self-Assessment in Large Language Models: A Psychometric Approach to AI Self-Efficacy

Artificial Intelligence 2025-11-27 v2

Abstract

Self-assessment is a key aspect of reliable intelligence, yet evaluations of large language models (LLMs) focus mainly on task accuracy. We adapted the 10-item General Self-Efficacy Scale (GSES) to elicit simulated self-assessments from ten LLMs across four conditions: no task, computational reasoning, social reasoning, and summarization. GSES responses were highly stable across repeated administrations and randomized item orders. However, models showed significantly different self-efficacy levels across conditions, with aggregate scores lower than human norms. All models achieved perfect accuracy on computational and social questions, whereas summarization performance varied widely. Self-assessment did not reliably reflect ability: several low-scoring models performed accurately, while some high-scoring models produced weaker summaries. Follow-up confidence prompts yielded modest, mostly downward revisions, suggesting mild overestimation in first-pass assessments. Qualitative analysis showed that higher self-efficacy corresponded to more assertive, anthropomorphic reasoning styles, whereas lower scores reflected cautious, de-anthropomorphized explanations. Psychometric prompting provides structured insight into LLM communication behavior but not calibrated performance estimates.

Keywords

Cite

@article{arxiv.2511.19872,
  title  = {Simulated Self-Assessment in Large Language Models: A Psychometric Approach to AI Self-Efficacy},
  author = {Daniel I Jackson and Emma L Jensen and Syed-Amad Hussain and Emre Sezgin},
  journal= {arXiv preprint arXiv:2511.19872},
  year   = {2025}
}

Comments

25 pages,5 tables, 3 figures