English

The Use of AI Tools to Develop and Validate Q-Matrices

Artificial Intelligence 2026-02-10 v1 Computation and Language

Abstract

Constructing a Q-matrix is a critical but labor-intensive step in cognitive diagnostic modeling (CDM). This study investigates whether AI tools (i.e., general language models) can support Q-matrix development by comparing AI-generated Q-matrices with a validated Q-matrix from Li and Suen (2013) for a reading comprehension test. In May 2025, multiple AI models were provided with the same training materials as human experts. Agreement among AI-generated Q-matrices, the validated Q-matrix, and human raters' Q-matrices was assessed using Cohen's kappa. Results showed substantial variation across AI models, with Google Gemini 2.5 Pro achieving the highest agreement (Kappa = 0.63) with the validated Q-matrix, exceeding that of all human experts. A follow-up analysis in January 2026 using newer AI versions, however, revealed lower agreement with the validated Q-matrix. Implications and directions for future research are discussed.

Keywords

Cite

@article{arxiv.2602.08796,
  title  = {The Use of AI Tools to Develop and Validate Q-Matrices},
  author = {Kevin Fan and Jacquelyn A. Bialo and Hongli Li},
  journal= {arXiv preprint arXiv:2602.08796},
  year   = {2026}
}

Comments

An earlier version of this study was presented at the Psychometric Society Meeting held in July 2025 in Minneapolis, USA

R2 v1 2026-07-01T10:28:08.269Z