English

Clustering Students and Inferring Skill Set Profiles with Skill Hierarchies

Applications 2021-04-07 v1

Abstract

Cognitive diagnosis models (CDMs) are a popular tool for assessing students' mastery of sets of skills. Given a set of KK skills tested on an assessment, students are classified into one of 2K2^K latent skill set profiles that represent whether they have mastered each skill or not. Traditional approaches to estimating these profiles are computationally intensive and become infeasible on large datasets. Instead, proxy skill estimates can be generated from the observed responses and then clustered, and these clusters can be assigned to different profiles. Building on previous work, we consider how to optimally perform this clustering when not all 2K2^K profiles are possible, e.g. because of hierarchical relationships among the skills, and when not all possible profiles are present in the population. We compare hierarchical clustering and several k-means variants, including semisupervised clustering using simulated student responses. The empty k-means algorithm paired with a novel method for generating starting centers yields the best overall performance.

Keywords

Cite

@article{arxiv.2104.02237,
  title  = {Clustering Students and Inferring Skill Set Profiles with Skill Hierarchies},
  author = {Alan Mishler and Rebecca Nugent},
  journal= {arXiv preprint arXiv:2104.02237},
  year   = {2021}
}

Comments

4 pages, 3 figures. Originally presented at the Doctoral Consortium of the 11th International Conference on Educational Data Mining, July, 2018, Buffalo, NY