English
Related papers

Related papers: An item response theory evaluation of the Light an…

200 papers

Large Language Models (LLMs) are increasingly used as proxy students in the development of Intelligent Tutoring Systems (ITSs) and in piloting test questions. However, to what extent these proxy students accurately emulate the behavior and…

Computation and Language · Computer Science 2025-07-14 KV Aditya Srivatsa , Kaushal Kumar Maurya , Ekaterina Kochmar

Item Response Theory (IRT) has been proposed within the field of Educational Psychometrics to assess student ability as well as test question difficulty and discrimination power. More recently, IRT has been applied to evaluate machine…

Machine Learning · Statistics 2023-08-01 Sevvandi Kandanaarachchi , Kate Smith-Miles

Evaluating the abilities of learners is a fundamental objective in the field of education. In particular, there is an increasing need to assess higher-order abilities such as expressive skills and logical thinking. Constructed-response…

Computation and Language · Computer Science 2025-06-26 Masaki Uto , Yuma Ito

This review presents an overview and analysis of the body of research on special relativity theory (SRT) education at the secondary and lower undergraduate level. There is currently a growing international interest in implementing SRT in…

Physics Education · Physics 2021-07-21 Paul Alstein , Kim Krijtenburg-Lewerissa , Wouter R. van Joolingen

Humans can progressively learn visual concepts from easy to hard questions. To mimic this efficient learning ability, we propose a competence-aware curriculum for visual concept learning in a question-answering manner. Specifically, we…

Computer Vision and Pattern Recognition · Computer Science 2020-07-29 Qing Li , Siyuan Huang , Yining Hong , Song-Chun Zhu

Deep learning based knowledge tracing model has been shown to outperform traditional knowledge tracing model without the need for human-engineered features, yet its parameters and representations have long been criticized for not being…

Machine Learning · Computer Science 2019-04-29 Chun-Kit Yeung

The proliferation of Large Language Models (LLMs) necessitates valid evaluation methods to guide downstream applications and actionable future improvements. The Item Response Theory (IRT) has recently emerged as a promising framework for…

Methodology · Statistics 2025-12-12 Zhiyu Xu , Jia Liu , Yixin Wang , Yuqi Gu

Item Response Theory (IRT) aims to assess latent abilities of respondents based on the correctness of their answers in aptitude test items with different difficulty levels. In this paper, we propose the $\beta^3$-IRT model, which models…

Machine Learning · Statistics 2019-06-04 Yu Chen , Telmo Silva Filho , Ricardo B. C. Prudêncio , Tom Diethe , Peter Flach

The increasing use of interactive learning strategies in Astro 101 classrooms has led some instructors to consider the usefulness of a textbook in such classes. These strategies provide students a learning modality very different from the…

Physics Education · Physics 2015-06-17 Alexander L. Rudolph

We have found that non-STEM majors taking either a conceptual physics or astronomy course at two regional comprehensive institutions score significantly lower pre-instruction on the Lawson's Classroom Test of Scientific Reasoning (LCTSR) in…

Physics Education · Physics 2015-05-30 J. Christopher Moore , Louis J. Rubbo

Evaluating models and datasets in computer vision remains a challenging task, with most leaderboards relying solely on accuracy. While accuracy is a popular metric for model evaluation, it provides only a coarse assessment by considering a…

Computer Vision and Pattern Recognition · Computer Science 2024-09-09 Rahul Ramachandran , Tejal Kulkarni , Charchit Sharma , Deepak Vijaykeerthy , Vineeth N Balasubramanian

The Knowledge Tracing (KT) task focuses on predicting a learner's future performance based on the historical interactions. The knowledge state plays a key role in learning process. However, considering that the knowledge state is influenced…

Artificial Intelligence · Computer Science 2024-12-30 Shanshan Wang , Xueying Zhang , Keyang Wang , Xun Yang , Xingyi Zhang

Estimating student proficiency is an important task for computer based learning systems. We compare a family of IRT-based proficiency estimation methods to Deep Knowledge Tracing (DKT), a recently proposed recurrent neural network model…

Artificial Intelligence · Computer Science 2016-05-24 Kevin H. Wilson , Yan Karklin , Bojian Han , Chaitanya Ekanadham

Analyses of heterogeneous treatment effects (HTE) are common in applied causal inference research. However, when outcomes are latent variables assessed via psychometric instruments such as educational tests, standard methods ignore the…

We propose a class of Item Response Theory models for items with ordinal polytomous responses, which extends an existing class of multidimensional models for dichotomously-scored items measuring more than one latent trait. In the proposed…

Methodology · Statistics 2012-01-24 Silvia Bacci , Francesco Bartolucci , Michela Gnaldi

Ishimoto, Davenport, and Wittmann have previously reported analyses of data from student responses to the Force and Motion Conceptual Evaluation (FMCE), in which they used item response curves (IRCs) to make claims about American and…

Physics Education · Physics 2021-10-04 Connor J. Richardson , Trevor I. Smith , Paul J. Walter

Evaluating large language models (LLMs) typically requires thousands of benchmark items, making the process expensive, slow, and increasingly impractical at scale. Existing evaluation protocols rely on average accuracy over fixed item sets,…

Computation and Language · Computer Science 2026-02-03 Peiyu Li , Xiuxiu Tang , Si Chen , Ying Cheng , Ronald Metoyer , Ting Hua , Nitesh V. Chawla

South Korea's educational system has faced criticism for its lack of focus on critical thinking and creativity, resulting in high levels of stress and anxiety among students. As part of the government's effort to improve the educational…

Applications · Statistics 2024-05-28 Seorim Yi , Minkyu Kim , Jaewoo Park , Minjeong Jeon , Ick Hoon Jin

Automated short answer grading (ASAG) with large language models (LLMs) is commonly evaluated with aggregate metrics such as macro-F1 and Cohen's kappa. However, these metrics provide limited insight into how grading performance varies…

Computation and Language · Computer Science 2026-05-14 Longwei Cong , Sonja Hahn , Sebastian Gombert , Leon Camus , Hendrik Drachsler , Ulf Kroehne

Most datasets suffer from partial or complete missing values, which has downstream limitations on the available models on which to test the data and on any statistical inferences that can be made from the data. Several imputation techniques…

Machine Learning · Statistics 2023-02-09 Adrienne Kline , Yuan Luo