English
Related papers

Related papers: Item Response Theory -- A Statistical Framework fo…

200 papers

In this paper, we apply Vuong's (1989) general approach of model selection to the comparison of nested and non-nested unidimensional and multidimensional item response theory (IRT) models. Vuong's approach of model selection is useful…

Applications · Statistics 2019-09-20 Lennart Schneider , R. Philip Chalmers , Rudolf Debelak , Edgar C. Merkle

Measurement non-invariance arises when the psychometric properties of a scale differ across subgroups, undermining the validity of group comparisons. At the item level, such non-invariance manifests as differential item functioning (DIF),…

Methodology · Statistics 2026-01-27 Gabriel Wallin , Qi Huang

An adaptive design adjusts dynamically as information is accrued and a consequence of applying an adaptive design is the potential for inducing small-sample bias in estimates. In psychometrics and psychophysics, a common class of studies…

Methodology · Statistics 2025-02-17 Simon Bang Kristensen , Katrine Bødkergaard , Bo Martin Bibby

Since the public release of Chat Generative Pre-Trained Transformer (ChatGPT), extensive discourse has emerged concerning the potential advantages and challenges of integrating Generative Artificial Intelligence (GenAI) into education. In…

Computers and Society · Computer Science 2024-07-30 Jan-Erik Kalmus , Anastasija Nikiforova

An intelligent tutoring system (ITS) aims to provide instructions and exercises tailored to the ability of a student. To do this, the ITS needs to estimate the ability based on student input. Rather than including frequent full-scale tests…

Methodology · Statistics 2024-11-12 Karl Sigfrid , Ellinor Fackle-Fornius , Frank Miller

I present a new approach for the interpretation of reaction time (RT) data from behavioral experiments. From a physical perspective, the entropy of the RT distribution provides a model-free estimate of the amount of processing performed by…

Neurons and Cognition · Quantitative Biology 2009-08-24 Fermín Moscoso del Prado Martín

Human self-report questionnaires are increasingly used in NLP to benchmark and audit large language models (LLMs), from persona consistency to safety and bias assessments. Yet these instruments presume honest responding; in evaluative…

Computation and Language · Computer Science 2026-04-29 Kensuke Okada , Yui Furukawa , Kyosuke Bunji

Reinforcement learning (RL) is concerned with how intelligence agents take actions in a given environment to maximize the cumulative reward they receive. In healthcare, applying RL algorithms could assist patients in improving their health…

Machine Learning · Statistics 2025-04-21 Chengchun Shi

Recent advancements in Large Language Models (LLMs) have led to their adaptation in various domains as conversational agents. We wonder: can personality tests be applied to these agents to analyze their behavior, similar to humans? We…

Reliance on stereotypes is a persistent feature of human decision-making and has been extensively documented in educational settings, where it can shape students' confidence, performance, and long-term human capital accumulation. While…

General Economics · Economics 2025-03-05 Elisa Baldazzi , Pietro Biroli , Marina Della Giusta , Florent Dubois

Information-theoretic (IT) measures are ubiquitous in artificial intelligence: entropy drives decision-tree splits and uncertainty quantification, cross-entropy is the default classification loss, mutual information underpins representation…

Artificial Intelligence · Computer Science 2026-04-28 Nikolaos Al. Papadopoulos , Konstantinos E. Psannis

Attention can be used to inform choice selection in contextual bandit tasks even when context features have not been previously experienced. One example of this is in dimensional shifts, where additional feature values are introduced and…

Machine Learning · Computer Science 2025-05-16 Tailia Malloy , Roderick Seow , Cleotilde Gonzalez

This paper introduces the generalized Hausman test as a novel method for detecting non-normality of the latent variable distribution of unidimensional Item Response Theory (IRT) models for binary data. The test utilizes the pairwise maximum…

Methodology · Statistics 2024-02-14 Lucia Guastadisegni , Silvia Cagnone , Irini Moustaki , Vassilis Vasdekis

Typical IRT rating-scale models assume that the rating category threshold parameters are the same over examinees. However, it can be argued that many rating data sets violate this assumption. To address this practical psychometric problem,…

Methodology · Statistics 2013-03-22 Ken Akira Fujimoto , George Karabatsos

This paper proposes a method for assessing differential item functioning (DIF) in item response theory (IRT) models. The method does not require pre-specification of anchor items, which is its main virtue. It is developed in two main steps,…

Methodology · Statistics 2025-01-08 Peter F. Halpin

While regression models capture the relationship between predictors and the response variable, they often lack intuitive accompanying methods to understand the influence of predictors on the outcome. To address this, we introduce an…

Methodology · Statistics 2026-02-06 Jihao You , Dan Tulpan , Jiaojiao Diao , Jennifer L. Ellis

Statistical evaluation aims to estimate the generalization performance of a model using held-out i.i.d.\ test data sampled from the ground-truth distribution. In supervised learning settings such as classification, performance metrics such…

Machine Learning · Computer Science 2026-04-08 Shashaank Aiyer , Yishay Mansour , Shay Moran , Han Shao

The rapid adoption of large language models (LLMs) in education raises profound challenges for assessment design. To adapt assessments to the presence of LLM-based tools, it is crucial to characterize the strengths and weaknesses of LLMs in…

Human-Computer Interaction · Computer Science 2026-04-16 Licol Zeinfeld , Alona Strugatski , Ziva Bar-Dov , Ron Blonder , Shelley Rap , Giora Alexandron

Ishimoto, Davenport, and Wittmann have previously reported analyses of data from student responses to the Force and Motion Conceptual Evaluation (FMCE), in which they used item response curves (IRCs) to make claims about American and…

Physics Education · Physics 2021-10-04 Connor J. Richardson , Trevor I. Smith , Paul J. Walter

The Implicit Association Test, IAT, is widely used to measure hidden (subconscious) human biases, implicit bias, of many topics: race, gender, age, ethnicity, religion stereotypes. There is a need to understand the reliability of these…

Applications · Statistics 2023-12-27 S. Stanley Young , Warren B. Kindzierski