English
Related papers

Related papers: Item Response Theory -- A Statistical Framework fo…

200 papers

Forensic science often involves the comparison of crime-scene evidence to a known-source sample to determine if the evidence and the reference sample came from the same source. Even as forensic analysis tools become increasingly objective…

Applications · Statistics 2019-10-17 Amanda Luby , Anjali Mazumder , Brian Junker

Measurement validity in Item Response Theory depends on appropriately modeling dependencies between items when these reflect meaningful theoretical structures rather than random measurement error. In ecological assessment, citizen…

Applications · Statistics 2025-07-22 Mingya Huang , Soham Ghosh

Large language models (LLMs) achieve high performance on mathematical reasoning, but these results can be inflated by training data leakage or superficial pattern matching rather than genuine reasoning. To this end, an adversarial…

Computation and Language · Computer Science 2026-02-03 Xinyuan Li , Murong Xu , Wenbiao Tao , Hanlun Zhu , Yike Zhao , Jipeng Zhang , Yunshi Lan

Generative AI is transforming the educational landscape, raising significant concerns about cheating. Despite the widespread use of multiple-choice questions in assessments, the detection of AI cheating in MCQ-based tests has been almost…

Artificial Intelligence · Computer Science 2024-12-13 Alona Strugatski , Giora Alexandron

Modeling item parameters as a function of item characteristics has a long history but has generally focused on models for item location. Explanatory item response models for item discrimination are available but rarely used. In this study,…

Methodology · Statistics 2025-06-24 Joshua B. Gilbert , Lijin Zhang , Esther Ulitzsch , Benjamin W. Domingue

Evaluation of large language models (LLMs) is increasingly critical, yet standard benchmarking methods rely on average accuracy, overlooking both the inherent stochasticity of LLM outputs and the heterogeneity of benchmark items. Item…

Machine Learning · Statistics 2026-05-11 Xinhao Qu , Qiang Heng , Hao Zeng , Xiaoqian Liu

Human evaluations play a central role in training and assessing AI models, yet these data are rarely treated as measurements subject to systematic error. This paper integrates psychometric rater models into the AI pipeline to improve the…

Artificial Intelligence · Computer Science 2026-02-27 Jodi M. Casabianca , Maggie Beiting-Parrish

Persona conditioning is widely used to steer large language model (LLM) behavior, but it is unclear whether it induces stable behavioral structure or superficial variation. We propose a framework to measure consistent behavioral tendencies…

Artificial Intelligence · Computer Science 2026-05-12 Alexandra Yost , Shreyans Jain , Shivam Raval , Grant Corser , Allen Roush , Nina Xu , Jacqueline Hammack , Ravid Shwartz-Ziv , Amirali Abdullah

One of the largest drivers of social inequality is unequal access to personal tutoring, with wealthier individuals able to afford it, while the majority cannot. Affordable, effective AI tutors offer a scalable solution. We focus on adaptive…

Artificial Intelligence · Computer Science 2025-10-01 Tom Quilter , Anastasia Ilick , Karen Poon , Richard Turner

Classic item response models assume that all items with the same difficulty have the same response probability among all respondents with the same ability. These assumptions, however, may very well be violated in practice, and it is not…

Methodology · Statistics 2021-08-23 Minjeong Jeon , Ick Hoon Jin , Michael Schweinberger , Samuel Baugh

A common approach when studying the quality of representation involves comparing the latent preferences of voters and legislators, commonly obtained by fitting an item-response theory (IRT) model to a common set of stimuli. Despite being…

Applications · Statistics 2023-08-08 Yuki Shiraito , James Lo , Santiago Olivella

Evaluating the abilities of learners is a fundamental objective in the field of education. In particular, there is an increasing need to assess higher-order abilities such as expressive skills and logical thinking. Constructed-response…

Computation and Language · Computer Science 2025-06-26 Masaki Uto , Yuma Ito

The rapid release of both language models and benchmarks makes it increasingly costly to evaluate every model on every dataset. In practice, models are often evaluated on different samples, making scores difficult to compare across studies.…

Computation and Language · Computer Science 2026-04-16 Eliya Habba , Itay Itzhak , Asaf Yehudai , Yotam Perlitz , Elron Bandel , Michal Shmueli-Scheuer , Leshem Choshen , Gabriel Stanovsky

Machine-learned models for author profiling in social media often rely on data acquired via self-reporting-based psychometric tests (questionnaires) filled out by social media users. This is an expensive but accurate data collection…

Computation and Language · Computer Science 2022-05-17 Anne Kreuter , Kai Sassenberg , Roman Klinger

Within the educational context, students' assessment tests are routinely validated through Item Response Theory (IRT) models which assume unidimensionality and absence of Differential Item Functioning (DIF). In this paper, we investigate if…

Applications · Statistics 2012-12-04 Michela Gnaldi , Francesco Bartolucci , Silvia Bacci

Since the introduction of network psychometrics, several connections to statistical models in "classical" psychometrics (i.e., IRT, SEM, GLM) as well as to approaches from other research fields have been established. In this paper, these…

Methodology · Statistics 2026-05-18 Kevin Kistermann , Vivato V. Andriamiarana , Augustin Kelava

The Rasch model is one of the most fundamental models in \emph{item response theory} and has wide-ranging applications from education testing to recommendation systems. In a universe with $n$ users and $m$ items, the Rasch model assumes…

Machine Learning · Computer Science 2023-10-31 Duc Nguyen , Anderson Zhang

As psychometric surveys are increasingly used to assess the traits of large language models (LLMs), the need for scalable survey item generation suited for LLMs has also grown. A critical challenge here is ensuring the construct validity of…

Computation and Language · Computer Science 2026-05-26 Sungjib Lim , Woojung Song , Eun-Ju Lee , Yohan Jo

In a recent review, Liu, Pek, & Maydeu-Olivares (2025b) classified reliability coefficients into two types: classical test theory (CTT) reliability and proportional reduction in mean squared error (PRMSE). This article focuses on…

Methodology · Statistics 2026-04-14 Youjin Sung , Yang Liu

A comprehensive class of models is proposed that can be used for continuous, binary, ordered categorical and count type responses. The difficulty of items is described by difficulty functions, which replace the item difficulty parameters…

Methodology · Statistics 2021-06-25 Gerhard Tutz
‹ Prev 1 3 4 5 6 7 10 Next ›