English
Related papers

Related papers: Item Response Theory -- A Statistical Framework fo…

200 papers

Reaction time (RT) is a fundamental measure in cognitive and neurophysiological assessment, yet most existing RT systems require active user engagement and controlled environments, limiting their use in real-world settings. This paper…

Quantitative Methods · Quantitative Biology 2026-03-13 Abhigyan Sarkar , Boris Rubinsky

Despite the availability of benchmark machine learning (ML) repositories (e.g., UCI, OpenML), there is no standard evaluation strategy yet capable of pointing out which is the best set of datasets to serve as gold standard to test different…

Language models (LMs) are increasingly used to simulate human-like responses in scenarios where accurately mimicking a population's behavior can guide decision-making, such as in developing educational materials and designing public…

Computation and Language · Computer Science 2024-07-23 Joy He-Yueya , Wanjing Anya Ma , Kanishk Gandhi , Benjamin W. Domingue , Emma Brunskill , Noah D. Goodman

Research on the test structure of the Force Concept Inventory (FCI) has largely been performed with exploratory methods such as factor analysis and cluster analysis. Multi-Dimensional Item Response Theory (MIRT) provides an alternative to…

Physics Education · Physics 2018-06-20 John Stewart , Cabot Zabriskie , Seth DeVore , Gay Stewart

Healthcare quality metrics refer to a variety of measures used to characterize what should have been done or not done for a patient or the health consequences of what was or was not done. When estimating healthcare quality, many metrics are…

Applications · Statistics 2025-02-13 Sharon-Lise Normand , Katya Zelevinsky , Marcela Horvitz-Lennon

Item factor analysis (IFA) refers to the factor models and statistical inference procedures for analyzing multivariate categorical data. IFA techniques are commonly used in social and behavioral sciences for analyzing item-level response…

Methodology · Statistics 2020-04-17 Yunxiao Chen , Siliang Zhang

This work presents a systematic study of objective evaluations of abstaining classifications using Information-Theoretic Measures (ITMs). First, we define objective measures for which they do not depend on any free parameter. This…

Computer Vision and Pattern Recognition · Computer Science 2012-08-16 Bao-Gang Hu , Ran He , XiaoTong Yuan

Statistical models such as those derived from Item Response Theory (IRT) enable the assessment of students on a specific subject, which can be useful for several purposes (e.g., learning path customization, drop-out prediction). However,…

Computation and Language · Computer Science 2020-05-07 Luca Benedetto , Andrea Cappelli , Roberto Turrin , Paolo Cremonesi

Opinion polls remain among the most efficient and widespread methods to capture psycho-social data at large scales. However, there are limitations on the logistics and structure of opinion polls that restrict the amount and type of…

Applications · Statistics 2018-10-16 Alvin Vista

Nonparametric item response models provide a flexible framework in psychological and educational measurements. Douglas (2001) established asymptotic identifiability for a class of models with nonparametric response functions for long…

Statistics Theory · Mathematics 2025-01-08 Yinqiu He

Monte Carlo simulations are the primary methodology for evaluating Item Response Theory (IRT) methods, yet marginal reliability - the fundamental metric of data informativeness - is rarely treated as an explicit design factor. Unlike in…

Methodology · Statistics 2026-01-14 JoonHo Lee

It is reasonable to consider, in many cases, that individuals' latent traits have a hierarchical structure such that more general traits are a suitable composition of more specific ones. Existing item response models that account for such…

Methodology · Statistics 2020-07-28 Juliane Venturelli S. L. , Flavio B. Gonçalves , Dalton F. Andrade

Item difficulty plays a crucial role in test performance, interpretability of scores, and equity for all test-takers, especially in large-scale assessments. Traditional approaches to item difficulty modeling rely on field testing and…

Computation and Language · Computer Science 2025-09-30 Sydney Peters , Nan Zhang , Hong Jiao , Ming Li , Tianyi Zhou , Robert Lissitz

High-quality test items are essential for educational assessments, particularly within Item Response Theory (IRT). Traditional validation methods rely on resource-intensive pilot testing to estimate item difficulty and discrimination. More…

Computation and Language · Computer Science 2025-08-08 Robin Schmucker , Steven Moore

Item response theory (IRT) models are widely used in psychometrics and educational measurement, being deployed in many high stakes tests such as the GRE aptitude test. IRT has largely focused on estimation of a single latent trait (e.g.…

Machine Learning · Statistics 2019-09-10 Ajay Shanker Tripathi , Benjamin W. Domingue

Thematic Apperception Test (TAT) is a psychometrically grounded, multidimensional assessment framework that systematically differentiates between cognitive-representational and affective-relational components of personality-like…

Computation and Language · Computer Science 2026-02-20 Anton Dzega , Aviad Elyashar , Ortal Slobodin , Odeya Cohen , Rami Puzis

The rapid proliferation of large language models (LLMs) in healthcare creates an urgent need for scalable and psychometrically sound evaluation methods. Conventional static benchmarks are costly to administer repeatedly, vulnerable to data…

Computation and Language · Computer Science 2026-03-26 Tianpeng Zheng , Zhehan Jiang , Jiayi Liu , Shicong Feng

Naming tests represent an essential tool in gauging the severity of aphasia and monitoring the trajectory of recovery for individuals afflicted with this debilitating condition. In these assessments, patients are presented with images…

Multimodal Large Language Models (MLLMs) have recently emerged as general architectures capable of reasoning over diverse modalities. Benchmarks for MLLMs should measure their ability for cross-modal integration. However, current benchmarks…

Computation and Language · Computer Science 2026-03-04 Shunki Uebayashi , Kento Masui , Kyohei Atarashi , Han Bao , Hisashi Kashima , Naoto Inoue , Mayu Otani , Koh Takeuchi

Reading comprehension is a key for individual success, yet the assessment of question difficulty remains challenging due to the extensive human annotation and large-scale testing required by traditional methods such as linguistic analysis…

Computation and Language · Computer Science 2025-02-26 Yoshee Jain , John Hollander , Amber He , Sunny Tang , Liang Zhang , John Sabatini
‹ Prev 1 4 5 6 7 8 10 Next ›