English
Related papers

Related papers: Item Response Theory -- A Statistical Framework fo…

200 papers

Interpretable Learning to Rank (LtR) is an emerging field within the research area of explainable AI, aiming at developing intelligible and accurate predictive models. While most of the previous research efforts focus on creating post-hoc…

Information Retrieval · Computer Science 2022-06-02 Claudio Lucchese , Franco Maria Nardini , Salvatore Orlando , Raffaele Perego , Alberto Veneri

Standard LLM evaluation practices compress diverse abilities into single scores, obscuring their inherently multidimensional nature. We present JE-IRT, a geometric item-response framework that embeds both LLMs and questions in a shared…

Artificial Intelligence · Computer Science 2025-09-30 Louie Hong Yao , Nicholas Jarvis , Tiffany Zhan , Saptarshi Ghosh , Linfeng Liu , Tianyu Jiang

In this paper, we demonstrate how Large Language Models (LLMs) can effectively learn to use an off-the-shelf information retrieval (IR) system specifically when additional context is required to answer a given question. Given the…

Computation and Language · Computer Science 2024-05-08 Tiziano Labruna , Jon Ander Campos , Gorka Azkune

Light phenomena conceptual assessment (LPCA) is a conceptual survey of light phenomena that has been recently established by the physics education research (PER) scholars. Studying the LPCA psychometric properties is always imperative to…

Physics Education · Physics 2023-12-27 Purwoko Haryadi Santoso , Edi Istiyono , Haryanto , Heri Retnawati

Relationships among teachers are known to influence their teaching-related perceptions. We study whether and how teachers' advising relationships (networks) are related to their perceptions of satisfaction, students, and influence over…

Methodology · Statistics 2026-02-18 Selena Wang , Plamena Powla , Tracy Sweet , Subhadeep Paul

Estimating student proficiency is an important task for computer based learning systems. We compare a family of IRT-based proficiency estimation methods to Deep Knowledge Tracing (DKT), a recently proposed recurrent neural network model…

Artificial Intelligence · Computer Science 2016-05-24 Kevin H. Wilson , Yan Karklin , Bojian Han , Chaitanya Ekanadham

Computerized Adaptive Testing (CAT) has proven effective for efficient LLM evaluation on multiple-choice benchmarks, but modern LLM evaluation increasingly relies on generation tasks where outputs are scored continuously rather than marked…

Computation and Language · Computer Science 2026-01-21 Esma Balkır , Alice Pernthaller , Marco Basaldella , José Hernández-Orallo , Nigel Collier

Linear Response theory aims to predict how added forcing alters the statistical properties of an unforced system. These kinds of questions have been studied predominantly for autonomous dynamical systems, yet many systems in the physical,…

Dynamical Systems · Mathematics 2026-04-07 Stefano Galatolo , Valerio Lucarini

Reward models are widely used as proxies for human preferences when aligning or evaluating LLMs. However, reward models are black boxes, and it is often unclear what, exactly, they are actually rewarding. In this paper we develop…

Computation and Language · Computer Science 2025-05-21 David Reber , Sean Richardson , Todd Nief , Cristina Garbacea , Victor Veitch

Latent variable models are popularly used to measure latent factors (e.g., abilities and personalities) from large-scale assessment data. Beyond understanding these latent factors, the covariate effect on responses controlling for latent…

Methodology · Statistics 2026-01-12 Jing Ouyang , Chengyu Cui , Kean Ming Tan , Gongjun Xu

Response times collected in computerised assessments provide information about the underlying response process and may exhibit within-person variation over the course of a test. We propose a latent variable model for log response times that…

Methodology · Statistics 2026-05-29 Gabriel Wallin , Nivedita Bhaktha

In standardized educational testing, test items are reused in multiple test administrations. To ensure the validity of test scores, the psychometric properties of items should remain unchanged over time. In this paper, we consider the…

Applications · Statistics 2021-10-26 Yunxiao Chen , Yi-Hsuan Lee , Xiaoou Li

Recent advances in large language models (LLMs) offer new opportunities for scalable, interactive mental health assessment, but excessive querying by LLMs burdens users and is inefficient for real-world screening across transdiagnostic…

Computation and Language · Computer Science 2025-11-21 Vasudha Varadarajan , Hui Xu , Rebecca Astrid Boehme , Mariam Marlan Mirstrom , Sverker Sikstrom , H. Andrew Schwartz

Large language models (LLMs) hold the potential to absorb and reflect personality traits and attitudes specified by users. In our study, we investigated this potential using robust psychometric measures. We adapted the most studied test in…

Computation and Language · Computer Science 2025-12-19 Anton Vasiliuk , Irina Abdullaeva , Polina Druzhinina , Anton Razzhigaev , Andrey Kuznetsov

The growing dependence on eTextbooks and Massive Open Online Courses (MOOCs) has led to an increase in the amount of students' learning data. By carefully analyzing this data, educators can identify difficult exercises, and evaluate the…

Data Structures and Algorithms · Computer Science 2022-11-28 Ahmed Abd Elrahman , Ahmed I. Taloba , Mohammed F. Farghally , Taysir Hassan A Soliman

Integrated Information Theory (IIT) is a prominent theory of consciousness that has at its centre measures that quantify the extent to which a system generates more information than the sum of its parts. While several candidate measures of…

Neurons and Cognition · Quantitative Biology 2019-01-30 Pedro A. M. Mediano , Anil K. Seth , Adam B. Barrett

Ordinal user-provided ratings across multiple items are frequently encountered in both scientific and commercial applications. Whilst recommender systems are known to do well on these type of data from a predictive point of view, their…

Methodology · Statistics 2025-03-05 Sjoerd Hermes

Effective educational measurement relies heavily on the curation of well-designed item pools (i.e., possessing the right psychometric properties). However, item calibration is time-consuming and costly, requiring a sufficient number of…

Computers and Society · Computer Science 2024-07-16 Yunting Liu , Shreya Bhandari , Zachary A. Pardos

This paper studies the item-to-item recommendation problem in recommender systems from a new perspective of metric learning via implicit feedback. We develop and investigate a personalizable deep metric model that captures both the internal…

Information Retrieval · Computer Science 2022-03-24 Trong Nghia Hoang , Anoop Deoras , Tong Zhao , Jin Li , George Karypis

Univariate isotonic regression (IR) has been used for nonparametric estimation in dose-response and dose-finding studies. One undesirable property of IR is the prevalence of piecewise-constant stretches in its estimates, whereas the…

Methodology · Statistics 2017-01-24 Assaf P. Oron , Nancy Flournoy