English
Related papers

Related papers: SPRITE: A Response Model For Multiple Choice Testi…

200 papers

We present a new probing dataset named PROST: Physical Reasoning about Objects Through Space and Time. This dataset contains 18,736 multiple-choice questions made from 14 manually curated templates, covering 10 physical reasoning concepts.…

Computation and Language · Computer Science 2021-06-08 Stéphane Aroca-Ouellette , Cory Paik , Alessandro Roncone , Katharina Kann

Incorporating Item Response Theory (IRT) into NLP tasks can provide valuable information about model performance and behavior. Traditionally, IRT models are learned using human response pattern (RP) data, presenting a significant bottleneck…

Computation and Language · Computer Science 2019-09-02 John P. Lalor , Hao Wu , Hong Yu

Randomized experiments are the gold standard for estimating the average treatment effect (ATE). While covariate adjustment can reduce the asymptotic variances of the unbiased Horvitz-Thompson estimators for the ATE, it suffers from…

Methodology · Statistics 2025-08-22 Xin Lu , Lei Shi , Hanzhong Liu , Peng Ding

School bullying victimization is a variable that cannot be measured directly. Taking into account that this variable has a lower bound, given by the absence of bullying victimization, this paper proposes IRT logistic models, where the…

Applications · Statistics 2019-04-03 Edilberto Cepeda-Cuervo

Users increasingly expect modern search systems to offer a unified interface that seamlessly retrieves information from diverse data sources and formats. However, current information retrieval (IR) evaluation benchmarks have not kept pace…

Information Retrieval · Computer Science 2026-05-13 Mehmet Deniz Türkmen , Suchana Datta , Dwaipayan Roy , Daniel Hienert , Philipp Mayr , Derek Greene

Recent deep learning methods for fMRI-based diagnosis have achieved promising accuracy by modeling functional connectivity networks. However, standard approaches often struggle with noisy interactions, and conventional post-hoc attribution…

Machine Learning · Computer Science 2026-02-25 Kunyu Zhang , Yanwu Yang , Jing Zhang , Xiangjie Shi , Shujian Yu

Accuracy-based evaluation of Large Language Models (LLMs) measures benchmark-specific performance rather than underlying medical competency: it treats all questions as equally informative, conflates model ability with item characteristics,…

Computation and Language · Computer Science 2026-04-07 Zhimeng Luo , Lixin Wu , Adam Frisch , Daqing He

In many categorical response regression applications, the response categories admit a multiresolution structure. That is, subsets of the response categories may naturally be combined into coarser response categories. In such applications,…

Methodology · Statistics 2023-08-25 Aaron J. Molstad , Keshav Motwani

In responding to rating questions, an individual may give answers either according to his/her knowledge/awareness or to his/her level of indecision/uncertainty, typically driven by a response style. As ignoring this dual behaviour may lead…

Methodology · Statistics 2021-04-08 Roberto Colombi , Sabrina Giordano , Anna Gottard , Maria Iannario

A common approach when studying the quality of representation involves comparing the latent preferences of voters and legislators, commonly obtained by fitting an item-response theory (IRT) model to a common set of stimuli. Despite being…

Applications · Statistics 2023-08-08 Yuki Shiraito , James Lo , Santiago Olivella

In this paper, we apply Vuong's (1989) general approach of model selection to the comparison of nested and non-nested unidimensional and multidimensional item response theory (IRT) models. Vuong's approach of model selection is useful…

Applications · Statistics 2019-09-20 Lennart Schneider , R. Philip Chalmers , Rudolf Debelak , Edgar C. Merkle

User preferences for items can be inferred from either explicit feedback, such as item ratings, or implicit feedback, such as rental histories. Research in collaborative filtering has concentrated on explicit feedback, resulting in the…

Machine Learning · Computer Science 2015-03-19 Andriy Mnih , Yee Whye Teh

A new random forest based model for solving the Multiple Instance Learning (MIL) problem under small tabular data, called Soft Tree Ensemble MIL (STE-MIL), is proposed. A new type of soft decision trees is considered, which is similar to…

Machine Learning · Computer Science 2023-02-14 Andrei V. Konstantinov , Lev V. Utkin

We consider the setting of iterative learning control, or model-based policy learning in the presence of uncertain, time-varying dynamics. In this setting, we propose a new performance metric, planning regret, which replaces the standard…

Machine Learning · Computer Science 2021-03-01 Naman Agarwal , Elad Hazan , Anirudha Majumdar , Karan Singh

Classic item response models assume that all items with the same difficulty have the same response probability among all respondents with the same ability. These assumptions, however, may very well be violated in practice, and it is not…

Methodology · Statistics 2021-08-23 Minjeong Jeon , Ick Hoon Jin , Michael Schweinberger , Samuel Baugh

Computerized adaptive testing is becoming increasingly popular due to advancement of modern computer technology. It differs from the conventional standardized testing in that the selection of test items is tailored to individual examinee's…

Statistics Theory · Mathematics 2009-06-11 Hua-Hua Chang , Zhiliang Ying

Identifying patient subgroups with different treatment responses is an important task to inform medical recommendations, guidelines, and the design of future clinical trials. Existing approaches for treatment effect estimation primarily…

Methodology · Statistics 2025-12-10 Vincent Jeanselme , Chang Ho Yoon , Fabian Falck , Brian Tom , Jessica Barrett

Model-based deep learning solutions to inverse problems have attracted increasing attention in recent years as they bridge state-of-the-art numerical performance with interpretability. In addition, the incorporated prior domain knowledge…

Machine Learning · Statistics 2023-09-18 Frederik Hoppe , Claudio Mayrink Verdun , Felix Krahmer , Hannah Laus , Holger Rauhut

Visual object tracking performance has been dramatically improved in recent years, but some severe challenges remain open, like distractors and occlusions. We suspect the reason is that the feature representations of the tracking targets…

Computer Vision and Pattern Recognition · Computer Science 2021-10-29 Mengmeng Wang , Xiaoqian Yang , Yong Liu

The learning of predictive models for data-driven decision support has been a prevalent topic in many fields. However, construction of models that would capture interactions among input variables is a challenging task. In this paper, we…

Machine Learning · Computer Science 2019-05-22 Jiapeng Liu , Milosz Kadzinski , Xiuwu Liao , Xiaoxin Mao