中文
相关论文

相关论文: Regularized Bayesian calibration and scoring of th…

200 篇论文

Robust validation of Machine Learning (ML) models is essential, but traditional data partitioning approaches often ignore the intrinsic quality of each instance. This study proposes the use of Item Response Theory (IRT) parameters to…

机器学习 · 计算机科学 2025-08-15 Lucas Cardoso , Vitor Santos , José Ribeiro Filho , Ricardo Prudêncio , Regiane Kawasaki , Ronnie Alves

Despite recent progress in Large Language Model (LLM) Agents for Software Engineering (SWE) tasks, end-to-end fine-tuning typically relies on verifiable terminal rewards such as whether all unit tests pass. While these binary signals…

机器学习 · 计算机科学 2026-04-21 Jiawei Huang , Qingping Yang , Renjie Zheng , Jiaze Chen

Evaluation of NLP methods requires testing against a previously vetted gold-standard test set and reporting standard metrics (accuracy/precision/recall/F1). The current assumption is that all items in a given test set are equal with regards…

计算与语言 · 计算机科学 2016-09-26 John P. Lalor , Hao Wu , Hong Yu

Multivariate Item Response Theory (MIRT) is sought-after widely by applied researchers looking for interpretable (sparse) explanations underlying response patterns in questionnaire data. There is, however, an unmet demand for such sparsity…

统计方法学 · 统计学 2025-03-10 Jiguang Li , Robert Gibbons , Veronika Rockova

Missingness is a common occurrence in educational assessment and psychological measurement. It could not be casually ignored as it may threaten the validity of the test if not handled properly. Considering the difference between omitted and…

统计方法学 · 统计学 2019-04-09 Jinxin Guo

Item (question) difficulties play a crucial role in educational assessments, enabling accurate and efficient assessment of student abilities and personalization to maximize learning outcomes. Traditionally, estimating item difficulties can…

计算与语言 · 计算机科学 2025-09-19 Alexander Scarlatos , Nigel Fernandez , Christopher Ormerod , Susan Lottridge , Andrew Lan

This paper presents a machine learning approach to multidimensional item response theory (MIRT), a class of latent factor models that can be used to model and predict student performance from observed assessment data. Inspired by…

机器学习 · 统计学 2025-01-08 Yoav Bergner , Peter F. Halpin , Jill-Jênn Vie

Comprehensive evaluations of language models (LM) during both development and deployment phases are necessary because these models possess numerous capabilities (e.g., mathematical reasoning, legal support, or medical diagnostic) as well as…

计算与语言 · 计算机科学 2025-03-18 Sang Truong , Yuheng Tu , Percy Liang , Bo Li , Sanmi Koyejo

For personalized ranking models, the well-calibrated probability of an item being preferred by a user has great practical value. While existing work shows promising results in image classification, probability calibration has not been much…

信息检索 · 计算机科学 2022-04-27 Wonbin Kweon , SeongKu Kang , Hwanjo Yu

Evaluation of large language models (LLMs) is increasingly critical, yet standard benchmarking methods rely on average accuracy, overlooking both the inherent stochasticity of LLM outputs and the heterogeneity of benchmark items. Item…

机器学习 · 统计学 2026-05-11 Xinhao Qu , Qiang Heng , Hao Zeng , Xiaoqian Liu

Stepped-wedge cluster-randomized trials (SW-CRTs) are widely used in healthcare and implementation science, providing an ethical advantage by ensuring all clusters eventually receive the intervention. The staggered rollout of treatment…

统计方法学 · 统计学 2026-04-03 Xi Fang , Xueqi Wang , Patrick J. Heagerty , Bingkai Wang , Fan Li

Item Response Theory (IRT) models aim to assess latent abilities of $n$ examinees along with latent difficulty characteristics of $m$ test items from categorical data that indicates the quality of their corresponding answers. Classical…

机器学习 · 计算机科学 2024-08-16 Susanne Frick , Amer Krivošija , Alexander Munteanu

Randomized experiments are considered the gold standard for estimating causal effects. However, out of the set of possible randomized assignments, some may be likely to produce poor effect estimates and misleading conclusions. Restricted…

统计方法学 · 统计学 2025-08-28 Maggie Wang , René F. Kizilcec , Michael Baiocchi

Computerized adaptive tests (CATs) play a crucial role in educational assessment and diagnostic screening in behavioral health. Unlike traditional linear tests that administer a fixed set of pre-assembled items, CATs adaptively tailor the…

统计方法学 · 统计学 2026-05-11 Jiguang Li , Robert Gibbons , Veronika Rockova

Individualized treatment rules (ITR) can improve health outcomes by recognizing that patients may respond differently to treatment and assigning therapy with the most desirable predicted outcome for each individual. Flexible and efficient…

统计方法学 · 统计学 2017-09-25 Brent R. Logan , Rodney Sparapani , Robert E. McCulloch , Purushottam W. Laud

Marginal maximum likelihood estimation (MMLE) in item response theory (IRT) is highly sensitive to aberrant responses, such as careless answering and random guessing, which can reduce estimation accuracy. To address this issue, this study…

统计方法学 · 统计学 2025-02-18 Yuki Itaya , Kenichi Hayashi

Item Response Theory (IRT) is a well known method for assessing responses from humans in education and psychology. In education, IRT is used to infer student abilities and characteristics of test items from student responses. Interactions…

人工智能 · 计算机科学 2023-07-20 Antti Keurulainen , Isak Westerlund , Oskar Keurulainen , Andrew Howes

Item response theory (IRT) models for categorical response data are widely used in the analysis of educational data, computerized adaptive testing, and psychological surveys. However, most IRT models rely on both the assumption that…

机器学习 · 统计学 2015-01-14 Ryan Ning , Andrew E. Waters , Christoph Studer , Richard G. Baraniuk

Item response theory (IRT) models typically rely on a normality assumption for subject-specific latent traits, which is often unrealistic in practice. Semiparametric extensions based on Dirichlet process mixtures offer a more flexible…

Robustness is of central importance in machine learning and has given rise to the fields of domain generalization and invariant learning, which are concerned with improving performance on a test distribution distinct from but related to the…

机器学习 · 计算机科学 2020-12-03 Robert Adragna , Elliot Creager , David Madras , Richard Zemel