English
Related papers

Related papers: Item response parameter estimation performance usi…

200 papers

Comprehensive evaluations of language models (LM) during both development and deployment phases are necessary because these models possess numerous capabilities (e.g., mathematical reasoning, legal support, or medical diagnostic) as well as…

Computation and Language · Computer Science 2025-03-18 Sang Truong , Yuheng Tu , Percy Liang , Bo Li , Sanmi Koyejo

Recent years have witnessed a surge in the number of large language models (LLMs), yet efficiently managing and utilizing these vast resources remains a significant challenge. In this work, we explore how to learn compact representations of…

Artificial Intelligence · Computer Science 2025-10-02 Jianhao Chen , Chenxu Wang , Gengrui Zhang , Peng Ye , Lei Bai , Wei Hu , Yuzhong Qu , Shuyue Hu

Longitudinal settings involving outcome, competing risks and censoring events occurring and recurring in continuous time are common in medical research, but are often analyzed with methods that do not allow for taking post-baseline…

Methodology · Statistics 2025-04-14 Helene C. W. Rytgaard , Mark J. van der Laan

The win ratio (WR) is a widely used metric to compare treatments in randomized clinical trials with hierarchically ordered endpoints. Counting-based approaches, such as Pocock's algorithm, are the standard for WR estimation. However, this…

Methodology · Statistics 2026-02-17 Yi Liu , Huiman Barnhart , Sean O'Brien , Yuliya Lokhnygina , Roland A. Matsouaka

Item (question) difficulties play a crucial role in educational assessments, enabling accurate and efficient assessment of student abilities and personalization to maximize learning outcomes. Traditionally, estimating item difficulties can…

Computation and Language · Computer Science 2025-09-19 Alexander Scarlatos , Nigel Fernandez , Christopher Ormerod , Susan Lottridge , Andrew Lan

Reading comprehension is a key for individual success, yet the assessment of question difficulty remains challenging due to the extensive human annotation and large-scale testing required by traditional methods such as linguistic analysis…

Computation and Language · Computer Science 2025-02-26 Yoshee Jain , John Hollander , Amber He , Sunny Tang , Liang Zhang , John Sabatini

Joint maximum likelihood (JML) estimation is one of the earliest approaches to fitting item response theory (IRT) models. This procedure treats both the item and person parameters as unknown but fixed model parameters and estimates them…

Methodology · Statistics 2019-06-17 Yunxiao Chen , Xiaoou Li , Siliang Zhang

The generalized partial credit model (GPCM) is a popular polytomous IRT model that has been widely used in large-scale educational surveys and health care services. Same as other IRT models, GPCM can be estimated via marginal maximum…

Applications · Statistics 2018-09-21 Yong Luo

Psychological assessments commonly rely on rating-scale items, which require respondents to condense complex experiences into predefined categories. Although rich, unstructured text is often captured alongside these scales, it rarely…

Computation and Language · Computer Science 2026-03-20 Joe Watson , Ivan O'Connor , Chia-Wen Chen , Luning Sun , Fang Luo , David Stillwell

Large language models (LLMs) achieve high performance on mathematical reasoning, but these results can be inflated by training data leakage or superficial pattern matching rather than genuine reasoning. To this end, an adversarial…

Computation and Language · Computer Science 2026-02-03 Xinyuan Li , Murong Xu , Wenbiao Tao , Hanlun Zhu , Yike Zhao , Jipeng Zhang , Yunshi Lan

This is a technical report which explores the estimation methodologies on hyper-parameters in Markov Random Field and Gaussian Hidden Markov Random Field. In first section, we briefly investigate a theoretical framework on…

Machine Learning · Statistics 2017-11-22 Namjoon Suh

In this paper, we study a generalization of the two-groups model in the presence of covariates --- a problem that has recently received much attention in the statistical literature due to its applicability in multiple hypotheses testing…

Methodology · Statistics 2019-02-01 Nabarun Deb , Sujayam Saha , Adityanand Guntuboyina , Bodhisattva Sen

In this paper, we provide a novel method for the estimation of unknown parameters of the Gaussian Mixture Model (GMM) in Positron Emission Tomography (PET). A vast majority of PET imaging methods are based on reconstruction model that is…

Signal Processing · Electrical Eng. & Systems 2023-06-30 Tomislav Matulić , Damir Seršić

Long questionnaires increase the response burden for patients and healthcare workers. In the treatment of Parkinson's disease, the MDS-UPDRS questionnaire to track disease progression may be underutilized due to time requirements. While…

Methodology · Statistics 2026-04-14 Karl Sigfrid , Ellinor Fackle-Fornius , Frank Miller

Self-assessment is a key aspect of reliable intelligence, yet evaluations of large language models (LLMs) focus mainly on task accuracy. We adapted the 10-item General Self-Efficacy Scale (GSES) to elicit simulated self-assessments from ten…

Artificial Intelligence · Computer Science 2025-11-27 Daniel I Jackson , Emma L Jensen , Syed-Amad Hussain , Emre Sezgin

This paper deals with nonparametric maximum likelihood estimation for Gaussian locally stationary processes. Our nonparametric MLE is constructed by minimizing a frequency domain likelihood over a class of functions. The asymptotic behavior…

Statistics Theory · Mathematics 2011-11-10 Rainer Dahlhaus , Wolfgang Polonik

Imputation is a popular technique for handling item nonresponse in survey sampling. Parametric imputation is based on a parametric model for imputation and is less robust against the failure of the imputation model. Nonparametric imputation…

Methodology · Statistics 2019-09-20 Danhyang Lee , Jae Kwang Kim

Natural Language Recommendation (NLRec) generates item suggestions based on the relevance between user-issued NL requests and NL item description passages. Existing NLRec approaches often use Dense Retrieval (DR) to compute item relevance…

Information Retrieval · Computer Science 2025-11-04 Yifan Liu , Qianfeng Wen , Jiazhou Liang , Mark Zhao , Justin Cui , Anton Korikov , Armin Toroghi , Junyoung Kim , Scott Sanner

In this article the package High-dimensional Metrics (\texttt{hdm}) is introduced. It is a collection of statistical methods for estimation and quantification of uncertainty in high-dimensional approximately sparse models. It focuses on…

Methodology · Statistics 2017-09-28 Victor Chernozhukov , Chris Hansen , Martin Spindler

In a recent review, Liu, Pek, & Maydeu-Olivares (2025b) classified reliability coefficients into two types: classical test theory (CTT) reliability and proportional reduction in mean squared error (PRMSE). This article focuses on…

Methodology · Statistics 2026-04-14 Youjin Sung , Yang Liu
‹ Prev 1 3 4 5 6 7 10 Next ›