English
Related papers

Related papers: Efficient Detection of Bad Benchmark Items with No…

200 papers

While the volume of remote sensing data is increasing daily, deep learning in Earth Observation faces lack of accurate annotations for supervised optimization. Crowdsourcing projects such as OpenStreetMap distribute the annotation load to…

Computer Vision and Pattern Recognition · Computer Science 2024-03-05 Chenying Liu , Conrad M Albrecht , Yi Wang , Qingyu Li , Xiao Xiang Zhu

Text-to-image (T2I) generation has advanced rapidly, making reliable evaluation critical as performance differences between models narrow. Existing evaluation practices typically apply uniform annotation mechanisms, such as Likert-scale or…

Computer Vision and Pattern Recognition · Computer Science 2026-05-14 Abdelrahman Eldesokey , Merey Ramazanova , Ahmad Sait , Ansar Khangeldin , Karen Sanchez , Tong Zhang , Bernard Ghanem

Standard uncertainty estimation techniques, such as dropout, often struggle to clearly distinguish reliable predictions from unreliable ones. We attribute this limitation to noisy classifier weights, which, while not impairing overall…

Machine Learning · Computer Science 2025-06-09 Haripriya Harikumar , Santu Rana

Generalized linear models are often misspecified due to overdispersion, heteroscedasticity and ignored nuisance variables. Existing quasi-likelihood methods for testing in misspecified models often do not provide satisfactory type-I error…

Methodology · Statistics 2020-05-13 Jesse Hemerik , Jelle J Goeman , Livio Finos

In order to use psychometric instruments to assess a multidimensional construct, we may decompose it in dimensions and, in order to assess each dimension, develop a set of items, so one may assess the construct as a whole, by assessing its…

Methodology · Statistics 2018-02-02 Diego Marcondes , Nilton Rogerio Marcondes

Much recent work on visual recognition aims to scale up learning to massive, noisily-annotated datasets. We address the problem of scaling- up the evaluation of such models to large-scale datasets with noisy labels. Current protocols for…

Computer Vision and Pattern Recognition · Computer Science 2018-07-03 Phuc Nguyen , Deva Ramanan , Charless Fowlkes

The integrated conditional moment (ICM) test is a classical and widely used method for assessing the adequacy of regression models. Although it performs well in fixed-dimension settings, its behavior changes dramatically when the predictor…

Methodology · Statistics 2026-04-17 Yue Hu , Haiqi Li , Xintao Xia

R2 score is the standard metric for evaluating regression tasks, offering a normalized magnitude-agnostic measure of accuracy that captures variance. However, R2 has three key limitations: it is limited to at most two dimensional inputs, it…

Machine Learning · Computer Science 2026-05-05 Jaesung Yoo , Stefan Lemke , Jian Zhong Guo , Kanaka Rajan , Adam Hantman

The field of AI-generated text detection has evolved from supervised classification to zero-shot statistical analysis. However, current approaches share a fundamental limitation: they aggregate token-level measurements into scalar scores,…

Computation and Language · Computer Science 2025-09-24 Alva West , Yixuan Weng , Minjun Zhu , Luodan Zhang , Zhen Lin , Guangsheng Bao , Yue Zhang

Item difficulty plays a crucial role in test performance, interpretability of scores, and equity for all test-takers, especially in large-scale assessments. Traditional approaches to item difficulty modeling rely on field testing and…

Computation and Language · Computer Science 2025-09-30 Sydney Peters , Nan Zhang , Hong Jiao , Ming Li , Tianyi Zhou , Robert Lissitz

Within the nonparametric regression model with unknown regression function $l$ and independent, symmetric errors, a new multiscale signed rank statistic is introduced and a conditional multiple test of the simple hypothesis $l=0$ against a…

Statistics Theory · Mathematics 2008-12-18 Angelika Rohde

State-of-the-art object detectors rely on regressing and classifying an extensive list of possible anchors, which are divided into positive and negative samples based on their intersection-over-union (IoU) with corresponding groundtruth…

Computer Vision and Pattern Recognition · Computer Science 2020-06-01 Hengduo Li , Zuxuan Wu , Chen Zhu , Caiming Xiong , Richard Socher , Larry S. Davis

When two variables are related by a known function, the coefficient of determination (denoted $R^2$) measures the proportion of the total variance in the observations that is explained by that function. This quantifies the strength of the…

Applications · Statistics 2013-03-11 Ben Murrell , Daniel Murrell , Hugh Murrell

The dissemination of Large Language Models (LLMs), trained at scale, and endowed with powerful text-generating abilities, has made it easier for all to produce harmful, toxic, faked or forged content. In response, various proposals have…

Computation and Language · Computer Science 2025-06-12 Matthieu Dubois , François Yvon , Pablo Piantanida

Evaluating and validating the performance of prediction models is a fundamental task in statistics, machine learning, and their diverse applications. However, developing robust performance metrics for competing risks time-to-event data…

Methodology · Statistics 2025-07-22 Zian Zhuang , Wen Su , Eric Kawaguchi , Gang Li

Motivated by global warming issues, we consider a time se- ries that consists of a nondecreasing trend observed with station- ary fluctuations, nonparametric estimation of the trend under monotonicity assumption is considered. The rescaled…

Statistics Theory · Mathematics 2008-12-18 Ou Zhao , Michael Woodroofe

This research proposes a new (old) metric for evaluating goodness of fit in topic models, the coefficient of determination, or $R^2$. Within the context of topic modeling, $R^2$ has the same interpretation that it does when used in a…

Information Retrieval · Computer Science 2019-11-27 Tommy Jones

The increase in the number of Internet of Things (IoT) devices has tremendously increased the attack surface of cyber threats thus making a strong intrusion detection system (IDS) with a clear explanation of the process essential towards…

Cryptography and Security · Computer Science 2026-01-22 Ashikuzzaman , Md. Shawkat Hossain , Jubayer Abdullah Joy , Md Zahid Akon , Md Manjur Ahmed , Md. Naimul Islam

Investigating the main determinants of the mechanical performance of metals is not a simple task. Already known physical inspired qualitative relations between 2D microstructure characteristics and 3D mechanical properties can act as the…

Applications · Statistics 2020-02-05 Martina Vittorietti , Javier Hidalgo , Jilt Sietsma , Wei Li , Geurt Jongbloed

We propose causal isotonic calibration, a novel nonparametric method for calibrating predictors of heterogeneous treatment effects. Furthermore, we introduce cross-calibration, a data-efficient variant of calibration that eliminates the…

Machine Learning · Statistics 2023-06-07 Lars van der Laan , Ernesto Ulloa-Pérez , Marco Carone , Alex Luedtke