English
Related papers

Related papers: On cross-validation for small area estimators

200 papers

Is there really much more to say about sparse autoencoders (SAEs)? Autoencoders in general, and SAEs in particular, represent deep architectures that are capable of modeling low-dimensional latent structure in data. Such structure could…

Machine Learning · Computer Science 2025-06-09 Yin Lu , Xuening Zhu , Tong He , David Wipf

In small area estimation, it is a smart strategy to rely on data measured over time. However, linear mixed models struggle to properly capture time dependencies when the number of lags is large. Given the lack of published studies…

The determination of the sample size required by a crossover trial typically depends on the specification of one or more variance components. Uncertainty about the value of these parameters at the design stage means that there is often a…

Methodology · Statistics 2018-03-28 Michael Grayling , Adrian Mander , James Wason

Cross-validation is a well-known and widely used bandwidth selection method in nonparametric regression estimation. However, this technique has two remarkable drawbacks: (i) the large variability of the selected bandwidths, and (ii) the…

Methodology · Statistics 2021-05-11 D. Barreiro-Ures , R. Cao , M. Francisco-Fernández

We propose a general semi-supervised inference framework focused on the estimation of the population mean. As usual in semi-supervised settings, there exists an unlabeled sample of covariate vectors and a labeled sample consisting of…

Methodology · Statistics 2018-08-15 Anru Zhang , Lawrence D. Brown , T. Tony Cai

Sparse autoencoders (SAEs) have recently emerged as a powerful tool for interpreting the features learned by large language models (LLMs). By reconstructing features with sparsely activated networks, SAEs aim to recover complex superposed…

Machine Learning · Computer Science 2026-03-05 Jingyi Cui , Qi Zhang , Yifei Wang , Yisen Wang

We derive high-dimensional Gaussian comparison results for the standard $V$-fold cross-validated risk estimates. Our results combine a recent stability-based argument for the low-dimensional central limit theorem of cross-validation with…

Statistics Theory · Mathematics 2023-11-15 Nicholas Kissel , Jing Lei

This paper addresses feature subset selection for Support Vector Machines (SVMs) based on the cross-validation criterion. Unlike statistical criteria such as the Akaike information criterion (AIC) and the Bayesian information criterion…

Optimization and Control · Mathematics 2026-05-11 Masaharu Mori , Shunnosuke Ikeda , Ryuta Tamura , Yuichi Takano , Ryuhei Miyashiro

Predictive modelling using machine learning has become very popular for spatial mapping of the environment. Models are often applied to make predictions far beyond sampling locations where new geographic locations might considerably differ…

Machine Learning · Statistics 2021-08-20 Hanna Meyer , Edzer Pebesma

While many statistical models and methods are now available for network analysis, resampling network data remains a challenging problem. Cross-validation is a useful general tool for model selection and parameter tuning, but is not directly…

Methodology · Statistics 2020-05-04 Tianxi Li , Elizaveta Levina , Ji Zhu

When analyzing spatially referenced event data, the criteria for declaring rates as "reliable" is still a matter of dispute. What these varying criteria have in common, however, is that they are rarely satisfied for crude estimates in small…

Methodology · Statistics 2023-02-10 Harrison Quick , Guangzi Song

This paper devises a fully Bayesian sample size determination method for hierarchical model-based small area estimation with a decision risk approach. A new loss function specified around a desired maximum posterior variance target…

Methodology · Statistics 2018-02-27 Peter Dutey-Magni

Scalable oversight protocols aim to empower evaluators to accurately verify AI models more capable than themselves. However, human evaluators are subject to biases that can lead to systematic errors. We conduct two studies examining the…

Human-Computer Interaction · Computer Science 2025-07-29 Gabriel Recchia , Chatrik Singh Mangat , Jinu Nyachhyon , Mridul Sharma , Callum Canavan , Dylan Epstein-Gross , Muhammed Abdulbari

Social and economic studies are often implemented as complex survey designs. For example, multistage, unequal probability sampling designs utilized by federal statistical agencies are typically constructed to maximize the efficiency of the…

Methodology · Statistics 2020-06-09 Matthew R. Williams , Terrance D. Savitsky

In this article we have suggested an improved estimator for estimating the population mean in simple random sampling using auxiliary information under the presence of measurement errors. The mean square error (MSE) of the proposed estimator…

Applications · Statistics 2013-12-05 Sachin Malik , Jayant Singh , Rajesh Singh

Sparse autoencoders (SAEs) are commonly used to interpret the internal activations of large language models (LLMs) by mapping them to human-interpretable concept representations. While existing evaluations of SAEs focus on metrics such as…

Machine Learning · Computer Science 2026-01-26 Aaron J. Li , Suraj Srinivas , Usha Bhalla , Himabindu Lakkaraju

With the rise in popularity of digital Atlases to communicate spatial variation, there is an increasing need for robust small-area estimates. However, current small-area estimation methods suffer from various modeling problems when data are…

Methodology · Statistics 2023-12-06 James Hogg , Jessica Cameron , Susanna Cramb , Peter Baade , Kerrie Mengersen

This paper considers the estimation of binary choice models when survey responses are possibly misclassified but one of the response category can be validated. Partial validation may occur when survey questions about participation include…

Econometrics · Economics 2025-12-17 Augustine Denteh , Pierre E. Nguimkeu

In real applications of small area estimation, one often encounters data with positive response values. The use of a parametric transformation for positive response values in the Fay-Herriot model is proposed for such a case. An…

Methodology · Statistics 2017-03-31 Shonosuke Sugasawa , Tatsuya Kubokawa

We consider the problem in precision health of grouping people into subpopulations based on their degree of vulnerability to a risk factor. These subpopulations cannot be discovered with traditional clustering techniques because their…

Machine Learning · Statistics 2018-12-11 Alexander New , Kristin P. Bennett