中文
相关论文

相关论文: A Calibrated Data-Driven Approach for Small Area E…

200 篇论文

A common problem faced by statistical institutes is that data may be missing from collected data sets. The typical way to overcome this problem is to impute the missing data. The problem of imputing missing data is complicated by the fact…

应用统计 · 统计学 2014-01-09 Jeroen Pannekoek , Natalie Shlomo , Ton De Waal

In this paper we propose a flexible nested error regression small area model with high dimensional parameter that incorporates heterogeneity in regression coefficients and variance components. We develop a new robust small area specific…

统计方法学 · 统计学 2022-01-26 Partha Lahiri , Nicola Salvati

Constant (naive) imputation is still widely used in practice as this is a first easy-to-use technique to deal with missing data. Yet, this simple method could be expected to induce a large bias for prediction purposes, as the imputed input…

统计理论 · 数学 2024-02-07 Alexis Ayme , Claire Boyer , Aymeric Dieuleveut , Erwan Scornet

In analyzing big data for finite population inference, it is critical to adjust for the selection bias in the big data. In this paper, we propose two methods of reducing the selection bias associated with the big data sample. The first…

统计方法学 · 统计学 2019-01-08 Jae Kwang Kim , Zhonglei Wang

Area-specific causal inference is important in many policy and survey applications, where the goal is to evaluate treatment effects for small geographic or demographic domains. Existing causal small area estimation methods, however,…

统计理论 · 数学 2026-05-06 Tsubasa Ito , Shonosuke Sugasawa

Methods for reasoning under uncertainty are a key building block of accurate and reliable machine learning systems. Bayesian methods provide a general framework to quantify uncertainty. However, because of model misspecification and the use…

机器学习 · 计算机科学 2018-07-03 Volodymyr Kuleshov , Nathan Fenner , Stefano Ermon

Big data mining is well known to be an important task for data science, because it can provide useful observations and new knowledge hidden in given large datasets. Proximity-based data analysis is particularly utilized in many real-life…

数据库 · 计算机科学 2022-11-29 Daichi Amagata , Yusuke Arai , Sumio Fujita , Takahiro Hara

Accurate quantification of uncertainty is crucial for real-world applications of machine learning. However, modern deep neural networks still produce unreliable predictive uncertainty, often yielding over-confident predictions. In this…

机器学习 · 计算机科学 2020-10-29 Peng Cui , Wenbo Hu , Jun Zhu

For small area estimation of area-level data, the Fay-Herriot model is extensively used as a model based method. In the Fay-Herriot model, it is conventionally assumed that the sampling variances are known whereas estimators of sampling…

统计方法学 · 统计学 2017-05-15 Shonosuke Sugasawa , Hiromasa Tamae , Tatsuya Kubokawa

It is now well known that neural networks can be wrong with high confidence in their predictions, leading to poor calibration. The most common post-hoc approach to compensate for this is to perform temperature scaling, which adjusts the…

计算机视觉与模式识别 · 计算机科学 2022-07-25 Tom Joy , Francesco Pinto , Ser-Nam Lim , Philip H. S. Torr , Puneet K. Dokania

We consider a small area estimation model under square-root transformation in the presence of functional measurement error. When measurement error is present, the Bayes predictor can no longer be used as it depends on the covariates even if…

统计方法学 · 统计学 2023-09-27 Ka Long Keith Ho , Masayo Y. Hirose , Malay Ghosh

Sampling biases in training data are a major source of algorithmic biases in machine learning systems. Although there are many methods that attempt to mitigate such algorithmic biases during training, the most direct and obvious way is…

机器学习 · 统计学 2022-04-15 Laura Niss , Yuekai Sun , Ambuj Tewari

Aggregate outcome variables collected through surveys and administrative records are often subject to systematic measurement error. For instance, in disaster loss databases, county-level losses reported may differ from the true damages due…

机器学习 · 计算机科学 2026-03-17 Saketh Vishnubhatla , Shu Wan , Andre Harrison , Adrienne Raglin , Huan Liu

Calibration is an essential step in radio interferometric data processing that corrects the data for systematic errors and in addition, subtracts bright foreground interference to reveal weak signals hidden in the residual. These weak and…

天体物理仪器与方法 · 物理学 2019-05-08 Sarod Yatawatta

Conformal inference is a statistical method used to construct prediction sets for point predictors, providing reliable uncertainty quantification with probability guarantees. This method utilizes historical labeled data to estimate the…

机器学习 · 计算机科学 2024-11-05 Xiaoyi Su , Zhixin Zhou , Rui Luo

Imputation of missing data in large regions of satellite imagery is necessary when the acquired image has been damaged by shadows due to clouds, or information gaps produced by sensor failure. The general approach for imputation of missing…

应用统计 · 统计学 2010-06-23 Valeria Rulloni , Oscar Bustos , Ana Georgina Flesia

Machine learning models are increasingly used to produce predictions that serve as input data in subsequent statistical analyses. For example, computer vision predictions of economic and environmental indicators based on satellite imagery…

统计方法学 · 统计学 2025-11-18 Dan M. Kluger , Kerri Lu , Tijana Zrnic , Sherrie Wang , Stephen Bates

Nested error regression models are useful tools for analysis of grouped data, especially in the case of small area estimation. This paper suggests a nested error regression model using uncertain random effects in which the random effect in…

统计方法学 · 统计学 2017-02-28 Shonosuke Sugasawa , Tatsuya Kubokawa

There is a growing need for flexible general frameworks that integrate individual-level data with external summary information for improved statistical inference. External information relevant for a risk prediction model may come in…

统计方法学 · 统计学 2023-04-11 Tian Gu , Jeremy M. G. Taylor , Bhramar Mukherjee

Proper scoring rules evaluate the quality of probabilistic predictions, playing an essential role in the pursuit of accurate and well-calibrated models. Every proper score decomposes into two fundamental components -- proper calibration…