English
Related papers

Related papers: Overlapping Batch Confidence Intervals on Statisti…

200 papers

Confidence measures for the generalization error are crucial when small training samples are used to construct classifiers. A common approach is to estimate the generalization error by resampling and then assume the resampled estimator…

Machine Learning · Computer Science 2012-06-18 Eric B. Laber , Susan A. Murphy

In eXplainable Artificial Intelligence (XAI), instance-based explanations for time series have gained increasing attention due to their potential for actionable and interpretable insights in domains such as healthcare. Addressing the…

Machine Learning · Computer Science 2026-01-21 Maciej Mozolewski , Betül Bayrak , Kerstin Bach , Grzegorz J. Nalepa

Bootstrap resampling is the foundation of many ensemble learning methods, and out-of-bag (OOB) error estimation is the most widely used internal measure of generalization performance. In the standard multinomial bootstrap, the number of…

Methodology · Statistics 2025-11-25 Cheng Peng

Statistical analysis of high-dimensional functional times series arises in various applications. Under this scenario, in addition to the intrinsic infinite-dimensionality of functional data, the number of functional variables can grow with…

Statistics Theory · Mathematics 2022-01-14 Qin Fang , Shaojun Guo , Xinghao Qiao

Randomized clinical trials are the gold standard when estimating the average treatment effect. However, they are usually not a random sample from the real-world population because of the inclusion/exclusion rules. Meanwhile, observational…

Methodology · Statistics 2024-12-11 Kuan Jiang , Wenjie Hu , Shu Yang , Xinxing Lai , Xiaohua Zhou

When the study variable is functional and storage capacities are limited or transmission costs are high, selecting with survey sampling techniques a small fraction of the observations is an interesting alternative to signal compression…

Statistics Theory · Mathematics 2013-02-15 Hervé Cardot , Camelia Goga , Pauline Lardin

The estimation of cumulative distribution functions (CDF) is an important learning task with a great variety of downstream applications, such as risk assessments in predictions and decision making. In this paper, we study functional…

Machine Learning · Computer Science 2024-03-11 Qian Zhang , Anuran Makur , Kamyar Azizzadenesheli

In the analysis of survey data it is of interest to estimate and quantify uncertainty about means or totals for each of several non-overlapping subpopulations, or areas. When the sample size for a given area is small, standard confidence…

Methodology · Statistics 2018-09-26 Kyle Burris , Peter Hoff

Conformal prediction builds marginally valid prediction intervals that cover the unknown outcome of a randomly drawn test point with a prescribed probability. However, in practice, data-driven methods are often used to identify specific…

Methodology · Statistics 2025-04-21 Ying Jin , Zhimei Ren

Moving object segmentation (MOS) on LiDAR point clouds is crucial for autonomous systems like self-driving vehicles. Previous supervised approaches rely heavily on costly manual annotations, while LiDAR sequences naturally capture temporal…

Computer Vision and Pattern Recognition · Computer Science 2025-10-03 Ziliang Miao , Runjian Chen , Yixi Cai , Buwei He , Wenquan Zhao , Wenqi Shao , Bo Zhang , Fu Zhang

The conditional average treatment effect (CATE) is widely used in personalized medicine to inform therapeutic decisions. However, state-of-the-art methods for CATE estimation (so-called meta-learners) often perform poorly in the presence of…

Machine Learning · Computer Science 2026-03-12 Valentyn Melnychuk , Dennis Frauen , Jonas Schweisthal , Stefan Feuerriegel

Interval-censored data are common in fields such as epidemiology and demography. When the failure event of interest is relatively rare and the collection of covariates is costly, researchers often adopt the case-cohort design to reduce…

Methodology · Statistics 2025-09-29 Yeyu Xiao , Yonghong Long

The construction of confidence intervals for the mean of a bounded random variable is a classical problem in statistics with numerous applications in machine learning and virtually all scientific fields. In particular, obtaining the…

Machine Learning · Computer Science 2025-11-12 Václav Voráček , Francesco Orabona

It can be argued that optimal prediction should take into account all available data. Therefore, to evaluate a prediction interval's performance one should employ conditional coverage probability, conditioning on all available observations.…

Statistics Theory · Mathematics 2021-03-02 Yunyi Zhang , Dimitris N. Politis

In time series analysis, statistics based on collections of estimators computed from sub-samples play a crucial role in an increasing variety of important applications. Proving results about the joint asymptotic distribution of such…

Statistics Theory · Mathematics 2013-05-27 Stanislav Volgushev , Xiaofeng Shao

Machine learning models are increasingly used to produce predictions that serve as input data in subsequent statistical analyses. For example, computer vision predictions of economic and environmental indicators based on satellite imagery…

Methodology · Statistics 2025-11-18 Dan M. Kluger , Kerri Lu , Tijana Zrnic , Sherrie Wang , Stephen Bates

Many modern datasets, such as those in ecology and geology, are composed of samples with spatial structure and dependence. With such data violating the usual independent and identically distributed (IID) assumption in machine learning and…

Methodology · Statistics 2023-10-18 Kevin Fry , Jonathan E. Taylor

In this article, we consider the problem of constructing the confidence interval and testing hypothesis for the common coefficient of variation (CV) of several normal populations. A new method is suggested using the concepts of generalized…

Statistics Theory · Mathematics 2014-05-05 Javad Behboodian , Ali Akbar Jafari

Estimating externally valid causal effects is a foundational problem in the social and biomedical sciences. Generalizing or transporting causal estimates from an experimental sample to a target population of interest relies on an overlap…

Methodology · Statistics 2024-03-29 Melody Huang

When multiple investigators analyze a common dataset, the data reuse induces dependence across testing procedures, affecting the distribution of errors. Existing techniques of managing dependent tests require either cross-study coordination…

Statistics Theory · Mathematics 2026-04-10 Reid Dale , Jordan Rodu , Maria E. Currie , Mike Baiocchi