English
Related papers

Related papers: Confidence intervals for the random forest general…

200 papers

As shown in recent research, deep neural networks can perfectly fit randomly labeled data, but with very poor accuracy on held out data. This phenomenon indicates that loss functions such as cross-entropy are not a reliable indicator of…

Machine Learning · Statistics 2019-06-13 Yiding Jiang , Dilip Krishnan , Hossein Mobahi , Samy Bengio

Confidence intervals are assessed according to two criteria, namely expected length and coverage probability. In an attempt to apply the decision-theoretic method to finding a good confidence interval, a loss function that is a linear…

Statistics Theory · Mathematics 2017-10-18 Paul Kabaila

In many experimental or quasi-experimental studies, outcomes of interest are only observed for subjects who select (or are selected) to engage in the activity generating the outcome. Outcome data is thus endogenously missing for units who…

Methodology · Statistics 2026-01-14 Cyrus Samii , Ye Wang , Junlong Aaron Zhou

We propose a method of estimating the uncertainty of a result obtained through extrapolation to the complete basis set limit. The method is based on an ensemble of random walks which simulate all possible extrapolation outcomes that could…

Data Analysis, Statistics and Probability · Physics 2025-08-26 Jakub Lang , Michał Przybytek , Michał Lesiuk

This paper is motivated by an open problem around deep networks, namely, the apparent absence of over-fitting despite large over-parametrization which allows perfect fitting of the training data. In this paper, we analyze this phenomenon in…

Machine Learning · Computer Science 2019-08-28 Hrushikesh Mhaskar , Tomaso Poggio

In this paper an easy to implement method of stochastically weighing short and long memory linear processes is introduced. The method renders asymptotically exact size confidence intervals for the population mean which are significantly…

Methodology · Statistics 2019-01-15 Masoud M Nasari , Mohamedou Ould-Haye

When digitizing a print bilingual dictionary, whether via optical character recognition or manual entry, it is inevitable that errors are introduced into the electronic version that is created. We investigate automating the process of…

Computation and Language · Computer Science 2014-11-03 Michael Bloodgood , Peng Ye , Paul Rodrigues , David Zajic , David Doermann

Recent work has demonstrated the utility of Random Forest (RF) proximities for various supervised machine learning tasks, including outlier detection, missing data imputation, and visualization. However, the utility of the RF proximities…

Machine Learning · Computer Science 2025-11-26 Ben Shaw , Adam Rustad , Sofia Pelagalli Maia , Jake S. Rhodes , Kevin R. Moon

We develop Clustered Random Forests, a random forests algorithm for clustered data, arising from independent groups that exhibit within-cluster dependence. The leaf-wise predictions for each decision tree making up clustered random forests…

Methodology · Statistics 2026-01-26 Elliot H. Young , Peter Bühlmann

Generalization in generative modeling is defined as the ability to learn an underlying distribution from a finite dataset and produce novel samples, with evaluation largely driven by held-out performance and perceived sample quality. In…

Machine Learning · Computer Science 2026-03-05 Jerome Garnier-Brun , Luca Biggio , Davide Beltrame , Marc Mézard , Luca Saglietti

We introduce an optimization-based reconstruction attack capable of completely or near-completely reconstructing a dataset utilized for training a random forest. Notably, our approach relies solely on information readily available in…

Machine Learning · Computer Science 2024-08-16 Julien Ferry , Ricardo Fukasawa , Timothée Pascal , Thibaut Vidal

We develop a finite-sample, design-based theory for random forests in which each tree is a randomized conditional predictor acting on fixed covariates and the forest is their Monte Carlo average. An exact variance identity separates Monte…

Machine Learning · Statistics 2026-03-03 Nathaniel S. O'Connell

Point estimation of class prevalences in the presence of data set shift has been a popular research topic for more than two decades. Less attention has been paid to the construction of confidence and prediction intervals for estimates of…

Machine Learning · Statistics 2019-07-23 Dirk Tasche

Random Forests (RFs) are widely used Machine Learning models in low-power embedded devices, due to their hardware friendly operation and high accuracy on practically relevant tasks. The accuracy of a RF often increases with the number of…

We demonstrate that, for a range of state-of-the-art machine learning algorithms, the differences in generalisation performance obtained using default parameter settings and using parameters tuned via cross-validation can be similar in…

Machine Learning · Computer Science 2017-03-21 Anthony Bagnall , Gavin C. Cawley

Generalized Bayes posterior distributions are formed by putting a fractional power on the likelihood before combining with the prior via Bayes's formula. This fractional power, which is often viewed as a remedy for potential model…

Methodology · Statistics 2023-04-12 Pei-Shien Wu , Ryan Martin

We introduce a random forest approach to enable spreads' prediction in the primary catastrophe bond market. We investigate whether all information provided to investors in the offering circular prior to a new issuance is equally important…

Pricing of Securities · Quantitative Finance 2020-01-29 Despoina Makariou , Pauline Barrieu , Yining Chen

We suggest general methods to construct asymptotically uniformly valid confidence intervals post-model-selection. The constructions are based on principles recently proposed by Berk et al. (2013). In particular the candidate models used can…

Statistics Theory · Mathematics 2017-11-15 François Bachoc , David Preinerstorfer , Lukas Steinberger

This paper examines the use of a residual bootstrap for bias correction in machine learning regression methods. Accounting for bias is an important obstacle in recent efforts to develop statistical inference for machine learning methods. We…

Machine Learning · Statistics 2015-06-02 Giles Hooker , Lucas Mentch

In this paper we first provide a method to compute confidence intervals for the center of a piecewise normal distribution given a sample from this distribution, under certain assumptions. We then extend this method to an asymptotic setting,…

Optimization and Control · Mathematics 2022-08-08 Shu Lu , Hongsheng Liu