English
Related papers

Related papers: Small Area Estimation with Random Forests and the …

200 papers

Small area estimation has become an important tool in official statistics, used to construct estimates of population quantities for domains with small sample sizes. Typical area-level models function as a type of heteroscedastic regression,…

Methodology · Statistics 2022-09-07 Paul A. Parker , Scott H. Holan , Ryan Janicki

This paper presents a new method for spatially adaptive local (constant) likelihood estimation which applies to a broad class of nonparametric models, including the Gaussian, Poisson and binary response models. The main idea of the method…

Statistics Theory · Mathematics 2007-12-18 Denis Belomestny , Vladimir Spokoiny

This paper proposes a sparse regression method that continuously interpolates between Forward Stepwise selection (FS) and the LASSO. When tuned appropriately, our solutions are much sparser than typical LASSO fits but, unlike FS fits,…

Methodology · Statistics 2024-11-20 Ivy Zhang , Robert Tibshirani

Combining machine learning with econometric analysis is becoming increasingly prevalent in both research and practice. A common empirical strategy involves the application of predictive modeling techniques to 'mine' variables of interest…

Econometrics · Economics 2020-12-22 Mochen Yang , Edward McFowland , Gordon Burtch , Gediminas Adomavicius

Longitudinal analysis is important in many disciplines, such as the study of behavioral transitions in social science. Only very recently, feature selection has drawn adequate attention in the context of longitudinal modeling. Standard…

Methodology · Statistics 2016-10-26 Tingyang Xu , Jiangwen Sun , Jinbo Bi

Feature subsampling is a core component of random forests and other ensemble methods. While recent theory suggests that this randomization acts solely as a variance reduction mechanism analogous to ridge regularization, these results…

Machine Learning · Statistics 2026-01-06 Xin Chen , Jason M. Klusowski , Yan Shuo Tan , Chang Yu

We introduce random spatial forests, a method of bagging regression trees allowing for spatial correlation. Our main contribution is the development of a computationally efficient tree building algorithm which selects each split of the tree…

Methodology · Statistics 2020-07-24 Travis Hee Wai , Michael T. Young , Adam A. Szpiro

Survey sampling plays an important role in the efficient allocation and management of resources. The essence of survey sampling lies in acquiring a sample of data points from a population and subsequently using this sample to estimate the…

Methodology · Statistics 2024-01-29 Jonne Pohjankukka , Sakari Tuominen , Jukka Heikkonen

We analyze binary data, available for a relatively large number (big data) of families (or households), which are within small areas, from a population-based survey. Inference is required for the finite population proportion of individuals…

Methodology · Statistics 2018-06-04 Balgobin Nandram , Lu Chen , Shuting Fu , Binod Manandhar

Fine-tuned Large Language Models (LLMs) often suffer from overconfidence and poor calibration, particularly when fine-tuned on small datasets. To address these challenges, we propose a simple combination of Low-Rank Adaptation (LoRA) with…

Computation and Language · Computer Science 2024-07-23 Emre Onal , Klemens Flöge , Emma Caldwell , Arsen Sheverdin , Vincent Fortuin

``Localization'' has proven to be a valuable tool in the Statistical Learning literature as it allows sharp risk bounds in terms of the problem geometry. Localized bounds seem to be much less exploited in the Stochastic Optimization…

Optimization and Control · Mathematics 2023-03-30 Roberto I. Oliveira , Philip Thompson

Area-level models for small area estimation typically rely on areal random effects to shrink design-based direct estimates towards a model-based predictor. Incorporating the spatial dependence of the random effects into these models can…

Methodology · Statistics 2024-04-22 Sho Kawano , Paul A. Parker , Zehang Richard Li

Oversampled adaptive sensing (OAS) is a recently proposed Bayesian framework which sequentially adapts the sensing basis. In OAS, estimation quality is, in each step, measured by conditional mean squared errors (MSEs), and the basis for the…

Information Theory · Computer Science 2018-11-16 Ralf R. Müller , Ali Bereyhi , Christoph F. Mecklenbräuker

Sparse convex clustering is to cluster observations and conduct variable selection simultaneously in the framework of convex clustering. Although a weighted $L_1$ norm is usually employed for the regularization term in sparse convex…

Machine Learning · Statistics 2020-05-27 Kaito Shimamura , Shuichi Kawano

Accurate wind power forecasts depend on reliable wind speed forecasts. Numerical Weather Predictions (NWPs) utilize huge amounts of computing time, but still have rather low spatial and temporal resolution. However, stochastic wind speed…

Applications · Statistics 2015-09-10 Daniel Ambach , Carsten Croonenbroeck

In materials science, data-driven methods accelerate material discovery and optimization while reducing costs and improving success rates. Symbolic regression is a key to extracting material descriptors from large datasets, in particular…

Machine Learning · Computer Science 2024-10-01 Xiaolin Jiang , Guanqi Liu , Jiaying Xie , Zhenpeng Hu

Systematic sampling is often used to select plot locations for forest inventory estimation. However, it is not possible to derive a design-unbiased variance estimator for a systematic sample using one random start. As a result, many forest…

Applications · Statistics 2018-10-22 Chad Babcock , Andrew O. Finley , Timothy G. Gregoire , Hans-Erik Andersen

We propose generalized random forests, a method for non-parametric statistical estimation based on random forests (Breiman, 2001) that can be used to fit any quantity of interest identified as the solution to a set of local moment…

Methodology · Statistics 2018-04-06 Susan Athey , Julie Tibshirani , Stefan Wager

We present and apply methodology to improve inference for small area parameters by using data from several sources. This work extends Cahoy and Sedransk (2023) who showed how to integrate summary statistics from several sources. Our…

Methodology · Statistics 2024-12-12 D Cahoy , J Sedransk

There has been much recent work on inference after model selection when the noise level is known, however, $\sigma$ is rarely known in practice and its estimation is difficult in high-dimensional settings. In this work we propose using the…

Statistics Theory · Mathematics 2017-02-13 Xiaoying Tian , Joshua R. Loftus , Jonathan E. Taylor