中文
相关论文

相关论文: Random forests for binary geospatial data

200 篇论文

Mixed-effects logistic regression is widely used for binary outcomes in hierarchical data, yet formal goodness-of-fit tests remain limited to random-intercept models and do not address sparse cluster settings. We extend a grouping-based…

统计方法学 · 统计学 2026-04-22 Ariel Linden

Covariate shift occurs prevalently in practice, where the input distributions of the source and target data are substantially different. Despite its practical importance in various learning problems, most of the existing methods only focus…

机器学习 · 统计学 2023-10-20 Xingdong Feng , Xin He , Caixing Wang , Chao Wang , Jingnan Zhang

Over the past decade, random forest models have become widely used as a robust method for high-dimensional data regression tasks. In part, the popularity of these models arises from the fact that they require little hyperparameter tuning…

机器学习 · 计算机科学 2020-03-18 Shipra Malhotra , John Karanicolas

Remote sensing observations are extensively used for analysis of environmental variables. These variables often exhibit spatial correlation, which has to be accounted for in the calibration models used in predictions, either by direct…

应用统计 · 统计学 2017-02-14 Virpi Junttila , Marko Laine

Standard geostatistical models assume second order stationarity of the underlying Random Function. In some instances, there is little reason to expect the spatial dependence structure to be stationary over the whole region of interest. In…

统计方法学 · 统计学 2014-12-04 Francky Fouedjio , Nicolas Desassis , Jacques Rivoirard

We propose to prune a random forest (RF) for resource-constrained prediction. We first construct a RF and then prune it to optimize expected feature cost & accuracy. We pose pruning RFs as a novel 0-1 integer program with linear constraints…

机器学习 · 统计学 2016-06-17 Feng Nan , Joseph Wang , Venkatesh Saligrama

Machine learning iterative imputation methods have been well accepted by researchers for imputing missing data, but they can be time-consuming when handling large datasets. To overcome this drawback, parallel computing strategies have been…

应用统计 · 统计学 2020-04-24 Shangzhi Hong , Yuqi Sun , Hanying Li , Henry S. Lynn

Gaussian Processes (GPs) are widely recognized as powerful non-parametric models for regression and classification. Traditional GP frameworks predominantly operate under the assumption that the inputs are either accurately known or subject…

系统与控制 · 电气工程与系统科学 2025-10-14 Muzaffar Qureshi , Tochukwu Elijah Ogri , Zachary I. Bell , Wanjiku A. Makumi , Rushikesh Kamalapurkar

Ensemble forecasting systems have advanced meteorology by providing probabilistic estimates of future states. Nonetheless, systematic biases often persist, making statistical post-processing essential. Traditional parametric post-processing…

应用统计 · 统计学 2026-02-17 Mária Lakatos

The random subspace method, known as the pillar of random forests, is good at making precise and robust predictions. However, there is not a straightforward way yet to combine it with deep learning. In this paper, we therefore propose…

机器学习 · 计算机科学 2020-09-16 Yun-Hao Cao , Jianxin Wu , Hanchen Wang , Joan Lasenby

For various applications, the relations between the dependent and independent variables are highly nonlinear. Consequently, for large scale complex problems, neural networks and regression trees are commonly preferred over linear models…

机器学习 · 计算机科学 2017-05-23 Samet Oymak , Mehrdad Mahdavi , Jiasi Chen

Large or very large spatial (and spatio-temporal) datasets have become common place in many environmental and climate studies. These data are often collected in non-Euclidean spaces (such as the planet Earth) and they often present…

统计理论 · 数学 2023-01-09 Mike Pereira , Nicolas Desassis , Denis Allard

In this paper we recast the problem of missing values in the covariates of a regression model as a latent Gaussian Markov random field (GMRF) model in a fully Bayesian framework. Our proposed approach is based on the definition of the…

统计计算 · 统计学 2019-12-24 Virgilio Gómez-Rubio , Michela Cameletti , Marta Blangiardo

The Gaussian process (GP) regression can be severely biased when the data are contaminated by outliers. This paper presents a new robust GP regression algorithm that iteratively trims the most extreme data points. While the new algorithm…

机器学习 · 计算机科学 2021-06-15 Zhao-Zhou Li , Lu Li , Zhengyi Shao

Studying the effects of air-pollution on health is a key area in environmental epidemiology. An accurate estimation of air-pollution effects requires spatio-temporally resolved datasets of air-pollution, especially, Fine Particulate Matter…

应用统计 · 统计学 2019-03-27 Ron Sarafian , Itai Kloog , Allan C. Just , Johnathan D. Rosenblatt

We discuss an application of Generalized Random Forests (GRF) proposed by Athey et al.(2019) to quantile regression for time series data. We extracted the theoretical results of the GRF consistency for i.i.d. data to time series data. In…

统计理论 · 数学 2022-11-07 Hiroshi Shiraishi , Tomoshige Nakamura , Ryotato Shibuki

We develop a finite-sample, design-based theory for random forests in which each tree is a randomized conditional predictor acting on fixed covariates and the forest is their Monte Carlo average. An exact variance identity separates Monte…

机器学习 · 统计学 2026-03-03 Nathaniel S. O'Connell

Regional data analysis is concerned with the analysis and modeling of measurements that are spatially separated by specifically accounting for typical features of such data. Namely, measurements in close proximity tend to be more similar…

统计方法学 · 统计学 2023-08-15 Christoph Muehlmann , François Bachoc , Klaus Nordhausen

Generative Bayesian Filtering (GBF) provides a powerful and flexible framework for performing posterior inference in complex nonlinear and non-Gaussian state-space models. Our approach extends Generative Bayesian Computation (GBC) to…

统计方法学 · 统计学 2025-11-07 Edoardo Marcelli , Sean O'Hagan , Veronika Rockova

We propose a method for transfer learning in nonparametric regression using a random forest (RF) with distance covariance-based feature weights, assuming the unknown source and target regression functions are sparsely different. Our method…

机器学习 · 统计学 2026-03-17 Chenze Li , Subhadeep Paul