English
Related papers

Related papers: Wasserstein Random Forests and Applications in Het…

200 papers

We propose a modification that corrects for split-improvement variable importance measures in Random Forests and other tree-based methods. These methods have been shown to be biased towards increasing the importance of features with more…

Machine Learning · Statistics 2020-03-25 Zhengze Zhou , Giles Hooker

Impractical assumptions, an inherently myopic nature, and the crucial role of the initial design, all together contribute to making theoretical convergence proofs of little value in real-life Bayesian Optimization applications. In this…

Optimization and Control · Mathematics 2026-02-13 Antonio Candelieri , Francesco Archetti

The purpose of the present work is to construct estimators for the random effects in a fractional diffusion model using a hybrid estimation method where we combine parametric and nonparametric thechniques. We precisely consider $n$…

Statistics Theory · Mathematics 2025-06-13 Nesrine Chebli , Hamdi Fathallah , Yousri Slaoui

One of the most fundamental problems in causal inference is the estimation of a causal effect when variables are confounded. This is difficult in an observational study, because one has no direct evidence that all confounders have been…

Machine Learning · Statistics 2014-11-03 Ricardo Silva , Robin Evans

Wasserstein distances are widely used in modern data analysis but pose significant computational and statistical challenges in high dimensions. The sliced Wasserstein distance alleviates these challenges by leveraging one-dimensional…

Statistics Theory · Mathematics 2026-05-21 David Rodríguez-Vítores , Eustasio del Barrio , Jean-Michel Loubes

Gradient boosting is a sequential ensemble method that fits a new weaker learner to pseudo residuals at each iteration. We propose Wasserstein gradient boosting, a novel extension of gradient boosting that fits a new weak learner to…

Methodology · Statistics 2024-08-30 Takuo Matsubara

In this paper, we propose a new random forest algorithm that constructs the trees using a novel adaptive split-balancing method. Rather than relying on the widely-used random feature selection, we propose a permutation-based balanced…

Machine Learning · Statistics 2024-09-04 Yuqian Zhang , Weijie Ji , Jelena Bradic

We introduce random survival forests, a random forests method for the analysis of right-censored survival data. New survival splitting rules for growing survival trees are introduced, as is a new missing data algorithm for imputing missing…

Applications · Statistics 2008-11-12 Hemant Ishwaran , Udaya B. Kogalur , Eugene H. Blackstone , Michael S. Lauer

In this work clustering schemes for uncertain and structured data are considered relying on the notion of Wasserstein barycenters, accompanied by appropriate clustering indices based on the intrinsic geometry of the Wasserstein space where…

Previous work on causal inference has primarily focused on averages and conditional averages of treatment effects, with significantly less attention on variability and uncertainty in individual treatment responses. In this paper, we…

Machine Learning · Computer Science 2026-02-10 Liyuan Xu , Bijan Mazaheri

Random forest regression is a powerful non-parametric method that adapts to local data characteristics through data-driven partitioning, making it effective across diverse application domains. However, the piecewise constant nature of…

Machine Learning · Computer Science 2026-05-19 Ziyi Liu , Phuc Luong , Mario Boley , Daniel F. Schmidt

Understanding proper distance measures between distributions is at the core of several learning tasks such as generative models, domain adaptation, clustering, etc. In this work, we focus on mixture distributions that arise naturally in…

Machine Learning · Computer Science 2019-10-30 Yogesh Balaji , Rama Chellappa , Soheil Feizi

The wealth of data being gathered about humans and their surroundings drives new machine learning applications in various fields. Consequently, more and more often, classifiers are trained using not only numerical data but also complex data…

Machine Learning · Computer Science 2022-04-13 Maciej Piernik , Dariusz Brzezinski , Pawel Zawadzki

We study data-driven decision problems where historical observations are generated by a time-evolving distribution whose consecutive shifts are bounded in Wasserstein distance. We address this nonstationarity using a distributionally robust…

Optimization and Control · Mathematics 2025-12-25 Dominic S. T. Keehan , Edward J. Anderson , Wolfram Wiesemann

We provide statistical theory for conditional and unconditional Wasserstein generative adversarial networks (WGANs) in the framework of dependent observations. We prove upper bounds for the excess Bayes risk of the WGAN estimators with…

Statistics Theory · Mathematics 2020-11-09 Moritz Haas , Stefan Richter

Distributionally-robust optimization is often studied for a fixed set of distributions rather than time-varying distributions that can drift significantly over time (which is, for instance, the case in finance and sociology due to…

Optimization and Control · Mathematics 2020-10-01 Iman Shames , Farhad Farokhi

We study nonparametric density estimation problems where error is measured in the Wasserstein distance, a metric on probability distributions popular in many areas of statistics and machine learning. We give the first minimax-optimal rates…

Statistics Theory · Mathematics 2020-04-30 Jonathan Niles-Weed , Quentin Berthet

Principal stratification analysis evaluates how causal effects of a treatment on a primary outcome vary across strata of units defined by their treatment effect on some intermediate quantity. This endeavor is substantially challenged when…

Methodology · Statistics 2024-03-21 Chanmin Kim , Corwin Zigler

Regression trees and random forests are popular and effective non-parametric estimators in practical applications. A recent paper by Athey and Wager shows that the random forest estimate at any point is asymptotically Gaussian; in this…

Econometrics · Economics 2021-02-02 Kevin Li

Recently, there has been great interest in estimating the conditional average treatment effect using flexible machine learning methods. However, in practice, investigators often have working hypotheses about effect heterogeneity across…

Methodology · Statistics 2023-09-13 Chan Park , Hyunseung Kang
‹ Prev 1 8 9 10 Next ›