中文
相关论文

相关论文: MDA for random forests: inconsistency, and a pract…

200 篇论文

Random Forests have become a widely used tool in machine learning since their introduction in 2001, known for their strong performance in classification and regression tasks. One key feature of Random Forests is the Random Forest…

统计理论 · 数学 2025-12-18 Nico Föge , Lena Schmid , Marc Ditzhaus , Markus Pauly

Standard supervised learning procedures are validated against a test set that is assumed to have come from the same distribution as the training data. However, in many problems, the test data may have come from a different distribution. We…

机器学习 · 统计学 2019-08-28 Tim Coleman , Kimberly Kaufeld , Mary Frances Dorn , Lucas Mentch

Unlike the ordinary least-squares (OLS) estimator for the linear model, a ridge regression linear model provides coefficient estimates via shrinkage, usually with improved mean-square and prediction error. This is true especially when the…

统计方法学 · 统计学 2015-06-25 George Karabatsos

Data augmentation (DA) techniques aim to increase data variability, and thus train deep networks with better generalisation. The pioneering AutoAugment automated the search for optimal DA policies with reinforcement learning. However,…

计算机视觉与模式识别 · 计算机科学 2020-07-31 Yonggang Li , Guosheng Hu , Yongtao Wang , Timothy Hospedales , Neil M. Robertson , Yongxin Yang

Sufficient dimension reduction reduces the dimensionality of data while preserving relevant regression information. In this article, we develop Minimum Average Deviance Estimation (MADE) methodology for sufficient dimension reduction. It…

统计方法学 · 统计学 2024-01-19 Kofi P. Adragni , Andrew M. Raim , Elias Al-Najjar

We consider a regression setting where observations are collected in different environments modeled by different data distributions. The field of out-of-distribution (OOD) generalization aims to design methods that generalize better to test…

机器学习 · 统计学 2026-03-12 Francesco Freni , Anya Fries , Linus Kühne , Markus Reichstein , Jonas Peters

In recent years, we have witnessed a surge of interests in learning a suitable distance metric from weakly supervised data. Most existing methods aim to pull all the similar samples closer while push the dissimilar ones as far as possible.…

机器学习 · 计算机科学 2021-02-05 Huiyuan Deng , Xiangzhu Meng , Lin Feng

Predictive uncertainty quantification is crucial in decision-making problems. We investigate how to adequately quantify predictive uncertainty with missing covariates. A bottleneck is that missing values induce heteroskedasticity on the…

统计方法学 · 统计学 2024-05-27 Margaux Zaffran , Julie Josse , Yaniv Romano , Aymeric Dieuleveut

In this article, we proposed a new probability distribution named as power Maxwell distribution (PMaD). It is another extension of Maxwell distribution (MaD) which would lead more flexibility to analyze the data with non-monotone failure…

应用统计 · 统计学 2018-07-04 Abhimanyu Singh Yadav , Hassan S. Bakouch , Sanjay Kumar Singh , Umesh Singh

Quadratic discriminant analysis (QDA) is a widely used statistical tool to classify observations from different multivariate Normal populations. The generalized quadratic discriminant analysis (GQDA) classification rule/classifier, which…

统计方法学 · 统计学 2020-04-15 Abhik Ghosh , Rita SahaRay , Sayan Chakrabarty , Sayan Bhadra

Tree-based ensemble methods, as Random Forests and Gradient Boosted Trees, have been successfully used for regression in many applications and research studies. Furthermore, these methods have been extended in order to deal with uncertainty…

机器学习 · 计算机科学 2018-11-20 Myriam Tami , Marianne Clausel , Emilie Devijver , Adrien Dulac , Eric Gaussier , Stefan Janaqi , Meriam Chebre

Anomaly detection (AD) plays a pivotal role in multimedia applications for detecting defective products and automating quality inspection. Deep learning (DL) models typically require large-scale annotated data, which are often highly…

计算机视觉与模式识别 · 计算机科学 2024-10-22 Eirini Cholopoulou , Dimitris K. Iakovidis

The maximum mean discrepancy (MMD) is a kernel-based distance between probability distributions useful in many applications (Gretton et al. 2012), bearing a simple estimator with pleasing computational and statistical properties. Being able…

机器学习 · 统计学 2022-11-16 Danica J. Sutherland , Namrata Deka

This study explores the classification error of Mixture Discriminant Analysis (MDA) in scenarios where the number of mixture components exceeds those present in the actual data distribution, a condition known as overspecification. We use a…

Accurate and rapid structural damage assessment (SDA) is crucial for post-disaster management, helping responders prioritise resources, plan rescues, and support recovery. Traditional field inspections, though precise, are limited by…

人工智能 · 计算机科学 2026-04-14 Wanli Ma , Sivasakthy Selvakumaran , Dain G. Farrimond , Adam A. Dennis , Samuel E. Rigby

This chapter assesses the sensitivity of multi-criteria decision-making (MCDM) methods to modifications within the decision or objective matrix (DOM) in the context of chemical engineering optimization applications. Employing eight common…

化学物理 · 物理学 2025-07-11 Seyed Reza Nabavi , Zhiyuan Wang , Gade Pandu Rangaiah

Random forests are an ensemble method relevant for many problems, such as regression or classification. They are popular due to their good predictive performance (compared to, e.g., decision trees) requiring only minimal tuning of…

统计方法学 · 统计学 2022-10-20 Nikolaus Umlauf , Nadja Klein

Remarkable progress has been made in difference-in-differences (DID) approaches to causal inference that estimate the average effect of a treatment on the treated (ATT). Of these, the semiparametric DID (SDID) approach incorporates a…

统计方法学 · 统计学 2026-03-09 Takamichi Baba , Yoshiyuki Ninomiya

Despite the simplicity and intuitive interpretation of Minimum Mean Squared Error (MMSE) estimators, their effectiveness in certain scenarios is questionable. Indeed, minimizing squared errors on average does not provide any form of…

最优化与控制 · 数学 2019-12-09 Dionysios S. Kalogerias , Luiz F. O. Chamon , George J. Pappas , Alejandro Ribeiro

Random forest is a classification algorithm well suited for microarray data: it shows excellent performance even when most predictive variables are noise, can be used when the number of variables is much larger than the number of…

定量方法 · 定量生物学 2007-05-23 Ramon Diaz-Uriarte , Sara Alvarez de Andres