English
Related papers

Related papers: The Spectral Condition Number Plot for Regularizat…

200 papers

The decision tree is a flexible machine learning model that finds its success in numerous applications. It is usually fitted in a recursively greedy manner using CART. In this paper, we investigate the convergence rate of CART under a…

Machine Learning · Statistics 2023-10-27 Rahul Mazumder , Haoyue Wang

The generalized Ridge penalty is a powerful tool for dealing with overfitting and for high-dimensional regressions. The generalized Ridge regression can be derived as the mean of a posterior distribution with a Normal prior and a given…

Methodology · Statistics 2022-08-10 Said Obakrim , Pierre Ailliot , Valérie Monbet , Nicolas Raillard

Feature selection is a standard approach to understanding and modeling high-dimensional classification data, but the corresponding statistical methods hinge on tuning parameters that are difficult to calibrate. In particular, existing…

Methodology · Statistics 2019-03-01 Wei Li , Johannes Lederer

Analyzing classification model performance is a crucial task for machine learning practitioners. While practitioners often use count-based metrics derived from confusion matrices, like accuracy, many applications, such as weather…

Human-Computer Interaction · Computer Science 2022-07-29 Peter Xenopoulos , Joao Rulff , Luis Gustavo Nonato , Brian Barr , Claudio Silva

We consider the sensitivity of real roots of polynomial systems with respect to perturbations of the coefficients. In particular - for a version of the condition number defined by Cucker, Krick, Malajovich, and Wschebor - we establish new…

Probability · Mathematics 2018-06-11 Alperen A. Ergür , J. Maurice Rojas , Grigoris Paouris

We give the first specific conjectures on how frequently graphs satisfy sufficient conditions for being uniquely characterized by spectral information. These conjectures arise from a theoretical framework that we developed based on…

Probability · Mathematics 2026-03-31 Nikita Lvov , Alexander Van Werde

We develop a new approximative estimation method for conditional Shapley values obtained using a linear regression model. We develop a new estimation method and outperform existing methodology and implementations. Compared to the sequential…

Methodology · Statistics 2025-04-28 Fredrik Lohne Aanes

The standard procedure to decide on the complexity of a CART regression tree is to use cross-validation with the aim of obtaining a predictor that generalises well to unseen data. The randomness in the selection of folds implies that the…

Methodology · Statistics 2025-10-29 Nils Engler , Mathias Lindholm , Filip Lindskog , Taariq Nazar

This paper shows that sequential statistical analysis techniques can be generalised to the problem of selecting between alternative forecasting methods using scoring rules. A return to basic principles is necessary in order to show that…

Statistics Theory · Mathematics 2025-05-15 David T. Frazier , Donald S. Poskitt

To use control charts in practice, the in-control state usually has to be estimated. This estimation has a detrimental effect on the performance of control charts, which is often measured for example by the false alarm probability or the…

Methodology · Statistics 2013-07-30 Axel Gandy , Jan Terje Kvaløy

The regsem package in R, an implementation of regularized structural equation modeling (RegSEM; Jacobucci, Grimm, and McArdle 2016), was recently developed with the goal of incorporating various forms of penalized likelihood estimation in a…

Methodology · Statistics 2017-09-11 Ross Jacobucci

Max-stable random fields play a central role in modeling extreme value phenomena. We obtain an explicit formula for the conditional probability in general max-linear models, which include a large class of max-stable random fields. As a…

Computation · Statistics 2010-11-29 Yizao Wang , Stilian A. Stoev

Spectral estimators are fundamental in lowrank matrix models and arise throughout machine learning and statistics, with applications including network analysis, matrix completion and PCA. These estimators aim to recover the leading…

Statistics Theory · Mathematics 2025-02-17 Hao Yan , Keith Levin

In this article a special class of nonlinear optimal control problems involving a bilinear term in the boundary condition is studied. These kind of problems arise for instance in the identification of an unknown space-dependent Robin…

Numerical Analysis · Mathematics 2024-12-20 Max Winkler

The thresholding covariance estimator has nice asymptotic properties for estimating sparse large covariance matrices, but it often has negative eigenvalues when used in real data analysis. To simultaneously achieve sparsity and positive…

Methodology · Statistics 2012-08-29 Lingzhou Xue , Shiqian Ma , Hui Zou

Randomized experiments are the gold standard for estimating the average treatment effect (ATE). While covariate adjustment can reduce the asymptotic variances of the unbiased Horvitz-Thompson estimators for the ATE, it suffers from…

Methodology · Statistics 2025-08-22 Xin Lu , Lei Shi , Hanzhong Liu , Peng Ding

Calibration of expensive simulation models involves an emulator based on simulation outputs generated across various parameter settings to replace the actual model. Noisy outputs of stochastic simulation models require many simulation…

Methodology · Statistics 2025-05-08 Özge Sürer

There are many advantages to use probability method for nonlinear system identification, such as the noises and outliers in the data set do not affect the probability models significantly; the input features can be extracted in probability…

Systems and Control · Computer Science 2018-06-08 Erick de la Rosa , Wen Yu

Spectral clustering is sensitive to how graphs are constructed from data particularly when proximal and imbalanced clusters are present. We show that Ratio-Cut (RCut) or normalized cut (NCut) objectives are not tailored to imbalanced data…

Machine Learning · Statistics 2013-09-11 Jing Qian , Venkatesh Saligrama

The estimation of covariance matrices of multiple classes with limited training data is a difficult problem. The sample covariance matrix (SCM) is known to perform poorly when the number of variables is large compared to the available…

Methodology · Statistics 2021-11-10 Elias Raninen , Esa Ollila