中文
相关论文

相关论文: Rademacher upper bounds for cross-validation error…

200 篇论文

We investigate optimal posteriors for recently introduced \cite{begin2016pac} chi-squared divergence based PAC-Bayesian bounds in terms of nature of their distribution, scalability of computations, and test set performance. For a finite…

统计理论 · 数学 2020-08-18 Puja Sahu , Nandyala Hemachandra

Approximate Leave-One-Out Cross-Validation (ALO-CV) is a method that has been proposed to estimate the generalization error of a regularized estimator in the high-dimensional regime where dimension and sample size are of the same order, the…

统计理论 · 数学 2026-02-13 Pierre C Bellec

Recently, contrastive learning has found impressive success in advancing the state of the art in solving various machine learning tasks. However, the existing generalization analysis is very limited or even not meaningful. In particular,…

机器学习 · 计算机科学 2023-03-01 Yunwen Lei , Tianbao Yang , Yiming Ying , Ding-Xuan Zhou

Mixed level orthogonal arrays are basic structures in experimental design. We develop three algorithms that compute Rao and Gilbert-Varshamov type bounds for mixed level orthogonal arrays. The computational complexity of the terms involved…

统计理论 · 数学 2009-05-03 Ferruh Ozbudak , Ali Devin Sezer

In this work, we tackle the problem of minimising the Conditional-Value-at-Risk (CVaR) of output quantities of complex differential models with random input data, using gradient-based approaches in combination with the Multi-Level Monte…

数值分析 · 数学 2023-10-16 Sundar Ganesh , Fabio Nobile

Bayesian coresets speed up posterior inference in the large-scale data regime by approximating the full-data log-likelihood function with a surrogate log-likelihood based on a small, weighted subset of the data. But while Bayesian coresets…

机器学习 · 统计学 2024-10-18 Trevor Campbell

The constrained Cramer-Rao bound (CCRB) is a lower bound on the mean-squared-error (MSE) of estimators that satisfy some unbiasedness conditions. Although the CCRB unbiasedness conditions are satisfied asymptotically by the constrained…

信息论 · 计算机科学 2019-02-20 Eyal Nitzan , Tirza Routtenberg , Joseph Tabrikian

Model selection is difficult to analyse yet theoretically and empirically important, especially for high-dimensional data analysis. Recently the least absolute shrinkage and selection operator (Lasso) has been applied in the statistical and…

机器学习 · 统计学 2016-06-02 Ning Xu , Jian Hong , Timothy C. G. Fisher

We study differentially private (DP) stochastic optimization (SO) with loss functions whose worst-case Lipschitz parameter over all data may be extremely large or infinite. To date, the vast majority of work on DP SO assumes that the loss…

机器学习 · 计算机科学 2024-10-01 Andrew Lowy , Meisam Razaviyayn

Real-world data is laden with outlying values. The challenge for machine learning is that the learner typically has no prior knowledge of whether the feedback it receives (losses, gradients, etc.) will be heavy-tailed or not. In this work,…

机器学习 · 统计学 2020-12-16 Matthew J. Holland

PAC-Bayesian set up involves a stochastic classifier characterized by a posterior distribution on a classifier set, offers a high probability bound on its averaged true risk and is robust to the training sample used. For a given posterior,…

机器学习 · 计算机科学 2019-12-17 Puja Sahu , Nandyala Hemachandra

Motivated by the learned iterative soft thresholding algorithm (LISTA), we introduce a general class of neural networks suitable for sparse reconstruction from few linear measurements. By allowing a wide range of degrees of weight-sharing…

机器学习 · 计算机科学 2022-01-19 Ekkehard Schnoor , Arash Behboodi , Holger Rauhut

We study the approximation of a (finite) continuous-time Markov chain by a Markov chain on a reduced state space, and we provide formal error bounds for the approximated transient distributions in the Wasserstein distance. These bounds…

概率论 · 数学 2025-12-19 Fabian Michel

State-of-the-art automated machine learning systems for tabular data often employ cross-validation; ensuring that measured performances generalize to unseen data, or that subsequent ensembling does not overfit. However, using k-fold…

机器学习 · 计算机科学 2024-08-05 Edward Bergman , Lennart Purucker , Frank Hutter

Cross-Validation (CV), and out-of-sample performance-estimation protocols in general, are often employed both for (a) selecting the optimal combination of algorithms and values of hyper-parameters (called a configuration) for producing the…

机器学习 · 计算机科学 2017-08-28 Ioannis Tsamardinos , Elissavet Greasidou , Michalis Tsagris , Giorgos Borboudakis

We study the minimax sample complexity of multicalibration in the batch setting. A learner observes $n$ i.i.d. samples from an unknown distribution and must output a (possibly randomized) predictor whose population multicalibration error,…

机器学习 · 计算机科学 2026-04-24 Natalie Collina , Jiuyao Lu , Georgy Noarov , Aaron Roth

Manifold regularization is a commonly used technique in semi-supervised learning. It enforces the classification rule to be smooth with respect to the data-manifold. Here, we derive sample complexity bounds based on pseudo-dimension for…

机器学习 · 计算机科学 2020-07-31 Alexander Mey , Tom Viering , Marco Loog

We propose a simple method for evaluating the model that has been chosen by an adaptive regression procedure, our main focus being the lasso. This procedure deletes each chosen predictor and refits the lasso to get a set of models that are…

统计方法学 · 统计学 2018-06-06 Leying Guan , Robert Tibshirani

We study the efficiency of V-fold cross-validation (VFCV) for model selection from the non-asymptotic viewpoint, and suggest an improvement on it, which we call ``V-fold penalization''. Considering a particular (though simple) regression…

统计理论 · 数学 2008-02-07 Sylvain Arlot

Importance sampling (IS) is commonly used for cross validation (CV) in Bayesian models, because it only involves reweighting existing posterior draws without needing to re-estimate the model by re-running Markov chain Monte Carlo (MCMC).…

统计计算 · 统计学 2025-08-12 Geonhee Han , Andrew Gelman