中文
相关论文

相关论文: Empirical Risk Minimization under Random Censorshi…

200 篇论文

Empirical risk minimization (ERM) incentivizes models to exploit shortcuts, i.e., spurious correlations between input attributes and labels that are prevalent in the majority of the training data but unrelated to the task at hand. This…

机器学习 · 计算机科学 2025-07-09 Michalis Korakakis , Andreas Vlachos , Adrian Weller

We consider the empirical risk minimization problem for linear supervised learning, with regularization by structured sparsity-inducing norms. These are defined as sums of Euclidean norms on certain subsets of variables, extending the usual…

机器学习 · 统计学 2011-11-23 Rodolphe Jenatton , Jean-Yves Audibert , Francis Bach

This paper explores continuous-time and state-space optimal stopping problems from a reinforcement learning perspective. We begin by formulating the stopping problem using randomized stopping times, where the decision maker's control is…

最优化与控制 · 数学 2026-03-12 Jodi Dianetti , Giorgio Ferrari , Renyuan Xu

Tuning parameters in supervised learning problems are often estimated by cross-validation. The minimum value of the cross-validation error can be biased downward as an estimate of the test error at that same value of the tuning parameter.…

应用统计 · 统计学 2009-08-21 Ryan J. Tibshirani , Robert Tibshirani

In many scientific experiments, the data annotating cost constraints the pace for testing novel hypotheses. Yet, modern machine learning pipelines offer a promising solution, provided their predictions yield correct conclusions. We focus on…

This paper deals with parameter estimation when the data are randomly right censored. The maximum likelihood estimates from censored samples are obtained by using the expectation-maximization (EM) and Monte Carlo EM (MCEM) algorithms. We…

统计计算 · 统计学 2012-03-20 Chanseok Park , Seong Beom Lee

We derive the asymptotic risk function of regularized empirical risk minimization (ERM) estimators tuned by $n$-fold cross-validation (CV). The out-of-sample prediction loss of such estimators converges in distribution to the squared-error…

统计理论 · 数学 2026-03-24 Karun Adusumilli , Maximilian Kasy , Ashia Wilson

Often, the data used to train ranking models is subject to label noise. For example, in web-search, labels created from clickstream data are noisy due to issues such as insufficient information in item descriptions on the SERP, query…

信息检索 · 计算机科学 2022-08-18 Dany Haddad

Across health applications, researchers model outcomes as a function of time to an event, but the event time is right-censored for participants who exit the study or otherwise do not experience the event during follow-up. When censoring…

统计方法学 · 统计学 2025-11-21 Jesus E. Vazquez , Yanyuan Ma , Karen Marder , Tanya P. Garcia

We propose a deep generative approach to nonparametric estimation of conditional survival and hazard functions with right-censored data. The key idea of the proposed method is to first learn a conditional generator for the joint conditional…

统计理论 · 数学 2022-05-20 Xingyu Zhou , Wen Su , Changyu Liu , Yuling Jiao , Xingqiu Zhao , Jian Huang

This guide provides a reference for high-probability regret bounds in empirical risk minimization (ERM). The presentation is modular: we begin with intuition and general proof strategies, then state broadly applicable guarantees under…

机器学习 · 统计学 2026-03-04 Lars van der Laan

Empirical Risk Minimization (ERM) is a standard technique in machine learning, where a model is selected by minimizing a loss function over constraint set. When the training dataset consists of private information, it is natural to use a…

机器学习 · 计算机科学 2016-11-22 Kunal Talwar , Abhradeep Thakurta , Li Zhang

Distributional regression aims to find the best candidate in a given parametric family of conditional distributions to model a given dataset. As each candidate in the distribution family can be identified by the corresponding distribution…

统计理论 · 数学 2026-05-18 Gitte Kremling , Gerhard Dikta

We develop a learning principle and an efficient algorithm for batch learning from logged bandit feedback. This learning setting is ubiquitous in online systems (e.g., ad placement, web search, recommendation), where an algorithm makes a…

机器学习 · 计算机科学 2015-05-22 Adith Swaminathan , Thorsten Joachims

It has been recently shown that, under the margin (or low noise) assumption, there exist classifiers attaining fast rates of convergence of the excess Bayes risk, that is, rates faster than $n^{-1/2}$. The work on this subject has suggested…

统计理论 · 数学 2009-09-29 Jean-Yves Audibert , Alexandre B. Tsybakov

The setting of a right-censored random sample subject to contamination is considered. In various fields, expert information is often available and used to overcome the contamination. This paper integrates expert knowledge into the…

统计方法学 · 统计学 2023-03-28 Martin Bladt , Christian Furrer

Given a collection of feature maps indexed by a set $\mathcal{T}$, we study the performance of empirical risk minimization (ERM) on regression problems with square loss over the union of the linear classes induced by these feature maps.…

机器学习 · 统计学 2024-11-20 Ayoub El Hanchi , Chris J. Maddison , Murat A. Erdogdu

We prove risk bounds for binary classification in high-dimensional settings when the sample size is allowed to be smaller than the dimensionality of the training set observations. In particular, we prove upper bounds for both 'compressive…

统计理论 · 数学 2017-09-29 Ata Kaban , Robert J. Durrant

This paper extends the standard chaining technique to prove excess risk upper bounds for empirical risk minimization with random design settings even if the magnitude of the noise and the estimates is unbounded. The bound applies to many…

机器学习 · 统计学 2016-09-08 Gábor Balázs , András György , Csaba Szepesvári

Label noise in data has long been an important problem in supervised learning applications as it affects the effectiveness of many widely used classification methods. Recently, important real-world applications, such as medical diagnosis…

机器学习 · 统计学 2021-12-02 Shunan Yao , Bradley Rava , Xin Tong , Gareth James