中文
相关论文

相关论文: Learning from MOM's principles: Le Cam's approach

200 篇论文

Data used in deep learning is notoriously problematic. For example, data are usually combined from diverse sources, rarely cleaned and vetted thoroughly, and sometimes corrupted on purpose. Intentional corruption that targets the weak spots…

机器学习 · 统计学 2021-11-09 Shih-Ting Huang , Johannes Lederer

We study the estimation capacity of the generalized Lasso, i.e., least squares minimization combined with a (convex) structural constraint. While Lasso-type estimators were originally designed for noisy linear regression problems, it has…

统计理论 · 数学 2019-09-12 Martin Genzel , Gitta Kutyniok

Anomalies and outliers are common in real-world data, and they can arise from many sources, such as sensor faults. Accordingly, anomaly detection is important both for analyzing the anomalies themselves and for cleaning the data for further…

机器学习 · 统计学 2018-11-13 Haitao Liu , Randy C. Paffenroth , Jian Zou , Chong Zhou

We obtain the minimax rate for a mean location model with a bounded star-shaped set $K \subseteq \mathbb{R}^n$ constraint on the mean, in an adversarially corrupted data setting with Gaussian noise. We assume an unknown fraction $\epsilon…

统计理论 · 数学 2026-03-06 Akshay Prasadan , Matey Neykov

For obtaining optimal first-order convergence guarantee for stochastic optimization, it is necessary to use a recurrent data sampling algorithm that samples every data point with sufficient frequency. Most commonly used data sampling…

最优化与控制 · 数学 2024-07-23 William G. Powell , Hanbaek Lyu

Le Cam's method (or the two-point method) is a commonly used tool for obtaining statistical lower bound and especially popular for functional estimation problems. This work aims to explain and give conditions for the tightness of Le Cam's…

统计理论 · 数学 2021-01-05 Yury Polyanskiy , Yihong Wu

This paper addresses the robust estimation of linear regression models in the presence of potentially endogenous outliers. Through Monte Carlo simulations, we demonstrate that existing $L_1$-regularized estimation methods, including the…

计量经济学 · 经济学 2024-08-08 Zhan Gao , Hyungsik Roger Moon

We want to reconstruct a signal based on inhomogeneous data (the amount of data can vary strongly), using the model of regression with a random design. Our aim is to understand the consequences of inhomogeneity on the accuracy of estimation…

统计理论 · 数学 2016-08-16 Stéphane Gaiffas

For many inference problems in statistics and econometrics, the unknown parameter is identified by a set of moment conditions. A generic method of solving moment conditions is the Generalized Method of Moments (GMM). However, classical GMM…

机器学习 · 统计学 2021-10-18 Dhruv Rohatgi , Vasilis Syrgkanis

We consider the problem of sampling from a strongly log-concave density in $\mathbb{R}^d$, and prove an information theoretic lower bound on the number of stochastic gradient queries of the log density needed. Several popular sampling…

机器学习 · 统计学 2021-07-06 Niladri S. Chatterji , Peter L. Bartlett , Philip M. Long

A generic out-of-sample error estimate is proposed for robust $M$-estimators regularized with a convex penalty in high-dimensional linear regression where $(X,y)$ is observed and $p,n$ are of the same order. If $\psi$ is the derivative of…

统计理论 · 数学 2023-03-31 Pierre C Bellec

Suppose a given observation matrix can be decomposed as the sum of a low-rank matrix and a sparse matrix (outliers), and the goal is to recover these individual components from the observed sum. Such additive decompositions have…

机器学习 · 统计学 2010-12-07 Daniel Hsu , Sham M. Kakade , Tong Zhang

In this work, we address the problem of estimating sparse communication channels in OFDM systems in the presence of carrier frequency offset (CFO) and unknown noise variance. To this end, we consider a convex optimization problem, including…

信息论 · 计算机科学 2013-11-15 Rodrigo Carvajal , Boris I. Godoy , Juan C. Agüero

We study the problem of estimating a $p$-dimensional $s$-sparse vector in a linear model with Gaussian design and additive noise. In the case where the labels are contaminated by at most $o$ adversarial outliers, we prove that the…

统计理论 · 数学 2019-11-20 Arnak S. Dalalyan , Philip Thompson

Tournament procedures, recently introduced in Lugosi & Mendelson (2016), offer an appealing alternative, from a theoretical perspective at least, to the principle of Empirical Risk Minimization in machine learning. Statistical learning by…

机器学习 · 统计学 2022-11-02 Pierre Laforgue , Stephan Clémençon , Patrice Bertail

This work presents a technique for statistically modeling errors introduced by reduced-order models. The method employs Gaussian-process regression to construct a mapping from a small number of computationally inexpensive `error indicators'…

数值分析 · 计算机科学 2015-04-16 Martin Drohmann , Kevin Carlberg

The goal of ordinal embedding is to represent items as points in a low-dimensional Euclidean space given a set of constraints in the form of distance comparisons like "item $i$ is closer to item $j$ than item $k$". Ordinal constraints like…

机器学习 · 统计学 2016-06-24 Lalit Jain , Kevin Jamieson , Robert Nowak

For statistical modeling wherein the data regime is unfavorable in terms of dimensionality relative to the sample size, finding hidden sparsity in the ground truth can be critical in formulating an accurate statistical model. The so-called…

最优化与控制 · 数学 2025-08-04 Matteo Bergamaschi , Andrea Cristofari , Vyacheslav Kungurtsev , Francesco Rinaldi

We consider the problem of estimating a sparse linear regression vector $\beta^*$ under a gaussian noise model, for the purpose of both prediction and model selection. We assume that prior knowledge is available on the sparsity pattern,…

This paper proposes a statistically optimal approach for learning a function value using a confidence interval in a wide range of models, including general non-parametric estimation of an expected loss described as a stochastic programming…

机器学习 · 统计学 2025-08-07 Arnab Ganguly , Tobias Sutter