English
Related papers

Related papers: Depth based trimmed means

200 papers

A weighted likelihood technique for robust estimation of a multivariate Wrapped Normal distribution for data points scattered on a p-dimensional torus is proposed. The occurrence of outliers in the sample at hand can badly compromise…

Methodology · Statistics 2021-07-01 Giovanni Saraceno , Claudio Agostinelli , Luca Greco

High-dimensional statistical inference deals with models in which the the number of parameters p is comparable to or larger than the sample size n. Since it is usually impossible to obtain consistent procedures unless $p/n\rightarrow0$, a…

Statistics Theory · Mathematics 2013-03-13 Sahand N. Negahban , Pradeep Ravikumar , Martin J. Wainwright , Bin Yu

We introduce a novel projection depth for data lying in a general Hilbert space, called the regularized projection depth, with a focus on functional data. By regularizing projection directions, the proposed depth does not suffer from the…

Methodology · Statistics 2025-12-24 Filip Bočinec , Stanislav Nagy , Hyemin Yeon

Robustness in terms of outliers is an important topic and has been formally studied for a variety of problems in machine learning and computer vision. Generalized median computation is a special instance of consensus learning and a common…

Machine Learning · Computer Science 2025-03-10 Andreas Nienkötter , Sandro Vega-Pons , Xiaoyi Jiang

The halfspace depth is a well studied tool of nonparametric statistics in multivariate spaces, naturally inducing a multivariate generalisation of quantiles. The halfspace depth of a point with respect to a measure is defined as the infimum…

Methodology · Statistics 2024-09-30 Dušan Pokorný , Petra Laketa , Stanislav Nagy

In high-dimensional multivariate regression problems, enforcing low rank in the coefficient matrix offers effective dimension reduction, which greatly facilitates parameter estimation and model interpretation. However, commonly-used…

Statistics Theory · Mathematics 2017-07-18 Yiyuan She , Kun Chen

Notions of depth in regression have been introduced and studied in the literature. Regression depth (RD) of Rousseeuw and Hubert (1999), the most famous one, is a direct extension of Tukey location depth (Tukey (1975)) to regression. Like…

Statistics Theory · Mathematics 2020-07-13 Yijun Zuo

We study the task of high-dimensional entangled mean estimation in the subset-of-signals model. Specifically, given $N$ independent random points $x_1,\ldots,x_N$ in $\mathbb{R}^D$ and a parameter $\alpha \in (0, 1)$ such that each $x_i$ is…

Data Structures and Algorithms · Computer Science 2025-01-10 Ilias Diakonikolas , Daniel M. Kane , Sihan Liu , Thanasis Pittas

Unsupervised representation learning methods are widely used for gaining insight into high-dimensional, unstructured, or structured data. In some cases, users may have prior topological knowledge about the data, such as a known cluster…

Machine Learning · Computer Science 2023-11-08 Edith Heiter , Robin Vandaele , Tijl De Bie , Yvan Saeys , Jefrey Lijffijt

Low-rank tensor models are widely used in statistics. However, most existing methods rely heavily on the assumption that data follows a sub-Gaussian distribution. To address the challenges associated with heavy-tailed distributions…

Methodology · Statistics 2025-09-16 Xiaoyu Zhang , Di Wang , Guodong Li , Defeng Sun

Tukey's halfspace depth can be seen as a stochastic program and as such it is not guarded against optimizer's curse, so that a limited training sample may easily result in a poor out-of-sample performance. We propose a generalized halfspace…

Optimization and Control · Mathematics 2024-05-13 Jevgenijs Ivanovs , Pavlo Mozharovskyi

Most of the modern literature on robust mean estimation focuses on designing estimators which obtain optimal sub-Gaussian concentration bounds under minimal moment assumptions and sometimes also assuming contamination. This work looks at…

Statistics Theory · Mathematics 2024-10-30 Lucas Resende

Generalized dimensions of multifractal measures are usually seen as static objects, related to the scaling properties of suitable partition functions, or moments of measures of cells. When these measures are invariant for the flow of a…

Dynamical Systems · Mathematics 2019-10-02 Théophile Caby , Davide Faranda , Giorgio Mantica , Sandro Vaienti , Pascal Yiou

The popularity of transformer-based text embeddings calls for better statistical tools for measuring distributions of such embeddings. One such tool would be a method for ranking texts within a corpus by centrality, i.e. assigning each text…

Computation and Language · Computer Science 2023-10-24 Parker Seegmiller , Sarah Masud Preum

The trimming scheme with a prefixed cutoff portion is known as a method of improving the robustness of statistical models such as multivariate Gaussian mixture models (MG- MMs) in small scale tests by alleviating the impacts of outliers.…

Computation and Language · Computer Science 2014-05-20 Dalei Wu , Haiqing Wu

This paper is devoted to the estimators of the mean that provide strong non-asymptotic guarantees under minimal assumptions on the underlying distribution. The main ideas behind proposed techniques are based on bridging the notions of…

Statistics Theory · Mathematics 2019-05-07 Stanislav Minsker

We give the first dimensionality reduction methods for the overconstrained Tukey regression problem. The Tukey loss function $\|y\|_M = \sum_i M(y_i)$ has $M(y_i) \approx |y_i|^p$ for residual errors $y_i$ smaller than a prescribed…

Data Structures and Algorithms · Computer Science 2019-05-15 Kenneth L. Clarkson , Ruosong Wang , David P. Woodruff

Distributed data aggregation is an important task, allowing the decentralized determination of meaningful global properties, that can then be used to direct the execution of other applications. The resulting values result from the…

Distributed, Parallel, and Cluster Computing · Computer Science 2011-10-05 Paulo Jesus , Carlos Baquero , Paulo Sérgio Almeida

Confidence intervals are a popular way to visualize and analyze data distributions. Unlike p-values, they can convey information both about statistical significance as well as effect size. However, very little work exists on applying…

Applications · Statistics 2017-01-23 Jussi Korpela , Emilia Oikarinen , Kai Puolamäki , Antti Ukkonen

Nowadays, massive datasets are typically dispersed across multiple locations, encountering dual challenges of high dimensionality and huge sample size. Therefore, it is necessary to explore sufficient dimension reduction (SDR) methods for…

Methodology · Statistics 2025-09-16 Hongying Li , Minyi Zhu , Yaqi Cao , Xinyi Xu