English
Related papers

Related papers: Improved performance guarantees for Tukey's median

200 papers

Aleatoric uncertainty captures the inherent randomness of the data, such as measurement noise. In Bayesian regression, we often use a Gaussian observation model, where we control the level of aleatoric uncertainty with a noise variance…

Machine Learning · Computer Science 2022-03-31 Sanyam Kapoor , Wesley J. Maddox , Pavel Izmailov , Andrew Gordon Wilson

The design of a metric between probability distributions is a longstanding problem motivated by numerous applications in Machine Learning. Focusing on continuous probability distributions on the Euclidean space $\mathbb{R}^d$, we introduce…

We study discrete dynamics governed by a difference inclusion whose increment is the sum of a selection from a set-valued map and a noise term. For any bounded realization, convergence follows once the inter-iterate diameter is controlled…

Optimization and Control · Mathematics 2026-05-15 Lexiao Lai , Mingzhi Song

We propose a robust and scalable procedure for general optimization and inference problems on manifolds leveraging the classical idea of `median-of-means' estimation. This is motivated by ubiquitous examples and applications in modern data…

Methodology · Statistics 2020-06-16 Lizhen Lin , Drew Lazar , Bayan Sarpabayeva , David B. Dunson

The computational complexity of some depths that satisfy the projection property, such as the halfspace depth or the projection depth, is known to be high, especially for data of higher dimensionality. In such scenarios, the exact depth is…

Statistics Theory · Mathematics 2021-05-28 Stanislav Nagy , Rainer Dyckerhoff , Pavlo Mozharovskyi

Centrality descriptors are widely used to rank nodes according to specific concept(s) of importance. Despite the large number of centrality measures available nowadays, it is still poorly understood how to identify the node which can be…

Methodology · Statistics 2020-01-13 Giulia Bertagnolli , Claudio Agostinelli , Manlio De Domenico

The estimation of information measures of continuous distributions based on samples is a fundamental problem in statistics and machine learning. In this paper, we analyze estimates of differential entropy in $K$-dimensional Euclidean space,…

Information Theory · Computer Science 2021-11-29 Georg Pichler , Pablo Piantanida , Günther Koliander

Estimated density is often interpreted as indicating how typical a sample is under a model. Yet deep models trained on one dataset can assign higher density to simpler out-of-distribution (OOD) data than to in-distribution test data. We…

Machine Learning · Computer Science 2026-04-03 Weyl Lu , Chenjie Hao , Yubei Chen

We characterize the fundamental limits of high-dimensional mean testing under arbitrary truncation, where samples are drawn from the conditional distribution $P(\cdot \mid S)$ for an unknown truncation set $S$ that may hide up to an…

Machine Learning · Statistics 2026-05-05 Yuhao Wang , Roberto Imbuzeiro Oliveira , Themis Gouleakis

In many problems from multivariate analysis, the parameter of interest is a shape matrix, that is, a normalized version of the corresponding scatter or dispersion matrix. In this paper, we propose a depth concept for shape matrices that…

Statistics Theory · Mathematics 2018-12-03 Davy Paindaveine , Germain Van Bever

$\renewcommand{\Re}{\mathbb{R}}$ We develop a general randomized technique for solving "implic it" linear programming problems, where the collection of constraints are defined implicitly by an underlying ground set of elements. In many…

Computational Geometry · Computer Science 2021-12-24 Timothy M. Chan , Sariel Har-Peled , Mitchell Jones

Data depth provides a centre-outward ordering for multivariate data. Recently, some univariate GoF tests based on data depth have been studied by Li (2018). This paper discusses some univariate goodness of fit tests based on centre-outward…

Methodology · Statistics 2024-05-14 Rahul Singh

Normalization and outlier detection belong to the preprocessing of gene expression data. We propose a natural normalization procedure based on statistical data depth which normalizes to the distribution of gene expressions of the most…

Methodology · Statistics 2022-06-29 Alicia Nieto-Reyes , Javier Cabrera

Kernel techniques are among the most popular and flexible approaches in data science allowing to represent probability measures without loss of information under mild conditions. The resulting mapping called mean embedding gives rise to a…

Machine Learning · Statistics 2024-11-27 Linda Chamakh , Zoltan Szabo

It is a standard assumption that datasets in high dimension have an internal structure which means that they in fact lie on, or near, subsets of a lower dimension. In many instances it is important to understand the real dimension of the…

Machine Learning · Statistics 2025-07-21 James A. D. Binnie , Paweł Dłotko , John Harvey , Jakub Malinowski , Ka Man Yim

We confirm a conjecture by Everett, Sinclair, and Dankelmann~[Some Centrality results new and old, J. Math. Sociology 28 (2004), 215--227] regarding the problem of maximizing closeness centralization in two-mode data, where the number of…

Combinatorics · Mathematics 2016-08-16 Matjaž Krnc , Jean-Sébastien Sereni , Riste Škrekovski , Zelealem B. Yilma

In this paper, we introduce a class of improved estimators for the mean parameter matrix of a multivariate normal distribution with an unknown variance-covariance matrix. In particular, the main results of [D.Ch\'etelat and M. T.…

Statistics Theory · Mathematics 2024-06-25 Arash A. Foroushani , Severien Nkurunziza

The box-and-whisker plot, introduced by Tukey (1977), is one of the most popular graphical methods in descriptive statistics. On the other hand, however, Tukey's boxplot is free of sample size, yielding the so-called "one-size-fits-all"…

Methodology · Statistics 2025-06-10 Hongmei Lin , Riquan Zhang , Tiejun Tong

The problem of nonlinear functional of parameters, such as differential entropy, has received much attention in information theory and statistics. In many situations, prior information about the parameters is available in the form of order…

Statistics Theory · Mathematics 2026-03-10 Somnath Mandal , Lakshmi Kanta Patra

The notion of data depth has long been in use to obtain robust location and scale estimates in a multivariate setting. The depth of an observation is a measure of its centrality, with respect to a data set or a distribution. The data depths…

Methodology · Statistics 2009-09-29 Sara López-Pintado , Rebecka Jornsten
‹ Prev 1 3 4 5 6 7 10 Next ›