English
Related papers

Related papers: Large distortion dimension reduction using random …

200 papers

Applications in machine learning and data mining require computing pairwise Lp distances in a data matrix A. For massive high-dimensional data, computing all pairwise distances of A can be infeasible. In fact, even storing A or all pairwise…

Machine Learning · Computer Science 2008-12-18 Ping Li

We propose two tests for the equality of covariance matrices between two high-dimensional populations. One test is on the whole variance--covariance matrices, and the other is on off-diagonal sub-matrices, which define the covariance…

Statistics Theory · Mathematics 2012-06-06 Jun Li , Song Xi Chen

Perturbing a deterministic $n$-dimensional matrix with small Gaussian noise is a cornerstone of smoothed analysis of algorithms [Spielman and Teng, JACM 2004], as it reduces the condition number of the input to $O(n)$, and with it the…

Data Structures and Algorithms · Computer Science 2026-04-28 Shabarish Chenakkod , Michał Dereziński , Xiaoyu Dong , Mark Rudelson

A well-known conjecture in analytic number theory states that for every pair of sets $X,Y\subset\mathbb{Z}/p\mathbb{Z}$, each of size at least $\log ^C p$ (for some constant $C$) we have that the number of pairs $(x,y)\in X\times Y$ such…

Combinatorics · Mathematics 2015-12-18 Rudi Mrazović

We use available measurements to estimate the unknown parameters (variance, smoothness parameter, and covariance length) of a covariance function by maximizing the joint Gaussian log-likelihood function. To overcome cubic complexity in the…

Computation · Statistics 2018-09-13 Alexander Litvinenko , Ying Sun , Marc G. Genton , David Keyes

We prove bounds for the almost sure value of the Hausdorff dimension of the limsup set of a sequence of balls in $\mathbf{R}^d$ whose centres are independent, identically distributed random variables. The formulas obtained involve the rate…

Classical Analysis and ODEs · Mathematics 2018-08-01 Fredrik Ekström , Tomas Persson

Let $\mathbf{A}\in \mathbb{R}^{n\times n}$ be a matrix with diagonal $\text{diag}(\mathbf{A})$ and let $\bar{\mathbf{A}}$ be $\mathbf{A}$ with its diagonal set to all zeros. We show that Hutchinson's estimator run for $m$ iterations returns…

Data Structures and Algorithms · Computer Science 2022-11-08 Prathamesh Dharangutte , Christopher Musco

Denote by $\mathcal{H}_k (n,p)$ the random $k$-graph in which each $k$-subset of $\{1... n\}$ is present with probability $p$, independent of other choices. More or less answering a question of Balogh, Bohman and Mubayi, we show: there is a…

Combinatorics · Mathematics 2016-08-17 Arran Hamm , Jeff Kahn

We consider high-dimensional inference for potentially misspecified Cox proportional hazard models based on low dimensional results by Lin and Wei [1989]. A de-sparsified Lasso estimator is proposed based on the log partial likelihood…

Statistics Theory · Mathematics 2018-11-02 Shengchun Kong , Zhuqing Yu , Xianyang Zhang , Guang Cheng

In a variety of applications it is important to extract information from a probability measure $\mu$ on an infinite dimensional space. Examples include the Bayesian approach to inverse problems and possibly conditioned) continuous time…

Probability · Mathematics 2016-06-02 Frank Pinski , Gideon Simpson , Andrew Stuart , Hendrik Weber

Starting with the large deviation principle (LDP) for the Erd\H{o}s-R\'enyi binomial random graph $\mathcal{G}(n,p)$ (edge indicators are i.i.d.), due to Chatterjee and Varadhan (2011), we derive the LDP for the uniform random graph…

Probability · Mathematics 2018-05-01 Amir Dembo , Eyal Lubetzky

For arbitrary two probability measures on real d-space with given means and variances (covariance matrices), we provide lower bounds for their total variation distance. In the one-dimensional case, a tight bound is given.

Probability · Mathematics 2022-12-27 Tomohiro Nishiyama

The likelihood-informed subspace (LIS) method offers a viable route to reducing the dimensionality of high-dimensional probability distributions arising in Bayesian inference. LIS identifies an intrinsic low-dimensional linear subspace…

Computation · Statistics 2021-10-22 Tiangang Cui , Xin T. Tong

Let $A \in \mathbb{R}^{n \times n}$ be invertible, $x \in \mathbb{R}^n$ unknown and $b =Ax $ given. We are interested in approximate solutions: vectors $y \in \mathbb{R}^n$ such that $\|Ay - b\|$ is small. We prove that for all $0<…

Numerical Analysis · Mathematics 2022-07-08 Stefan Steinerberger

This paper proposes two linear projection methods for supervised dimension reduction using only the first and second-order statistics. The methods, each catering to a different parameter regime, are derived under the general Gaussian model…

Information Theory · Computer Science 2024-08-13 Biao Chen , Joshua Kortje

Let $X$ be a symmetric, isotropic random vector in $\mathbb{R}^m$ and let $X_1...,X_n$ be independent copies of $X$. We show that under mild assumptions on $\|X\|_2$ (a suitable thin-shell bound) and on the tail-decay of the marginals…

Functional Analysis · Mathematics 2022-07-13 Daniel Bartl , Shahar Mendelson

For probability distributions on $\mathbb{R}^n$, we study the optimal sample size N = N(n,p) that suffices to uniformly approximate the pth moments of all one-dimensional marginals. Under the assumption that the marginals have bounded 4p…

Probability · Mathematics 2014-05-21 Roman Vershynin

The problem of finding \emph{distance} between \emph{pattern} of length $m$ and \emph{text} of length $n$ is a typical way of generalizing pattern matching to incorporate dissimilarity score. For both Hamming and $L_1$ distances only a…

Data Structures and Algorithms · Computer Science 2018-05-04 Przemysław Uznański

Every student in statistics or data science learns early on that when the sample size largely exceeds the number of variables, fitting a logistic model produces estimates that are approximately unbiased. Every student also learns that there…

Statistics Theory · Mathematics 2022-06-08 Pragya Sur , Emmanuel J. Candes

We study general singular value shrinkage estimators in high-dimensional regression and classification, when the number of features and the sample size both grow proportionally to infinity. We allow models with general covariance matrices…

Statistics Theory · Mathematics 2020-04-01 Panagiotis Lolas