English
Related papers

Related papers: Minimum intrinsic dimension scaling for entropic o…

200 papers

Entropic optimal transport (OT) and the Sinkhorn algorithm have made it practical for machine learning practitioners to perform the fundamental task of calculating transport distance between statistical distributions. In this work, we focus…

Optimization and Control · Mathematics 2024-03-11 Xun Tang , Holakou Rahmanian , Michael Shavlovsky , Kiran Koshy Thekumparampil , Tesi Xiao , Lexing Ying

There are several ways to measure the compressibility of a random measure; they include general approaches such as using the rate-distortion curve, as well as more specific notions, such as the Renyi information dimension (RID). The RID…

Information Theory · Computer Science 2022-03-09 Mohammad-Amin Charusaie , Arash Amini , Stefano Rini

The aim of this paper is two-fold: First, we obtain a better understanding of the intrinsic distance of diffusion processes. Precisely, (i) for all $n\ge1$, the diffusion matrix $A$ is weak upper semicontinuous on $\Omega$ if and only if…

Classical Analysis and ODEs · Mathematics 2013-05-28 Pekka Koskela , Nageswari Shanmugalingam , Yuan Zhou

When data is plentiful, the loss achieved by well-trained neural networks scales as a power-law $L \propto N^{-\alpha}$ in the number of network parameters $N$. This empirical scaling law holds for a wide variety of data modalities, and may…

Machine Learning · Computer Science 2020-04-24 Utkarsh Sharma , Jared Kaplan

The seminal result of Johnson and Lindenstrauss on random embeddings has been intensively studied in applied and theoretical computer science. Despite that vast body of literature, we still lack of complete understanding of statistical…

Machine Learning · Computer Science 2021-04-13 Maciej Skorski

This article is an exposition on some recent theoretical advances in learning latent structured models, with a primary focus on the fundamental roles that optimal transport distances play in the statistical theory. We aim at what may be the…

Statistics Theory · Mathematics 2026-01-19 XuanLong Nguyen , Yun Wei

Statistical data depth plays an important role in the analysis of multivariate data sets. The main outcome is a center-outward ordering of the observations that can be used both to highlight features of the underlying distribution of the…

Statistics Theory · Mathematics 2026-03-11 Giacomo Francisci , Claudio Agostinelli

We investigate MIMO eigenmode transmission using statistical channel state information at the transmitter. We consider a general jointly-correlated MIMO channel model, which does not require separable spatial correlations at the transmitter…

Information Theory · Computer Science 2016-11-17 Xiqi Gao , Bin Jiang , Xiao Li , Alex B. Gershman , Matthew R. McKay

A new data-enabled control technique for uncertain linear time-invariant systems, recently conceived by Coulson et\ al., builds upon the direct optimization of controllers over input/output pairs drawn from a large dataset. We adopt an…

Systems and Control · Electrical Eng. & Systems 2020-09-29 Filippo Fabiani , Paul J. Goulart

In this paper, we address the problem of estimating transport surplus (a.k.a. matching affinity) in high dimensional optimal transport problems. Classical optimal transport theory specifies the matching affinity and determines the optimal…

Methodology · Statistics 2017-01-02 Arnaud Dupuy , Alfred Galichon , Yifei Sun

An additive autoencoder for dimension reduction, which is composed of a serially performed bias estimation, linear trend estimation, and nonlinear residual estimation, is proposed and analyzed. Computational experiments confirm that an…

Machine Learning · Computer Science 2022-10-14 Tommi Kärkkäinen , Jan Hänninen

The intrinsic dimension (ID) is a powerful tool to detect and quantify correlations from data. Recently, it has been successfully applied to study statistical and many-body systems in equilibrium, yet its application to systems away from…

Statistical Mechanics · Physics 2025-10-21 Roberto Verdel , Devendra Singh Bhakuni , Santiago Acevedo

The Manifold Hypothesis is a widely accepted tenet of Machine Learning which asserts that nominally high-dimensional data are in fact concentrated near a low-dimensional manifold, embedded in high-dimensional space. This phenomenon is…

Methodology · Statistics 2025-03-24 Nick Whiteley , Annie Gray , Patrick Rubin-Delanchy

In this paper, we target the problem of sufficient dimension reduction with symmetric positive definite matrices valued responses. We propose the intrinsic minimum average variance estimation method and the intrinsic outer product gradient…

Methodology · Statistics 2023-02-28 B. Chen , S. Dai , Z. Yu

Existing generalization bounds fail to explain crucial factors that drive the generalization of modern neural networks. Since such bounds often hold uniformly over all parameters, they suffer from over-parametrization and fail to account…

Machine Learning · Statistics 2023-11-14 Songyan Hou , Parnian Kassraie , Anastasis Kratsios , Andreas Krause , Jonas Rothfuss

We introduce a new method for obtaining quantitative convergence rates for the central limit theorem (CLT) in a high dimensional setting. Using our method, we obtain several new bounds for convergence in transportation distance and entropy,…

Probability · Mathematics 2020-09-08 Ronen Eldan , Dan Mikulincer , Alex Zhai

Recent work has found that neural networks with stronger generalization tend to exhibit higher representational alignment with one another across architectures and training paradigms. In this work, we show that models with stronger…

Machine Learning · Computer Science 2026-02-02 Junjie Yu , Wenxiao Ma , Chen Wei , Jianyu Zhang , Haotian Deng , Zihan Deng , Quanying Liu

The global structure of the minimal spanning tree (MST) is expected to be universal for a large class of underlying random discrete structures. However, very little is known about the intrinsic geometry of MSTs of most standard models, and…

Probability · Mathematics 2021-06-01 Louigi Addario-Berry , Sanchayan Sen

We propose a fundamental metric for measuring the distance between two distributions. This metric, referred to as the decision-focused (DF) divergence, is tailored to stochastic linear optimization problems in which the objective…

Statistics Theory · Mathematics 2026-02-04 Suhan Liu , Mo Liu

Minimum divergence problems under integral constraints appear throughout statistics and probability, including sequential inference, bandit theory, and distributionally robust optimization. In many such settings, dual representations are…

Information Theory · Computer Science 2026-03-24 Shubhanshu Shekhar , Shubhada Agrawal