English
Related papers

Related papers: Distributional Shrinkage II: Higher-Order Scores E…

200 papers

We give a principled method for decomposing the predictive uncertainty of a model into aleatoric and epistemic components with explicit semantics relating them to the real-world data distribution. While many works in the literature have…

Machine Learning · Computer Science 2024-12-30 Gustaf Ahdritz , Aravind Gollakota , Parikshit Gopalan , Charlotte Peale , Udi Wieder

This paper investigates the problem of finding an optimal nonbinary index assignment from (M) quantization levels of a maximum entropy scalar quantizer to (M)-PSK symbols transmitted over a symmetric memoryless channel with additive noise…

Information Theory · Computer Science 2021-07-01 Yunxiang Yao , Wai Ho Mow

The need to reason about uncertainty in large, complex, and multi-modal datasets has become increasingly common across modern scientific environments. The ability to transform samples from one distribution $P$ to another distribution $Q$…

Machine Learning · Statistics 2018-11-30 Diego A. Mesa , Justin Tantiongloc , Marcela Mendoza , Todd P. Coleman

In this paper, we study the problem of inference in high-order structured prediction tasks. In the context of Markov random fields, the goal of a high-order inference task is to maximize a score function on the space of labels, and the…

Machine Learning · Computer Science 2023-10-23 Chuyang Ke , Jean Honorio

We develop a computationally tractable method for estimating the optimal map between two distributions over $\mathbb{R}^d$ with rigorous finite-sample guarantees. Leveraging an entropic version of Brenier's theorem, we show that our…

Statistics Theory · Mathematics 2024-05-14 Aram-Alexandre Pooladian , Jonathan Niles-Weed

In 1991, Brenier proved a theorem that generalizes the polar decomposition for square matrices -- factored as PSD $\times$ unitary -- to any vector field $F:\mathbb{R}^d\rightarrow \mathbb{R}^d$. The theorem, known as the polar…

Machine Learning · Statistics 2025-05-28 Nina Vesseron , Marco Cuturi

We provide a new proof of the known partial regularity result for the optimal transportation map (Brenier map) between two sets. Contrary to the existing regularity theory for the Monge-Amp{\`e}re equation, which is based on the maximum…

Analysis of PDEs · Mathematics 2017-10-25 Michael Goldman , F Otto

Scoring rules are an important tool for evaluating the performance of probabilistic forecasting schemes. In the binary case, scoring rules (which are strictly proper) allow for a decomposition into terms related to the resolution and to the…

Atmospheric and Oceanic Physics · Physics 2015-05-13 Jochen Bröcker

We present a modular semantic account of Bayesian inference algorithms for probabilistic programming languages, as used in data science and machine learning. Sophisticated inference algorithms are often explained in terms of composition of…

We study polynomial optimization problems whose objective has a composition or tensor train structure. These polynomials can be evaluated as a sequence of maps, giving rise to intermediate variables (``states'') of dimension lower than the…

Optimization and Control · Mathematics 2026-04-21 Llorenç Balada Gaggioli , Didier Henrion , Milan Korda

We study the following one-way asymmetric transmission problem, also a variant of model-based compressed sensing: a resource-limited encoder has to report a small set $S$ from a universe of $N$ items to a more powerful decoder (server). The…

Data Structures and Algorithms · Computer Science 2018-07-30 Alexandr Andoni , Javad Ghaderi , Daniel Hsu , Dan Rubenstein , Omri Weinstein

Over the past few years, numerous computational models have been developed to solve Optimal Transport (OT) in a stochastic setting, where distributions are represented by samples and where the goal is to find the closest map to the ground…

Statistics Theory · Mathematics 2023-02-01 Adrien Vacher , François-Xavier Vialard

We study the problem of estimating the score function using both implicit score matching and denoising score matching. Assuming that the data distribution exhibiting a low-dimensional structure, we prove that implicit score matching is able…

Statistics Theory · Mathematics 2026-01-01 Konstantin Yakovlev , Anna Markovich , Nikita Puchkin

Score matching enables the estimation of the gradient of a data distribution, a key component in denoising diffusion models used to recover clean data from corrupted inputs. In prior work, a heuristic weighting function has been used for…

Machine Learning · Computer Science 2025-08-05 Juyan Zhang , Rhys Newbury , Xinyang Zhang , Tin Tran , Dana Kulic , Michael Burke

Consider two correlated sources $X$ and $Y$ generated from a joint distribution $p_{X,Y}$. Their G\'acs-K\"orner Common Information, a measure of common information that exploits the combinatorial structure of the distribution $p_{X,Y}$,…

Information Theory · Computer Science 2016-04-15 Salman Salamatian , Asaf Cohen , Muriel Médard

We develop a data-driven optimal shrinkage algorithm for matrix denoising in the presence of high-dimensional noise with a separable covariance structure; that is, the noise is colored and dependent across samples. The algorithm, coined…

Applications · Statistics 2024-05-14 Pei-Chun Su , Hau-Tieng Wu

We present a systematic analysis of estimation errors for a class of optimal transport based algorithms for filtering and data assimilation. Along the way, we extend previous error analyses of Brenier maps to the case of conditional Brenier…

Statistics Theory · Mathematics 2025-10-23 Mohammad Al-Jarrah , Bamdad Hosseini , Niyizhen Jin , Michele Martino , Amirhossein Taghvaei

Denoising score matching plays a pivotal role in the performance of diffusion-based generative models. However, the empirical optimal score--the exact solution to the denoising score matching--leads to memorization, where generated samples…

Machine Learning · Statistics 2025-05-07 Yu-Han Wu , Pierre Marion , Gérard Biau , Claire Boyer

We study the second-order asymptotics of information transmission using random Gaussian codebooks and nearest neighbor (NN) decoding over a power-limited stationary memoryless additive non-Gaussian noise channel. We show that the dispersion…

Information Theory · Computer Science 2016-10-20 Jonathan Scarlett , Vincent Y. F. Tan , Giuseppe Durisi

This paper demonstrates how to recover causal graphs from the score of the data distribution in non-linear additive (Gaussian) noise models. Using score matching algorithms as a building block, we show how to design a new generation of…