English
Related papers

Related papers: Uniform Convergence Beyond Glivenko-Cantelli

200 papers

Mixture proportion estimation (MPE) aims to estimate class priors from unlabeled data. This task is a critical component in weakly supervised learning, such as PU learning, learning with label noise, and domain adaptation. Existing MPE…

Machine Learning · Computer Science 2026-04-09 Yushi Hirose , Akito Narahara , Takafumi Kanamori

Measurement models (MMs) stand at the highest structural level of quantum measurement theory. MMs can be employed to construct instruments which stand at the next level. An instrument is thought of as an apparatus that is used to measure…

Quantum Physics · Physics 2023-08-22 Stan Gudder

Random integers, sampled uniformly from $[1,x]$, share similarities with random permutations, sampled uniformly from $S_n$. These similarities include the Erd\H{o}s--Kac theorem on the distribution of the number of prime factors of a random…

Number Theory · Mathematics 2024-10-04 Dor Elboim , Ofir Gorodetsky

Energy-based models (EBMs) are generative models inspired by statistical physics with a wide range of applications in unsupervised learning. Their performance is best measured by the cross-entropy (CE) of the model distribution relative to…

Machine Learning · Computer Science 2023-12-14 Davide Carbone , Mengjian Hua , Simon Coste , Eric Vanden-Eijnden

We show how information on the uniformity properties of a point set employed in numerical multidimensional integration can be used to improve the error estimate over the usual Monte Carlo one. We introduce a new measure of (non-)uniformity…

High Energy Physics - Phenomenology · Physics 2009-10-28 Jiri Hoogland , Ronald Kleiss

Researchers have developed ways to generalize the mean and variance to situations in which a data metric is available. We apply the tools developed in Pennec (2006) to categorical data, and show the generality of this approach by…

Applications · Statistics 2014-10-07 Roger Bilisoly

Contextualized embeddings vary by context, even for the same token, and form a distribution in the embedding space. To analyze this distribution, we focus on the norm of the mean embedding and the variance of the embeddings. In this study,…

Computation and Language · Computer Science 2024-12-18 Hiroaki Yamagiwa , Hidetoshi Shimodaira

It is well known that the independence of the sample mean and the sample variance characterizes the normal distribution. By using Anosov's theorem, we further investigate the analogous characteristic properties in terms of the sample mean…

Statistics Theory · Mathematics 2021-12-14 Chin-Yuan Hu , Gwo Dong Lin

We employ a parameter-free distribution estimation framework where estimators are random distributions and utilize the Kullback-Leibler (KL) divergence as a loss function. Wu and Vos [J. Statist. Plann. Inference 142 (2012) 1525-1536] show…

Statistics Theory · Mathematics 2015-09-21 Paul Vos , Qiang Wu

In this work, we revisit the problem of uniformity testing of discrete probability distributions. A fundamental problem in distribution testing, testing uniformity over a known domain has been addressed over a significant line of works, and…

Data Structures and Algorithms · Computer Science 2017-08-17 Tuğkan Batu , Clément L. Canonne

Underdamped Langevin Monte Carlo (ULMC) is an algorithm used to sample from unnormalized densities by leveraging the momentum of a particle moving in a potential well. We provide a novel analysis of ULMC, motivated by two central questions:…

Statistics Theory · Mathematics 2023-02-17 Matthew Zhang , Sinho Chewi , Mufan Bill Li , Krishnakumar Balasubramanian , Murat A. Erdogdu

Generalized estimating equations (GEE) are of great importance in analyzing clustered data without full specification of multivariate distributions. A recent approach jointly models the mean, variance, and correlation coefficients of…

Methodology · Statistics 2025-01-13 Zhenyu Xu , Jason P. Fine , Wenling Song , Jun Yan

We present an operator-free, measure-theoretic approach to the conditional mean embedding (CME) as a random variable taking values in a reproducing kernel Hilbert space. While the kernel mean embedding of unconditional distributions has…

Machine Learning · Computer Science 2021-01-11 Junhyung Park , Krikamol Muandet

Uncertainty quantification by ensemble learning is explored in terms of an application from computational optical form measurements. The application requires to solve a large-scale, nonlinear inverse problem. Ensemble learning is used to…

Machine Learning · Computer Science 2021-03-03 Lara Hoffmann , Ines Fortmeier , Clemens Elster

A hidden Markov model is called observable if distinct initial laws give rise to distinct laws of the observation process. Observability implies stability of the nonlinear filter when the signal process is tight, but this need not be the…

Probability · Mathematics 2009-08-10 Ramon van Handel

Distribution testing can be described as follows: $q$ samples are being drawn from some unknown distribution $P$ over a known domain $[n]$. After the sampling process, a decision must be made about whether $P$ holds some property, or is far…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-12-04 Uri Meir

We consider the estimation of the mixing distribution of a normal distribution where both the shift and scale are unobserved random variables. We argue that in general, the model is not identifiable. We give an elegant non-constructive…

Statistics Theory · Mathematics 2024-08-20 Ya'acov Ritov

The Median of Means (MoM) is a mean estimator that has gained popularity in the context of heavy-tailed data. In this work, we analyze its performance in the task of simultaneously estimating the mean of each function in a class…

Machine Learning · Statistics 2025-06-23 Mikael Møller Høgsgaard , Andrea Paudice

This paper proposes a new notion of typical sequences on a wide class of abstract alphabets (so-called standard Borel spaces), which is based on approximations of memoryless sources by empirical distributions uniformly over a class of…

Information Theory · Computer Science 2016-11-17 Maxim Raginsky

The problem of characterizing a multivariate distribution of a random vector using examination of univariate combinations of vector components is an essential issue of multivariate analysis. The likelihood principle plays a prominent role…

Methodology · Statistics 2019-10-29 Albert Vexler