English
Related papers

Related papers: Sample-based distance-approximation for subsequenc…

200 papers

We consider the following basic inference problem: there is an unknown high-dimensional vector $w \in \mathbb{R}^n$, and an algorithm is given access to labeled pairs $(x,y)$ where $x \in \mathbb{R}^n$ is a measurement and $y = w \cdot x +…

Computational Complexity · Computer Science 2019-11-05 Xue Chen , Anindya De , Rocco A. Servedio

We initiate a systematic investigation of distribution testing in the framework of algorithmic replicability. Specifically, given independent samples from a collection of probability distributions, the goal is to characterize the sample…

Machine Learning · Computer Science 2025-07-04 Ilias Diakonikolas , Jingyi Gao , Daniel Kane , Sihan Liu , Christopher Ye

This paper studies distributed binary test of statistical independence under communication (information bits) constraints. While testing independence is very relevant in various applications, distributed independence test is particularly…

Statistics Theory · Mathematics 2021-11-29 Sebastian Espinosa , Jorge F. Silva , Pablo Piantanida

The $L^k$-Wasserstein distance $\mathbb{W}_k (k\ge 1)$ and the probability distance $\mathbb{W}_\psi$ induced by a concave function $\psi$, are estimated between different diffusion processes with singular coefficients. As applications, the…

Probability · Mathematics 2023-11-07 Xing Huang , Panpan Ren , Feng-Yu Wang

Given a large set $U$ where each item $a\in U$ has weight $w(a)$, we want to estimate the total weight $W=\sum_{a\in U} w(a)$ to within factor of $1\pm\varepsilon$ with some constant probability $>1/2$. Since $n=|U|$ is large, we want to do…

Data Structures and Algorithms · Computer Science 2021-10-29 Lorenzo Beretta , Jakub Tětek

The distance from calibration, introduced by B{\l}asiok, Gopalan, Hu, and Nakkiran (STOC 2023), has recently emerged as a central measure of miscalibration for probabilistic predictors. We study the fundamental problems of computing and…

Data Structures and Algorithms · Computer Science 2026-03-20 Mingda Qiao

We consider the problem of estimating the joint distribution of $n$ independent random variables. Our approach is based on a family of candidate probabilities that we shall call a model and which is chosen to either contain the true…

Statistics Theory · Mathematics 2021-06-01 Yannick Baraud

Given two distributions $\mathcal{P}$ and $\mathcal{Q}$ over a high-dimensional domain $\{0,1\}^n$, and a parameter $\varepsilon$, the goal of distance estimation is to determine the statistical distance between $\mathcal{P}$ and…

Data Structures and Algorithms · Computer Science 2025-09-09 Gunjan Kumar , Kuldeep S. Meel , Yash Pote

We study the problem of closeness testing for continuous distributions and its implications for causal discovery. Specifically, we analyze the sample complexity of distinguishing whether two multidimensional continuous distributions are…

Machine Learning · Computer Science 2025-03-11 Fateme Jamshidi , Sina Akbari , Negar Kiyavash

Let $T\$ be a stopping time associated with a sequence of independent random variables $Z_{1},Z_{2},...$ . By applying a suitable change in the probability measure we present relations between the moment or probability generating functions…

Statistics Theory · Mathematics 2011-06-28 M. V. Boutsikas , A. C. Rakitzis , D. L. Antzoulakos

In threshold-based anomaly detection, we want to tune the threshold of a detector to achieve an acceptable false alarm rate. However, tuning the threshold is often a non-trivial task due to unknown detector output distributions. A detector…

Systems and Control · Electrical Eng. & Systems 2022-05-02 David Umsonst , Justin Ruths , Henrik Sandberg

We study properties of a sample covariance estimate $\widehat \Sigma$ given a finite sample of $n$ i.i.d. centered random elements in $\R^d$ with the covariance matrix $\Sigma$. We derive dimension-free bounds on the squared Frobenius norm…

Probability · Mathematics 2024-09-09 Nikita Puchkin , Fedor Noskov , Vladimir Spokoiny

Sampling algorithms play an important role in controlling the quality and runtime of diffusion model inference. In recent years, a number of works~\cite{chen2023sampling,chen2023ode,benton2023error,lee2022convergence} have proposed schemes…

Machine Learning · Computer Science 2024-10-18 Shivam Gupta , Linda Cai , Sitan Chen

Approximate Bayesian Computation (ABC) is a popular computational method for likelihood-free Bayesian inference. The term "likelihood-free" refers to problems where the likelihood is intractable to compute or estimate directly, but where it…

Statistics Theory · Mathematics 2014-07-21 Stuart Barber , Jochen Voss , Mark Webster

We study the fundamental problem of sampling independent events, called subset sampling. Specifically, consider a set of $n$ events $S=\{x_1, \ldots, x_n\}$, where each event $x_i$ has an associated probability $p(x_i)$. The subset sampling…

Data Structures and Algorithms · Computer Science 2023-09-22 Lu Yi , Hanzhi Wang , Zhewei Wei

Discrepancy measures between probability distributions, often termed statistical distances, are ubiquitous in probability theory, statistics and machine learning. To combat the curse of dimensionality when estimating these distances from…

Statistics Theory · Mathematics 2021-12-21 Sloan Nietert , Ziv Goldfeld , Kengo Kato

Trace distance and infidelity (induced by square root fidelity), as basic measures of the closeness of quantum states, are commonly used in quantum state discrimination, certification, and tomography. However, the sample complexity for…

Quantum Physics · Physics 2024-10-29 Qisheng Wang , Zhicheng Zhang

Given samples from two distributions over an $n$-element set, we wish to test whether these distributions are statistically close. We present an algorithm which uses sublinear in $n$, specifically, $O(n^{2/3}\epsilon^{-8/3}\log n)$,…

Data Structures and Algorithms · Computer Science 2010-11-05 Tugkan Batu , Lance Fortnow , Ronitt Rubinfeld , Warren D. Smith , Patrick White

In multiple classification, one aims to determine whether a testing sequence is generated from the same distribution as one of the M training sequences or not. Unlike most of existing studies that focus on discrete-valued sequences with…

Machine Learning · Statistics 2024-10-30 Lina Zhu , Lin Zhou

In data-driven learning and inference tasks, the high cost of acquiring samples from the target distribution often limits performance. A common strategy to mitigate this challenge is to augment the limited target samples with data from a…

Statistics Theory · Mathematics 2025-02-06 Barron Han , Danil Akhtiamov , Reza Ghane , Babak Hassibi