中文
相关论文

相关论文: On the Concentration of the Missing Mass

200 篇论文

Satellite conjunctions involving "near misses" of space objects are becoming increasingly likely. One approach to risk analysis for them involves the computation of the collision probability, but this has been regarded as having some…

应用统计 · 统计学 2022-03-11 Soumaya Elkantassi , Anthony Davison

We assume that the forecast error follows a probability distribution which is symmetric and monotonically non-increasing on non-negative real numbers, and if there is a mismatch between observed and predicted value, then we suffer a loss.…

统计理论 · 数学 2023-07-06 Naoya Yamaguchi , Yuka Yamaguchi , Maiya Hori

Events in the world may be caused by other, unobserved events. We consider sequences of events in continuous time. Given a probability model of complete sequences, we propose particle smoothing---a form of sequential importance…

机器学习 · 计算机科学 2019-05-15 Hongyuan Mei , Guanghui Qin , Jason Eisner

The recent paper [27] provides a statistical analysis for efficient detection of signal components when missing data samples are present. Here we focus our attention to some complex-valued discrete random variables $X_l(m,N)$ ($0\le l\le…

统计理论 · 数学 2018-03-07 Romeo Meštrović

Randomness is viewed through an analogy between a physical quantity, density of gas, and a mathematical construct -- probability density. Boltzmann's deduction of equilibrium distribution of ideal gas placed in an external potential field…

概率论 · 数学 2012-08-27 M. Grendar, , M. Grendar

We study the problem of sampling from a distribution $\mu$ with density $\propto e^{-V}$ for some potential function $V:\mathbb R^d\to \mathbb R$ with query access to $V$ and $\nabla V$. We start with the following standard assumptions: (1)…

数据结构与算法 · 计算机科学 2026-02-10 Yuchen He , Zhehan Lei , Jianan Shao , Chihao Zhang

Missing data imputation, where a model is trained on observed data to estimate unobserved values, is a fundamental problem in machine learning. In this paper, we rigorously formulate imputation model learning as a mean-squared error risk…

机器学习 · 统计学 2026-05-14 Luke Shannon , Song Liu , Katarzyna Reluga

Data represented by probability measures arise as empirical distributions, posterior distributions, and feature-based representations of complex objects. We study heterogeneity in a population of probability measures through the expected…

统计方法学 · 统计学 2026-03-17 Kisung You

Importance sampling is a popular technique in Bayesian inference: by reweighting samples drawn from a proposal distribution we are able to obtain samples and moment estimates from a Bayesian posterior over latent variables. Recent work,…

统计计算 · 统计学 2024-06-19 Sam Bowyer , Thomas Heap , Laurence Aitchison

Time series in real-world applications often have missing observations, making typical analytical methods unsuitable. One method for dealing with missing data is the concept of amplitude modulation. While this principle works with any data,…

统计方法学 · 统计学 2024-04-19 Simon Nik

We study Martin-L\"{o}f random (ML-random) points on computable probability measures on sample and parameter spaces (Bayes models). We consider variants of conditional randomness defined by ML-randomness on Bayes models and those of…

信息论 · 计算机科学 2023-04-24 Hayato Takahashi

Random probabilities are a key component to many nonparametric methods in Statistics and Machine Learning. To quantify comparisons between different laws of random probabilities several works are starting to use the elegant Wasserstein over…

统计理论 · 数学 2024-05-27 Marta Catalano , Hugo Lavenant

Let X_1,...., X_n be a collection of iid discrete random variables, and Y_1,..., Y_m a set of noisy observations of such variables. Assume each observation Y_a to be a random function of some a random subset of the X_i's, and consider the…

信息论 · 计算机科学 2007-09-04 Andrea Montanari

Consider a Gaussian memoryless multiple source with $m$ components with joint probability distribution known only to lie in a given class of distributions. A subset of $k \leq m$ components are sampled and compressed with the objective of…

信息论 · 计算机科学 2018-03-16 Vinay Praneeth Boda

We consider deconvolution from repeated observations with unknown error distribution. So far, this model has mostly been studied under the additional assumption that the errors are symmetric. We construct an estimator for the non-symmetric…

统计理论 · 数学 2014-07-15 Johanna Kappus , Fabienne Comte

A discrete-time stochastic process derived from a model of basketball is used to generalize any discrete distribution. The generalized distributions can have one or two more parameters than the parent distribution. Those derived from…

应用统计 · 统计学 2020-06-25 Rose Baker

We obtain a Bernstein type Gaussian concentration inequality for martingales. Our inequality improves the Azuma-Hoeffding inequality for moderate deviations $x$. Following the work of McDiarmid (1989), Talagrand (1996) and Boucheron, Lugosi…

概率论 · 数学 2017-10-17 Xiequan Fan

We introduce a sharpness functional for probabilistic models that quantifies sharpness as an intrinsic property of the probability distribution. The measure is derived based on a rank-based concentration principle that tracks upward…

统计方法学 · 统计学 2026-04-03 Pekka Syrjänen

Datasets are rarely a realistic approximation of the target population. Say, prevalence is misrepresented, image quality is above clinical standards, etc. This mismatch is known as sampling bias. Sampling biases are a major hindrance for…

Let $X_1,X_2,...,X_n$ be a sequence of independent or locally dependent random variables taking values in $\mathbb{Z}_+$. In this paper, we derive sharp bounds, via a new probabilistic method, for the total variation distance between the…

统计理论 · 数学 2010-10-11 Michael V. Boutsikas , Eutichia Vaggelatou