Related papers: Optimal estimation of high-order missing masses, a…
Many applications, including natural language processing, sensor networks, collaborative filtering, and federated learning, call for estimating discrete distributions from data collected in batches, some of which may be untrustworthy,…
Our primary aim is to find an estimate of the expected shortfall in various situations: (1) Nonparametric situation, when the probability distribution of the incurred loss is unknown, only satisfying some general conditions. Then, following…
Given a random sample from a distribution with density function that depends on an unknown parameter $\theta$, we are interested in accurately estimating the true parametric density function at a future observation from the same…
Let $X$ be an observable random variable with unknown distribution function $F(x) = \mathbb{P}(X \leq x), - \infty < x < \infty$, and let \[\ \theta = \sup\left \{ r \geq 0:~ \mathbb{E}|X|^{r} < \infty \right \}. \] We call $\theta$ the…
We seek an entropy estimator for discrete distributions with fully empirical accuracy bounds. As stated, this goal is infeasible without some prior assumptions on the distribution. We discover that a certain information moment assumption…
A probabilistic query may not be estimable from observed data corrupted by missing values if the data are not missing at random (MAR). It is therefore of theoretical interest and practical importance to determine in principle whether a…
The problem of estimating the covariance matrix $\Sigma$ of a $p$-variate distribution based on its $n$ observations arises in many data analysis contexts. While for $n>p$, the classical sample covariance matrix $\hat{\Sigma}_n$ is a good…
We study the problem of parameters estimation in Indirect Observability contexts, where $X_t \in R^r$ is an unobservable stationary process parametrized by a vector of unknown parameters and all observable data are generated by an…
We consider the problem of identifying, from statistics, a distribution of discrete random variables $X_1,\ldots,X_n$ that is a mixture of $k$ product distributions. The best previous sample complexity for $n \in O(k)$ was $(1/\zeta)^{O(k^2…
Finite state space hidden Markov models are flexible tools to model phenomena with complex time dependencies: any process distribution can be approximated by a hidden Markov model with enough hidden states.We consider the problem of…
In this paper we develop rate--optimal estimation procedures in the problem of estimating the $L_p$--norm, $p\in (0, \infty)$ of a probability density from independent observations. The density is assumed to be defined on $R^d$, $d\geq 1$…
As one of the most commonly seen data challenges, missing data, in particular, multiple, non-monotone missing patterns, complicates estimation and inference due to the fact that missingness mechanisms are often not missing at random, and…
In this paper, we study the hypothesis testing problem of, among $n$ random variables, determining $k$ random variables which have different probability distributions from the rest $(n-k)$ random variables. Instead of using separate…
This paper considers the problem of estimating a high-dimensional vector of parameters $\boldsymbol{\theta} \in \mathbb{R}^n$ from a noisy observation. The noise vector is i.i.d. Gaussian with known variance. For a squared-error loss…
This paper deals with the estimation of rare event probabilities using importance sampling (IS), where an optimal proposal distribution is computed with the cross-entropy (CE) method. Although, IS optimized with the CE method leads to an…
The halo mass function, encoding the comoving number density of dark matter halos of a given mass, plays a key role in understanding the formation and evolution of galaxies. As such, it is a key goal of current and future deep optical…
We consider estimating the transition probability matrix of a finite-state finite-observation alphabet hidden Markov model with known observation probabilities. The main contribution is a two-step algorithm; a method of moments estimator…
The purpose of this paper is to estimate the intensity of a Poisson process $N$ by using thresholding rules. In this paper, the intensity, defined as the derivative of the mean measure of $N$ with respect to $ndx$ where $n$ is a fixed…
The aim of this paper is to recover the regression function with sup norm loss. We construct an asymptotically sharp estimator which converges with the spatially dependent rate r\_{n, \mu}(x) = P \big(\log n / (n \mu(x)) \big)^{s / (2s +…
Mixture models are regularly used in density estimation applications, but the problem of estimating the mixing distribution remains a challenge. Nonparametric maximum likelihood produce estimates of the mixing distribution that are…