English
Related papers

Related papers: Data Amplification: Instance-Optimal Property Esti…

200 papers

The authors propose a robust semi-parametric empirical likelihood method to integrate all available information from multiple samples with a common center of measurements. Two different sets of estimating equations are used to improve the…

Methodology · Statistics 2012-10-03 Hsiao-Hsuan Wang , Yuehua Wu , Yuejiao Fu , Xiaogang Wang

Estimating information-theoretic quantities such as entropy and mutual information is central to many problems in statistics and machine learning, but challenging in high dimensions. This paper presents estimators of entropy via inference…

Machine Learning · Statistics 2022-12-13 Feras A. Saad , Marco Cusumano-Towner , Vikash K. Mansinghka

Maximum entropy estimation is of broad interest for inferring properties of systems across many different disciplines. In this work, we significantly extend a technique we previously introduced for estimating the maximum entropy of a set of…

Data Analysis, Statistics and Probability · Physics 2016-01-05 Elliot A. Martin , Jaroslav Hlinka , Alexander Meinke , Filip Děchtěrenko , Jörn Davidsen

Abundance estimation from capture-recapture data is of great importance in many disciplines. Analysis of capture-recapture data is often complicated by the existence of one-inflation and heterogeneity problems. Simultaneously taking these…

Methodology · Statistics 2025-07-15 Yang Liu , Pengfei Li , Yukun Liu , Riquan Zhang

We consider the problem of estimating the distribution function, the density and the hazard rate of the (unobservable) event time in the current status model. A well studied and natural nonparametric estimator for the distribution function…

Statistics Theory · Mathematics 2010-01-13 Piet Groeneboom , Geurt Jongbloed , Birgit I. Witte

In many situations, sample data is obtained from a noisy or imperfect source. In order to address such corruptions, this paper introduces the concept of a sampling corrector. Such algorithms use structure that the distribution is purported…

Data Structures and Algorithms · Computer Science 2018-04-03 Clément Canonne , Themis Gouleakis , Ronitt Rubinfeld

Entropy and its various generalizations are important in many fields, including mathematical statistics, communication theory, physics and computer science, for characterizing the amount of information associated with a probability…

Maximum-likelihood exponent maps have been studied as a technique to increase the understanding and improve the fit of power-law exponents to experimental and numerical simulation data, especially when they exhibit both upper and lower…

Statistical Mechanics · Physics 2012-07-02 Jordi Baró , Eduard Vives

Suppose that we wish to estimate a finite-dimensional summary of one or more function-valued features of an underlying data-generating mechanism under a nonparametric model. One approach to estimation is by plugging in flexible estimates of…

Methodology · Statistics 2020-08-28 Hongxiang Qiu , Alex Luedtke , Marco Carone

We introduce a new generalization of the Pseudo-Lindley distribution by applying alpha power transformation. The obtained distribution is referred as the Pseudo-Lindley alpha power transformed distribution (\textit{PL-APT}). Some tractable…

Statistics Theory · Mathematics 2022-01-20 Modou Ngom , Moumouni Diallo , Adja Mbarka Fall , Gane Samb Lo

Robust estimators of location and dispersion are often used in the elliptical model to obtain an uncontaminated and highly representative subsample by trimming the data outside an ellipsoid based in the associated Mahalanobis distance. Here…

Statistics Theory · Mathematics 2016-08-14 Juan A. Cuesta-Albertos , Carlos Matrán , Agustín Mayo-Iscar

We define a one-parameter family of entropies, each assigning a real number to any probability measure on a compact metric space (or, more generally, a compact Hausdorff space with a notion of similarity between points). These entropies…

Metric Geometry · Mathematics 2020-12-17 Tom Leinster , Emily Roff

Semi-continuous data comes from a distribution that is a mixture of the point mass at zero and a continuous distribution with support on the positive real line. A clear example is the daily rainfall data. In this paper, we present a novel…

Methodology · Statistics 2021-06-17 Sai K. Popuri , Nagaraj K. Neerchal , Amita Mehta , Ahmad Mousavi

In this paper, we study the functional linear multiplicative model based on the least product relative error criterion. Under some regularization conditions, we establish the consistency and asymptotic normality of the estimator. Further,…

Statistics Theory · Mathematics 2023-01-04 Qian Yan , Hanyu Li

Given data drawn from an unknown distribution, $D$, to what extent is it possible to ``amplify'' this dataset and output an even larger set of samples that appear to have been drawn from $D$? We formalize this question as follows: an…

Machine Learning · Computer Science 2024-08-27 Brian Axelrod , Shivam Garg , Vatsal Sharan , Gregory Valiant

The challenges posed by complex stochastic models used in computational ecology, biology and genetics have stimulated the development of approximate approaches to statistical inference. Here we focus on Synthetic Likelihood (SL), a…

Methodology · Statistics 2017-06-09 Matteo Fasiolo , Simon N. Wood , Florian Hartig , Mark V. Bravington

When the experimental data set is contaminated, we usually employ robust alternatives to common location and scale estimators such as the sample median and Hodges-Lehmann estimators for location and the sample median absolute deviation and…

Methodology · Statistics 2020-08-11 Chanseok Park , Haewon Kim , Min Wang

The empirical distribution function assigns mass $1/n$ to each of the $n$ observations in a sample. As these are highly variable, estimation error may be reduced by replacing them with estimated observations that are asymptotically less…

Methodology · Statistics 2026-05-26 Tommaso Lando , Lorenzo Tedesco

We study random points on the real line generated by the eigenvalues in unitary invariant random matrix ensembles or by more general repulsive particle systems. As the number of points tends to infinity, we prove convergence of the…

Probability · Mathematics 2015-11-11 Kristina Schubert , Martin Venker

Data observed at high sampling frequency are typically assumed to be an additive composite of a relatively slow-varying continuous-time component, a latent stochastic process or a smooth random function, and measurement error. Supposing…

Statistics Theory · Mathematics 2018-12-21 Jinyuan Chang , Aurore Delaigle , Peter Hall , Cheng Yong Tang