English
Related papers

Related papers: Consistent estimation of the missing mass for feat…

200 papers

The empirical distribution function assigns mass $1/n$ to each of the $n$ observations in a sample. As these are highly variable, estimation error may be reduced by replacing them with estimated observations that are asymptotically less…

Methodology · Statistics 2026-05-26 Tommaso Lando , Lorenzo Tedesco

We present new results for consistency of maximum likelihood estimators with a focus on multivariate mixed models. Our theory builds on the idea of using subsets of the full data to establish consistency of estimators based on the full…

Statistics Theory · Mathematics 2019-02-13 Karl Oskar Ekvall , Galin L. Jones

Existing approaches to model uncertainty typically either compare models using a quantitative model selection criterion or evaluate posterior model probabilities having set a prior. In this paper, we propose an alternative strategy which…

Methodology · Statistics 2025-03-26 Vik Shirvaikar , Stephen G. Walker , Chris Holmes

Predictive mean matching (PMM) is a popular imputation strategy that imputes missing values by borrowing observed values from other cases with similar expectations. We show that, unlike other imputation strategies, PMM is not guaranteed to…

Methodology · Statistics 2025-07-01 Paul T. von Hippel

Statisticians often face the choice between using probability models or a paradigm defined by minimising a loss function. Both approaches are useful and, if the loss can be re-cast into a proper probability model, there are many tools to…

Methodology · Statistics 2022-03-29 Jack Jewson , David Rossell

Motivated by recently emerging problems in machine learning and statistics, we propose data models which relax the familiar i.i.d. assumption. In essence, we seek to understand what it means for data to come from a set of probability…

Statistics Theory · Mathematics 2025-01-08 Christian Fröhlich , Robert C. Williamson

This article reviews recent advances in fixed effect estimation of panel data models for long panels, where the number of time periods is relatively large. We focus on semiparametric models with unobserved individual and time effects, where…

Econometrics · Economics 2018-03-29 Iván Fernández-Val , Martin Weidner

Predicting scalar outcomes using functional predictors is a classic problem in functional data analysis. In many applications, however, only specific locations or time-points of the functional predictors have an impact on the outcome. Such…

Statistics Theory · Mathematics 2020-07-14 Dominik Poß , Dominik Liebl , Alois Kneip , Hedwig Eisenbarth , Tor D. Wager , Lisa Feldman Barrett

We consider a multivariate density model where we estimate the excess mass of the unknown probability density $f$ at a given level $\nu>0$ from $n$ i.i.d. observed random variables. This problem has several applications such as…

Statistics Theory · Mathematics 2009-09-29 Cristina Butucea , Mathilde Mougeot , Karine Tribouley

Using a time series model to mimic an observed time series has a long history. However, with regard to this objective, conventional estimation methods for discrete-time dynamical models are frequently found to be wanting. In fact, they are…

Statistics Theory · Mathematics 2015-03-19 Yingcun Xia , Howell Tong

To address model uncertainty under flexible loss functions in prediction problems, we propose a model averaging method that accommodates various loss functions, including asymmetric linear and quadratic loss functions, as well as many other…

Methodology · Statistics 2025-01-23 Dieqi Gu , Qingfeng Liu , Xinyu Zhang

Missing data is an universal problem in statistics. We develop a unified framework for estimating parameters defined by general estimating equations under a missing-at-random (MAR) mechanism, based on generalized entropy calibration…

Methodology · Statistics 2026-03-31 Mst Moushumi Pervin , Hengfang Wang , Jae Kwang Kim

This paper shows that one cannot learn the probability of rare events without imposing further structural assumptions. The event of interest is that of obtaining an outcome outside the coverage of an i.i.d. sample from a discrete…

Machine Learning · Statistics 2015-03-13 Elchanan Mossel , Mesrob I. Ohannessian

An infinite urn scheme is defined by a probability mass function $(p_j)_{j\geq1}$ over positive integers. A random allocation consists of a sample of $N$ independent drawings according to this probability distribution where $N$ may be…

Statistics Theory · Mathematics 2016-09-29 Anna Ben-Hamou , Stéphane Boucheron , Mesrob I. Ohannessian

We address a parametric joint detection-estimation problem for discrete signals of the form $x(t) = \sum_{n=1}^{N} \alpha_n e^{-i \lambda_n t } + \epsilon_t$, $t \in \mathbb{N}$, with an additive noise represented by independent centered…

Classical Analysis and ODEs · Mathematics 2018-08-14 Illya M. Karabash , Jürgen Prestin

As technology advanced, collecting data via automatic collection devices become popular, thus we commonly face data sets with lengthy variables, especially when these data sets are collected without specific research goals beforehand. It…

Machine Learning · Statistics 2022-05-10 Wan-Ping Nicole Chen , Yuan-chin Ivan Chang

A probabilistic query may not be estimable from observed data corrupted by missing values if the data are not missing at random (MAR). It is therefore of theoretical interest and practical importance to determine in principle whether a…

Machine Learning · Statistics 2016-11-16 Jin Tian

The brilliant method due to Good and Turing allows for estimating objects not occurring in a sample. The problem, known under names "sample coverage" or "missing mass" goes back to their cryptographic work during WWII, but over years has…

Machine Learning · Statistics 2021-04-16 Maciej Skorski

The Hidden Markov Model (HMM) is a widely-used statistical model for handling sequential data. However, the presence of missing observations in real-world datasets often complicates the application of the model. The EM algorithm and Gibbs…

Machine Learning · Statistics 2026-01-06 Dongrong Li , Tianwei Yu , Xiaodan Fan

In this paper, a linear model with grouped explanatory variables is considered. The idea is to perform an automatic detection of different successive groups of the unknown coefficients under the assumption that the number of groups is of…

Methodology · Statistics 2019-05-22 Gabriela Ciuperca , Matus Maciak , Francois Wahl