English
Related papers

Related papers: Minimax Predictive Density for Sparse Count Data

200 papers

Regression for count data is widely performed by models such as Poisson, negative binomial (NB) and zero-inflated regression. A challenge often faced by practitioners is the selection of the right model to take into account dispersion,…

Methodology · Statistics 2018-08-02 Hadeel S. Klakattawi , Veronica Vinciotti , Keming Yu

We study the maximum likelihood estimator of density of $n$ independent observations, under the assumption that it is well approximated by a mixture with a large number of components. The main focus is on statistical properties with respect…

Statistics Theory · Mathematics 2017-01-19 Arnak S. Dalalyan , Mehdi Sebbar

In this technical report, we consider conditional density estimation with a maximum likelihood approach. Under weak assumptions, we obtain a theoretical bound for a Kullback-Leibler type loss for a single model maximum likelihood estimate.…

Statistics Theory · Mathematics 2012-07-11 Serge Cohen , Erwan Le Pennec

This paper describes a new Bayesian interpretation of a class of skew--Student $t$ distributions. We consider a hierarchical normal model with unknown covariance matrix and show that by imposing different restrictions on the parameter…

Methodology · Statistics 2018-05-25 Abdolnasser Sadeghkhani

Count data is prevalent in various fields like ecology, medical research, and genomics. In high-dimensional settings, where the number of features exceeds the sample size, feature selection becomes essential. While frequentist methods like…

Methodology · Statistics 2024-10-22 The Tien Mai

A popular approach within the signal processing and machine learning communities consists in modelling signals as sparse linear combinations of atoms selected from a learned dictionary. While this paradigm has led to numerous empirical…

Machine Learning · Computer Science 2015-08-25 Rémi Gribonval , Rodolphe Jenatton , Francis Bach

We reconsider a nonparametric density model based on Gaussian processes. By augmenting the model with latent P\'olya--Gamma random variables and a latent marked Poisson process we obtain a new likelihood which is conjugate to the model's…

Machine Learning · Statistics 2018-05-30 Christian Donner , Manfred Opper

Tensor factorizations with nonnegative constraints have found application in analyzing data from cyber traffic, social networks, and other areas. We consider application data best described as being generated by a Poisson process (e.g.,…

Numerical Analysis · Mathematics 2018-08-23 Samantha Hansen , Todd Plantenga , Tamara G. Kolda

Estimating the Shannon entropy of a discrete distribution from which we have only observed a small sample is challenging. Estimating other information-theoretic metrics, such as the Kullback-Leibler divergence between two sparsely sampled…

Data Analysis, Statistics and Probability · Physics 2023-02-24 Angelo Piga , Lluc Font-Pomarol , Marta Sales-Pardo , Roger Guimerà

Based on independently distributed $X_1 \sim N_p(\theta_1, \sigma^2_1 I_p)$ and $X_2 \sim N_p(\theta_2, \sigma^2_2 I_p)$, we consider the efficiency of various predictive density estimators for $Y_1 \sim N_p(\theta_1, \sigma^2_Y I_p)$, with…

Statistics Theory · Mathematics 2017-09-25 Éric Marchand , Abdolnasser Sadeghkhani

The variational framework for learning inducing variables (Titsias, 2009a) has had a large impact on the Gaussian process literature. The framework may be interpreted as minimizing a rigorously defined Kullback-Leibler divergence between…

Machine Learning · Statistics 2015-12-07 Alexander G. de G. Matthews , James Hensman , Richard E. Turner , Zoubin Ghahramani

In genomics, differential abundance and expression analyses are complicated by the compositional nature of sequence count data, which reflect only relative-not absolute-abundances or expression levels. Many existing methods attempt to…

Methodology · Statistics 2025-12-16 Won Gu , Francesca Chiaromonte , Justin D. Silverman

Modeling sparse data such as microbiome and transcriptomics (RNA-seq) data is very challenging due to the exceeded number of zeros and skewness of the distribution. Many probabilistic models have been used for modeling sparse data,…

Methodology · Statistics 2021-12-30 Hani Aldirawi , Jie Yang

Consider a situation of analyzing high-dimensional count data containing an excess of near-zero counts with a small number of moderate or large counts. Assuming that the observations are modeled by a Poisson distribution, we are interested…

Statistics Theory · Mathematics 2025-11-27 Sayantan Paul , Arijit Chakrabarti

In this paper, we propose a theoretical analysis of the algorithm ISDE, introduced in previous work. From a dataset, ISDE learns a density written as a product of marginal density estimators over a partition of the features. We show that…

Statistics Theory · Mathematics 2022-05-09 Louis Pujol

Quantile regression, a robust method for estimating conditional quantiles, has advanced significantly in fields such as econometrics, statistics, and machine learning. In high-dimensional settings, where the number of covariates exceeds…

Machine Learning · Statistics 2024-09-04 The Tien Mai

Many statistical studies are concerned with the analysis of observations organized in a matrix form whose elements are count data. When these observations are assumed to follow a Poisson or a multinomial distribution, it is of interest to…

Statistics Theory · Mathematics 2022-01-04 Jérémie Bigot , Charles Deledalle

We study full Bayesian procedures for sparse linear regression when errors have a symmetric but otherwise unknown distribution. The unknown error distribution is endowed with a symmetrized Dirichlet process mixture of Gaussians. For the…

Statistics Theory · Mathematics 2019-03-26 Minwoo Chae , Lizhen Lin , David B. Dunson

The Bayesian predictive density has complex representation and does not belong to any finite-dimensional statistical model except for in limited situations. In this paper, we introduce its simple approximate representation employing its…

Statistics Theory · Mathematics 2020-10-30 Michiko Okudo , Fumiyasu Komaki

A popular approach within the signal processing and machine learning communities consists in modelling signals as sparse linear combinations of atoms selected from a learned dictionary. While this paradigm has led to numerous empirical…

Machine Learning · Statistics 2012-10-03 Rodolphe Jenatton , Rémi Gribonval , Francis Bach