English
Related papers

Related papers: On the distribution of isometric log-ratio transfo…

200 papers

Compositional data, representing proportions constrained to the simplex, arise in diverse fields such as geosciences, ecology, genomics, and microbiome research. Existing nonparametric density estimation methods often rely on…

Methodology · Statistics 2025-10-10 Jiajin Xie , Yong Wang , Eduardo García-Portugués

Classification with imbalanced data is a common challenge in data analysis, where certain classes (minority classes) account for a small fraction of the training data compared with other classes (majority classes). Classical statistical…

Statistics Theory · Mathematics 2025-02-18 Jingyang Lyu , Kangjie Zhou , Yiqiao Zhong

Microbiome data are complex in nature, involving high dimensionality, compositionally, zero inflation, and taxonomic hierarchy. Compositional data reside in a simplex that does not admit the standard Euclidean geometry. Most existing…

Methodology · Statistics 2020-11-12 Gen Li , Yan Li , Kun Chen

We examine the normal approximation of the modified likelihood root, an inferential tool from higher-order asymptotic theory, for the linear exponential and location-scale family. We show that the $r^\star$ statistic can be thought of as a…

Methodology · Statistics 2022-01-13 Yanbo Tang , Nancy Reid

We study the joint distribution of the number of occurrences of members of a collection of nonoverlapping motifs in digital data. We deal with finite and countably infinite collections. For infinite collections, the setting requires that we…

Probability · Mathematics 2016-07-05 Jeffrey Gaither , Hosam Mahmoud , Mark Daniel Ward

Sampling from multimodal distributions is a challenging task in scientific computing. When a distribution has an exact symmetry between the modes, direct jumps among them can accelerate the samplings significantly. However, the…

Numerical Analysis · Mathematics 2024-01-05 Lexing Ying

Count data take on non-negative integer values and are challenging to properly analyze using standard linear-Gaussian methods such as linear regression and principal components analysis. Generalized linear models enable direct modeling of…

Methodology · Statistics 2020-01-14 F. William Townes

Bimodal truncated count distributions are frequently observed in aggregate survey data and in user ratings when respondents are mixed in their opinion. They also arise in censored count data, where the highest category might create an…

Methodology · Statistics 2014-01-24 Pragya Sur , Galit Shmueli , Smarajit Bose , Paromita Dubey

Many scientific datasets are compositional in nature. Important biological examples include species abundances in ecology, cell-type compositions derived from single-cell sequencing data, and amplicon abundance data in microbiome research.…

Machine Learning · Computer Science 2024-05-29 Elisabeth Ailer , Christian L. Müller , Niki Kilbertus

While calibration of probabilistic predictions has been widely studied, this paper rather addresses calibration of likelihood functions. This has been discussed, especially in biometrics, in cases with only two exhaustive and mutually…

Machine Learning · Computer Science 2025-09-04 Paul-Gauthier Noé , Andreas Nautsch , Driss Matrouf , Pierre-Michel Bousquet , Jean-François Bonastre

The Poisson distribution is the default choice of likelihood for probabilistic models of count data. However, due to the equidispersion contraint of the Poisson, such models may have predictive uncertainty that is artificially inflated.…

Methodology · Statistics 2025-07-15 Jimmy Lederman , Aaron Schein

In the analysis of count data often the equidispersion assumption is not suitable, hence the Poisson regression model is inappropriate. As a generalization of the Poisson distribution, the COM-Poisson distribution can deal with under-,…

Many scientific and industrial processes produce data that is best analysed as vectors of relative values, often called compositions or proportions. The Dirichlet distribution is a natural distribution to use for composition or proportion…

Methodology · Statistics 2020-04-15 Sean van der Merwe

Traditional methods for the analysis of compositional data consider the log-ratios between all different pairs of variables with equal weight, typically in the form of aggregated contributions. This is not meaningful in contexts where it is…

Methodology · Statistics 2022-01-27 Christopher Rieser , Peter Filzmoser

The assessment of diversity and similarity is relevant in monitoring the status of ecosystems. The respective indicators are based on the taxonomic composition of biological communities of interest, currently estimated through the…

Applications · Statistics 2018-10-12 Fabio Divino , Johanna Ärje , Antti Penttinen , Kristian Meissner , Salme Kärkkäinen

In the paper, multivariate probability distributions are considered that are representable as scale mixtures of multivariate elliptically contoured stable distributions. It is demonstrated that these distributions form a special subclass of…

Probability · Mathematics 2019-12-05 Victor Korolev , Alexander Zeifman

This paper develops a general inferential framework for discrete copulas on finite supports in any dimension. The copula of a multivariate discrete distribution is defined as Csiszar's I-projection (i.e., the minimum-Kullback-Leibler…

Statistics Theory · Mathematics 2025-06-17 Gery Geenens , Ivan Kojadinovic , Tommaso Martini

Recently Liu and Wang derived the likelihood ratio test (LRT) statistic and its asymptotic distribution for testing equality of two multinomial distributions vs. the alternative that the second distribution is larger in terms of increasing…

Statistics Theory · Mathematics 2007-06-13 Arthur Cohen , John Kolassa , Harold Sackrowitz

In this document we achieve exact and asymptotic enumeration of words, compositions over a finite group, and/or integer compositions characterized by local restrictions and, separately, subsequence pattern avoidance. We also count…

Combinatorics · Mathematics 2019-04-19 Andrew MacFie

In many applications, data cluster. Failing to take the cluster structure into consideration generally leads to underestimated variances of point estimators and inflated type I errors in hypothesis tests. Many circumstance-dependent…

Methodology · Statistics 2025-07-21 Jiahua Chen , Pengfei Li , Yukun Liu , James V. Zidek