English
Related papers

Related papers: On the distribution of isometric log-ratio transfo…

200 papers

Posterior distribution over a countable set M of continuous data-sampling distributions piles up at L-projection of the true distribution r on M, provided that the L-projection is unique. If there are several L-projections of r on M, then…

Probability · Mathematics 2007-10-10 M. Grendar

Probabilistic finite mixture models are widely used for unsupervised clustering. These models can often be improved by adapting them to the topology of the data. For instance, in order to classify spatially adjacent data points similarly,…

Computer Vision and Pattern Recognition · Computer Science 2022-02-09 Jonathan Vacher , Claire Launay , Ruben Coen-Cagli

We study the isotonic regression estimator over a general countable pre-ordered set. We obtain the limiting distribution of the estimator and study its properties. It is proved that, under some general assumptions, the limiting distribution…

Statistics Theory · Mathematics 2018-11-06 Dragi Anevski , Vladimir Pastukhov

Empirical analyses of ordinal outcomes using repeated cross-sectional data rely on marginal distributions, leaving the joint distribution unobserved and the sources of distributional change unidentified. This paper develops a framework to…

Econometrics · Economics 2026-04-28 Rami V. Tabri

Recent research has established sufficient conditions for finite mixture models to be identifiable from grouped observations. These conditions allow the mixture components to be nonparametric and have substantial (or even total) overlap.…

Machine Learning · Statistics 2020-06-16 Alexander Ritchie , Robert A. Vandermeulen , Clayton Scott

Natural language data follows a power-law distribution, with most knowledge and skills appearing at very low frequency. While a common intuition suggests that reweighting or curating data towards a uniform distribution may help models…

Artificial Intelligence · Computer Science 2026-04-28 Zixuan Wang , Xingyu Dang , Jason D. Lee , Kaifeng Lyu

Integrative analysis of datasets generated by multiple cohorts is a widely-used approach for increasing sample size, precision of population estimators, and generalizability of analysis results in epidemiological studies. However, often…

Learning from a limited number of samples is challenging since the learned model can easily become overfitted based on the biased distribution formed by only a few training examples. In this paper, we calibrate the distribution of these…

Machine Learning · Computer Science 2021-08-17 Shuo Yang , Lu Liu , Min Xu

Compositional data, which is data consisting of fractions or probabilities, is common in many fields including ecology, economics, physical science and political science. If these data would otherwise be normally distributed, their spread…

Methodology · Statistics 2022-07-26 Matthew P. Adams

In his 1986 book, Aitchison explains that compositional data is regularly mishandled in statistical analyses, a pattern that continues to this day. The Dirichlet Type I distribution is a multivariate distribution commonly used to model a…

Statistics Theory · Mathematics 2018-04-06 Sean van der Merwe , Daan de Waal

Deep generative models parametrized up to a normalizing constant (e.g. energy-based models) are difficult to train by maximizing the likelihood of the data because the likelihood and/or gradients thereof cannot be explicitly or efficiently…

Machine Learning · Computer Science 2022-12-26 Frederic Koehler , Alexander Heckett , Andrej Risteski

Identifying which taxa in our microbiota are associated with traits of interest is important for advancing science and health. However, the identification is challenging because the measured vector of taxa counts (by amplicon sequencing) is…

Genomics · Quantitative Biology 2020-03-31 Barak Brill , Amnon Amir , Ruth Heller

It is generally known that counting statistics is not correctly described by a Gaussian approximation. Nevertheless, in neutron scattering, it is common practice to apply this approximation to the counting statistics; also at low counting…

Data Analysis, Statistics and Probability · Physics 2020-06-09 Jakob Lassa , Magnus Egede Bøggild , Per Hedegård , Kim Lefmann

Binomial data with unknown sizes often appear in biological and medical sciences and are usually overdispersed. All previous methods used parametric models and only considered overdispersion due to the variation of sizes. The proposed…

Statistics Theory · Mathematics 2007-06-13 Wei Zhang

We investigate the asymptotic distributions of coordinates of regression M-estimates in the moderate $p/n$ regime, where the number of covariates $p$ grows proportionally with the sample size $n$. Under appropriate regularity conditions, we…

Statistics Theory · Mathematics 2016-12-20 Lihua Lei , Peter J. Bickel , Noureddine El Karoui

Nonparametric density estimation for compositional data supported on the simplex is examined under a missing at random mechanism. Rather than imputing missing values and estimating the density from a completed data set, we adopt a strategy…

Methodology · Statistics 2026-03-10 Hanen Daayeb , Wissem Jedidi , Salah Khardani , Guanjie Lyu , Frédéric Ouimet

Recent developments in extracting and processing biological and clinical data are allowing quantitative approaches to studying living systems. High-throughput sequencing, expression profiles, proteomics, and electronic health records are…

Quantitative Methods · Quantitative Biology 2010-10-22 Vladimir Trifonov , Laura Pasqualucci , Riccardo Dalla-Favera , Raul Rabadan

Mixture distributions are a workhorse model for multimodal data in information theory, signal processing, and machine learning. Yet even when each component density is simple, the differential entropy of the mixture is notoriously hard to…

Information Theory · Computer Science 2026-02-18 Namyoon Lee

Few-shot learning is a rapidly evolving area of research in machine learning where the goal is to classify unlabeled data with only one or "a few" labeled exemplary samples. Neural networks are typically trained to minimize a distance…

Computer Vision and Pattern Recognition · Computer Science 2022-12-09 Samuel Hess , Gregory Ditzler

Given an m-dimensional compact submanifold $\mathbf{M}$ of Euclidean space $\mathbf{R}^s$, the concept of mean location of a distribution, related to mean or expected vector, is generalized to more general $\mathbf{R}^s$-valued functionals…

Statistics Theory · Mathematics 2007-08-07 Harrie Hendriks , Zinoviy Landsman