English
Related papers

Related papers: A Dirichlet Regression Model for Compositional Dat…

200 papers

The Regression Discontinuity (RD) design is a widely used non-experimental method for causal inference and program evaluation. While its canonical formulation only requires a score and an outcome variable, it is common in empirical work to…

Methodology · Statistics 2022-08-25 Matias D. Cattaneo , Luke Keele , Rocio Titiunik

We discuss a bivariate beta distribution that can model arbitrary beta-distributed marginals with a positive correlation. The distribution is constructed from six independent gamma-distributed random variates. We show how the parameters of…

Statistics Theory · Mathematics 2021-06-03 Susanne Trick , Frank Jäkel , Constantin A. Rothkopf

The paper revisits the $\alpha$--regression framework for compositional data. The model uses a flexible power transformation parameterized by $\alpha$ to interpolate between raw data analysis and log--ratio methods, naturally handling zeros…

Methodology · Statistics 2026-05-14 Michail Tsagris , Yannis Pantazis

Compositional generalization is a crucial step towards developing data-efficient intelligent machines that generalize in human-like ways. In this work, we tackle a challenging form of distribution shift, termed compositional shift, where…

Machine Learning · Computer Science 2025-07-14 Divyat Mahajan , Mohammad Pezeshki , Charles Arnal , Ioannis Mitliagkas , Kartik Ahuja , Pascal Vincent

A frequent challenge encountered with ecological data is how to interpret, analyze, or model data having a high proportion of zeros. Much attention has been given to zero-inflated count data, whereas models for non-negative continuous data…

Methodology · Statistics 2022-05-23 Becky Tang , Henry A Frye , Alan E. Gelfand , John A Silander

Covariate-adaptive randomization is widely used in clinical trials to balance prognostic factors, and regression adjustments are often adopted to further enhance the estimation and inference efficiency. In practice, the covariates may…

Methodology · Statistics 2025-08-15 Wanjia Fu , Yingying Ma , Hanzhong Liu

Change point detection algorithms have numerous applications in fields of scientific and economic importance. We consider the problem of change point detection on compositional multivariate data (each sample is a probability mass function),…

Applications · Statistics 2019-01-16 Prabuchandran K. J. , Nitin Singh , Pankaj Dayama , Vinayaka Pandit

In regression problems where there is no known true underlying model, conformal prediction methods enable prediction intervals to be constructed without any assumptions on the distribution of the underlying data, except that the training…

Methodology · Statistics 2023-01-31 Wenyu Chen , Kelli-Jean Chun , Rina Foygel Barber

The zero-inflated logistic regression model accommodates binary responses with excess zeros, which often arise from a latent mixture of susceptible and insusceptible subpopulations or asymmetric misclassification of the response. The model…

Methodology · Statistics 2026-04-23 Yui Tomo , Shinto Eguchi , Daisuke Yoneoka

Collecting robotic manipulation data is expensive, making it impractical to acquire demonstrations for the combinatorially large space of tasks that arise in multi-object, multi-robot, and multi-environment settings. While recent generative…

Compositional data, which are vectors of proportions constrained to the probability simplex, arise frequently in modern scientific applications, including microbiome relative abundances across body sites and cell-type mixture weights…

Methodology · Statistics 2026-05-08 Shuangjie Zhang , Bani K. Mallick , Yang Ni

Regression models for limited continuous dependent variables having a non-negligible probability of attaining exactly their limits are presented. The models differ in the number of parameters and in their flexibility. Fractional data being…

Applications · Statistics 2012-05-31 Fabio Sigrist , Werner A. Stahel

Compositional data analysis is carried out either by neglecting the compositional constraint and applying standard multivariate data analysis, or by transforming the data using the logs of the ratios of the components. In this work we…

Methodology · Statistics 2011-06-17 Michail T. Tsagris , Simon Preston , Andrew T. A. Wood

This paper considers the problem of kernel regression and classification with possibly unobservable response variables in the data, where the mechanism that causes the absence of information is unknown and can depend on both predictors and…

Statistics Theory · Mathematics 2022-12-07 Majid Mojirsheibani , William Pouliot , Andre Shakhbandaryan

Shuffled regression concerns settings in which covariates and responses are observed without their correct pairing. In dependent-data problems, a second form of missing correspondence can arise when responses are also detached from the…

Statistics Theory · Mathematics 2026-03-23 Anik Burman , Sayantan Choudhury , Debangan Dey

This paper is motivated by the recent interest in the analysis of high dimen- sional microbiome data. A key feature of this data is the presence of `structural zeros' which are microbes missing from an observation vector due to an…

Applications · Statistics 2016-05-23 Abhishek Kaul , Ori Davidov , Shyamal D. Peddada

Compositional data represent a specific family of multivariate data, where the information of interest is contained in the ratios between parts rather than in absolute values of single parts. The analysis of such specific data is…

Score matching is a vital tool for learning the distribution of data with applications across many areas including diffusion processes, energy based modelling, and graphical model estimation. Despite all these applications, little work…

Machine Learning · Statistics 2025-06-03 Josh Givens , Song Liu , Henry W J Reeve

In tasks like semantic parsing, instruction following, and question answering, standard deep networks fail to generalize compositionally from small datasets. Many existing approaches overcome this limitation with model architectures that…

Computation and Language · Computer Science 2023-07-06 Ekin Akyürek , Jacob Andreas

Count data are common in medical research. When these data have more zeros than expected by the most used count distributions, it is common to employ a zero-inflated regression model. However, the interpretability of these models is much…

Methodology · Statistics 2025-09-30 Gustavo H. A. Pereira , Jeremias Leão , Manoel Santos-Neto , Jianwen Cai