English
Related papers

Related papers: A data-based power transformation for compositiona…

200 papers

Compositional generalization is a crucial step towards developing data-efficient intelligent machines that generalize in human-like ways. In this work, we tackle a challenging form of distribution shift, termed compositional shift, where…

Machine Learning · Computer Science 2025-07-14 Divyat Mahajan , Mohammad Pezeshki , Charles Arnal , Ioannis Mitliagkas , Kartik Ahuja , Pascal Vincent

Applications such as the analysis of microbiome data have led to renewed interest in statistical methods for compositional data, i.e., multivariate data in the form of probability vectors that contain relative proportions. In particular,…

Methodology · Statistics 2021-09-13 Shiqing Yu , Mathias Drton , Ali Shojaie

Compositional data arise in many real-life applications and versatile methods for properly analyzing this type of data in the regression context are needed. When parametric assumptions do not hold or are difficult to verify, non-parametric…

Methodology · Statistics 2023-09-07 Michail Tsagris , Abdulaziz Alenazi , Connie Stewart

Compositional data consist of known compositions vectors whose components are positive and defined in the interval (0,1) representing proportions or fractions of a "whole". The sum of these components must be equal to one. Compositional…

Applications · Statistics 2015-07-02 Taciana K. O. Shimizu , Francisco Louzada , Adriano K. Suzuki , Ricardo S. Ehlers

Compositional data are commonly known as multivariate observations carrying relative information. Even though the case of vector or even two-factorial compositional data (compositional tables) is already well described in the literature,…

Methodology · Statistics 2022-01-26 Kamila Fačevicová , Peter Filzmoser , Karel Hron

This article present a method of mutual transformation between count model and composition model. Offer the mathematical view of classical radio and log-radio in compositional data analysis and expand the idea of mixture model of counts…

Statistics Theory · Mathematics 2025-09-12 Guanyi Wu

In compositional data, an observation is a vector with non-negative components which sum to a constant, typically 1. Data of this type arise in many areas, such as geology, archaeology, biology, economics and political science amongst…

Methodology · Statistics 2015-11-25 Michail Tsagris

Compositional data (i.e., data comprising random variables that sum up to a constant) arises in many applications including microbiome studies, chemical ecology, political science, and experimental designs. Yet when compositional data serve…

Methodology · Statistics 2025-01-03 Ritwik Bhaduri , Siyuan Ma , Lucas Janson

Compositional data, also referred to as simplicial data, naturally arise in many scientific domains such as geochemistry, microbiology, and economics. In such domains, obtaining sensible lower-dimensional representations and modes of…

In real world applications dealing with compositional datasets, it is easy to face the presence of structural zeros. The latter arise when, due to physical limitations, one or more variables are intrinsically zero for a subset of the…

Methodology · Statistics 2025-10-28 Francesco Porro , Fabio Rapallo , Sara Sommariva

Kendall transformation is a conversion of an ordered feature into a vector of pairwise order relations between individual values. This way, it preserves ranking of observations and represents it in a categorical form. Such transformation…

Machine Learning · Computer Science 2023-08-15 Miron Bartosz Kursa

Power transforms are popular parametric methods for making data more Gaussian-like, and are widely used as preprocessing steps in statistical analysis and machine learning. However, we find that direct implementations of power transforms…

Machine Learning · Computer Science 2026-04-16 Xuefeng Xu , Graham Cormode

Compositional data are met in many different fields, such as economics, archaeometry, ecology, geology and political sciences. Regression where the dependent variable is a composition is usually carried out via a log-ratio transformation of…

Methodology · Statistics 2017-06-08 Michail Tsagris , Connie Stewart

Correspondence analysis is a dimension reduction method for visualization of nonnegative data sets, in particular contingency tables ; but it depends on the marginals of the data set. Two transformations of the data have been proposed to…

Methodology · Statistics 2025-10-20 Vartan Choulakian

Functional data typically contains amplitude and phase variation. In many data situations, phase variation is treated as a nuisance effect and is removed during preprocessing, although it may contain valuable information. In this note, we…

Methodology · Statistics 2021-01-01 Clara Happ , Fabian Scheipl , Alice-Agnes Gabriel , Sonja Greven

Compositional data, representing proportions constrained to the simplex, arise in diverse fields such as geosciences, ecology, genomics, and microbiome research. Existing nonparametric density estimation methods often rely on…

Methodology · Statistics 2025-10-10 Jiajin Xie , Yong Wang , Eduardo García-Portugués

High-dimensional compositional data, such as those from human microbiome studies, pose unique statistical challenges due to the simplex constraint and excess zeros. While dimension reduction is indispensable for analyzing such data,…

Methodology · Statistics 2025-09-09 Junyoung Park , Cheolwoo Park , Jeongyoun Ahn

Many variables in the social, physical, and biosciences, including neuroscience, are non-normally distributed. To improve the statistical properties of such data, or to allow parametric testing, logarithmic or logit transformations are…

Methodology · Statistics 2018-01-08 Sacha Jennifer van Albada , Peter A. Robinson

Testing differences in mean vectors is a fundamental task in the analysis of high-dimensional compositional data. Existing methods may suffer from low power if the underlying signal pattern is in a situation that does not favor the deployed…

Methodology · Statistics 2025-03-11 Danning Li , Lingzhou Xue , Haoyi Yang , Xiufan Yu

The operating status of power systems is influenced by growing varieties of factors, resulting from the developing sizes and complexity of power systems; in this situation, the modelbased methods need be revisited. A data-driven method, as…

Methodology · Statistics 2016-07-07 Xinyi Xu , Xing He , Qian Ai , Robert C. Qiu