English
Related papers

Related papers: LinDA: linear models for differential abundance an…

200 papers

We introduce a novel approach to compositional data analysis based on $L^{\infty}$-normalization, addressing challenges posed by zero-rich high-throughput data. Traditional methods like Aitchison's transformations require excluding zeros,…

Computation · Statistics 2025-03-28 Pawel Gajer , Jacques Ravel

In the regression problem, L1 and L2 are the most commonly used loss functions, which produce mean predictions with different biases. However, the predictions are neither robust nor adequate enough since they only capture a few conditional…

Machine Learning · Computer Science 2019-11-14 Faen Zhang , Xinyu Fan , Hui Xu , Pengcheng Zhou , Yujian He , Junlong Liu

We propose a simple data augmentation protocol aimed at providing a compositional inductive bias in conditional and unconditional sequence models. Under this protocol, synthetic training examples are constructed by taking real training…

Computation and Language · Computer Science 2020-05-20 Jacob Andreas

Observational studies often benefit from an abundance of observational units. This can lead to studies that -- while challenged by issues of internal validity -- have inferences derived from sample sizes substantially larger than randomized…

Methodology · Statistics 2020-08-24 Rachael C. Aikens , Dylan Greaves , Michael Baiocchi

Microbiome interventions provide valuable data about microbial ecosystem structure and dynamics. Despite their ubiquity in microbiome research, few rigorous data analysis approaches are available. In this study, we extend transfer…

Applications · Statistics 2023-06-13 Kris Sankaran , Pratheepa Jeganathan

A network lasso enables us to construct a model for each sample, which is known as multi-task learning. Existing methods for multi-task learning cannot be applied to compositional data due to their intrinsic properties. In this paper, we…

Methodology · Statistics 2023-01-04 Akira Okazaki , Shuichi Kawano

The method of instrumental variables provides a fundamental and practical tool for causal inference in many empirical studies where unmeasured confounding between the treatments and the outcome is present. Modern data such as the genetical…

Methodology · Statistics 2022-10-28 Ziang Niu , Yuwen Gu , Wei Li

In many real-world regression tasks, the data distribution is heavily skewed, and models learn predominantly from abundant majority samples while failing to predict minority labels accurately. While imbalanced classification has been…

Machine Learning · Computer Science 2025-09-30 Shayan Alahyari

Data imbalance is common in production data, where controlled production settings require data to fall within a narrow range of variation and data are collected with quality assessment in mind, rather than data analytic insights. This…

Machine Learning · Statistics 2021-12-17 Rune D. Kjærsgaard , Manja G. Grønberg , Line K. H. Clemmensen

We present a new method based on Functional Data Analysis (FDA) for detecting associations between one or more scalar covariates and a longitudinal response, while correcting for other variables. Our methods exploit the temporal structure…

Applications · Statistics 2014-04-30 Matthew Reimherr , Dan Nicolae

We consider the complex data modeling problem motivated by the zero-inflated and overdispersed data from microbiome studies. Analyzing how microbiome abundance is associated with human biological features, such as BMI, is of great…

Methodology · Statistics 2025-03-31 Zirui Wang , Tianying Wang

The Dirichlet-multinomial (DM) distribution plays a fundamental role in modern statistical methodology development and application. Recently, the DM distribution and its variants have been used extensively to model multivariate count data…

Methodology · Statistics 2023-02-27 Matthew D. Koslovsky

Understanding the complex interactions within the microbiome is crucial for developing effective diagnostic and therapeutic strategies. Traditional machine learning models often lack interpretability, which is essential for clinical and…

Machine Learning · Computer Science 2024-10-22 Swagatam Haldar , Christoph Stein-Thoeringer , Vadim Borisov

Compositional data, which are vectors of proportions constrained to the probability simplex, arise frequently in modern scientific applications, including microbiome relative abundances across body sites and cell-type mixture weights…

Methodology · Statistics 2026-05-08 Shuangjie Zhang , Bani K. Mallick , Yang Ni

Addressing the simultaneous identification of contributory variables while controlling the false discovery rate (FDR) in high-dimensional data is a crucial statistical challenge. In this paper, we propose a novel model-free variable…

Methodology · Statistics 2024-04-23 Yixin Han , Xu Guo , Changliang Zou

There are two popular general approaches for the analysis and visualization of a contingency table and a compositional data set: Correspondence analysis (CA) and log ratio analysis (LRA). LRA includes two independently well developed…

Methodology · Statistics 2020-09-14 J. Allard , S. Champigny , V. Choulakian , S. Mahdi

Due to their type of mathematical construction, the use of standard financial ratios in studies analysing the financial health of a group of firms leads to a series of statistical problems that can invalidate the results obtained. These…

Statistical Finance · Quantitative Finance 2025-05-28 Salvador Linares-Mustarós , Maria Àngels Farreras-Noguer , Núria Arimany-Serrat , Germà Coenders

We argue that the selective inclusion of data points based on latent objectives is common in practical situations, such as music sequences. Since this selection process often distorts statistical analysis, previous work primarily views it…

Machine Learning · Computer Science 2024-07-02 Yujia Zheng , Zeyu Tang , Yiwen Qiu , Bernhard Schölkopf , Kun Zhang

Compositional data arise in many real-life applications and versatile methods for properly analyzing this type of data in the regression context are needed. When parametric assumptions do not hold or are difficult to verify, non-parametric…

Methodology · Statistics 2023-09-07 Michail Tsagris , Abdulaziz Alenazi , Connie Stewart

Discriminant analysis is a useful classification method. Variable selection for discriminant analysis is becoming more and more im- portant in a high-dimensional setting. This paper is concerned with the binary-class problems of main and…

Methodology · Statistics 2017-02-16 Deqiang Zheng , Jinzhu Jia , Xiangzhong Fang , Xiuhua Guo