English
Related papers

Related papers: A New Approach to Compositional Data Analysis usin…

200 papers

Coupling arguments are a central tool for bounding the deviation between two stochastic processes, but traditionally have been limited to Wasserstein metrics. In this paper, we apply the shifted composition rule--an information-theoretic…

Statistics Theory · Mathematics 2024-12-25 Jason M. Altschuler , Sinho Chewi

Biological sequencing data consist of read counts, e.g. of specified taxa and often exhibit sparsity (zero-count inflation) and overdispersion (extra-Poisson variability). As most sequencing techniques provide an arbitrary total count,…

Applications · Statistics 2024-07-01 Noora Kartiosuo , Jaakko Nevalainen , Olli Raitakari , Katja Pahkala , Kari Auranen

Ongoing advances in microbiome profiling have allowed unprecedented insights into the molecular activities of microbial communities. This has fueled a strong scientific interest in understanding the critical role the microbiome plays in…

Methodology · Statistics 2024-11-18 Satabdi Saha , Liangliang Zhang , Kim-Anh Do , Christine B. Peterson

We introduce a compositional data-driven methodology with noisy data for designing fully-decentralized safety controllers applicable to large-scale interconnected networks, encompassing a vast number of subsystems with unknown mathematical…

Systems and Control · Electrical Eng. & Systems 2025-06-18 Omid Akbarzadeh , Amy Nejati , Abolfazl Lavaei

This work is closely related to the theories of set estimation and manifold estimation. Our object of interest is a, possibly lower-dimensional, compact set $S \subset {\mathbb R}^d$. The general aim is to identify (via stochastic…

Statistics Theory · Mathematics 2017-11-06 Catherine Aaron , Alejandro Cholaquidis , Antonio Cuevas

Modern inference and learning often hinge on identifying low-dimensional structures that approximate large scale data. Subspace clustering achieves this through a union of linear subspaces. However, in contemporary applications data is…

Machine Learning · Computer Science 2018-08-03 Daniel L. Pimentel-Alarcón , Usman Mahmood

In many applications where collecting data is expensive, for example neuroscience or medical imaging, the sample size is typically small compared to the feature dimension. It is challenging in this setting to train expressive, non-linear…

Machine Learning · Computer Science 2019-04-23 Sergul Aydore , Bertrand Thirion , Gael Varoquaux

A new approach to data compression is developed and applied to multimedia content. This method separates messages into components suitable for both lossless coding and 'lossy' or statistical coding techniques, compressing complex objects by…

Information Theory · Computer Science 2011-12-26 John Scoville

Functional linear discriminant analysis offers a simple yet efficient method for classification, with the possibility of achieving a perfect classification. Several methods are proposed in the literature that mostly address the…

Methodology · Statistics 2020-12-14 Juhyun Park , Jeongyoun Ahn , Yongho Jeon

We propose a computationally efficient method to construct nonparametric, heteroscedastic prediction bands for uncertainty quantification, with or without any user-specified predictive model. Our approach provides an alternative to the…

Machine Learning · Statistics 2023-01-18 Tengyuan Liang

Nonnegative matrix factorization (NMF) approximates a nonnegative matrix, $X$, by the product of two nonnegative factors, $WH$, where $W$ has $r$ columns and $H$ has $r$ rows. In this paper, we consider NMF using the component-wise L1 norm…

Machine Learning · Computer Science 2026-04-01 Giovanni Seraghiti , Kévin Dubrulle , Arnaud Vandaele , Nicolas Gillis

Health care claims data refer to information generated from interactions within health systems. They have been used in health services research for decades to assess effectiveness of interventions, determine the quality of medical care,…

Applications · Statistics 2018-01-16 Jacob Spertus , Samrachana Adhikari , Sharon-Lise Normand

Modern data analysis frequently involves variables with highly non-Gaussian marginal distributions. However, commonly used analysis methods are most effective with roughly Gaussian data. This paper introduces an automatic transformation…

Methodology · Statistics 2016-01-11 Qing Feng , Jan Hannig , J. S. Marron

In this paper we revisit the classical method of partitioning classification and study its convergence rate under relaxed conditions, both for observable (non-privatised) and for privatised data. We consider the problem of classification in…

Machine Learning · Statistics 2025-09-09 Balázs Csanád Csáji , László Györfi , Ambrus Tamás , Harro Walk

This paper establishes Lipschitz stability for the simultaneous recovery of a variable density coefficient and the initial displacement in a damped biharmonic wave equation. The data consist of the boundary Cauchy data for the Laplacian of…

Analysis of PDEs · Mathematics 2026-05-18 Minghui Bi , Yixian Gao

Compositional data consist of known compositions vectors whose components are positive and defined in the interval (0,1) representing proportions or fractions of a "whole". The sum of these components must be equal to one. Compositional…

Applications · Statistics 2015-07-02 Taciana K. O. Shimizu , Francisco Louzada , Adriano K. Suzuki , Ricardo S. Ehlers

Dimension reduction of high-dimensional microbiome data facilitates subsequent analysis such as regression and clustering. Most existing reduction methods cannot fully accommodate the special features of the data such as count-valued and…

Methodology · Statistics 2023-05-02 Tianchen Xu , Ryan T. Demmer , Gen Li

The Lipschitz constant is a key measure for certifying the robustness of neural networks to input perturbations. However, computing the exact constant is NP-hard, and standard approaches to estimate the Lipschitz constant involve solving a…

Machine Learning · Computer Science 2026-04-14 Yuezhu Xu , S. Sivaranjani

Many biological high-throughput data sets, such as targeted amplicon-based and metagenomic sequencing data, are compositional in nature. A common exploratory data analysis task is to infer statistical associations between the…

Methodology · Statistics 2020-07-28 Aditya Mishra , Christian L. Muller

This paper discusses the incorporation of local sparsity information, e.g. in each pixel of an image, via minimization of the $\ell^{1,\infty}$-norm. We discuss the basic properties of this norm when used as a regularization functional and…

Optimization and Control · Mathematics 2015-06-12 Pia Heins , Michael Moeller , Martin Burger