English
Related papers

Related papers: Testing for differential abundance in compositiona…

200 papers

We propose a modelling framework which allows for the estimation of abundances from trace counts. This indirect method of estimating abundance is attractive due to the relative affordability with which it may be carried out, and the…

Methodology · Statistics 2022-06-14 Niamh Mimnagh , Iuri Ferreira , Luciano Verdade , Rafael de Andrade Moral

There is an implicit assumption in software testing that more diverse and varied test data is needed for effective testing and to achieve different types and levels of coverage. Generic approaches based on information theory to measure and…

Software Engineering · Computer Science 2017-09-19 Robert Feldt , Simon Poulding

This paper studies Difference-in-Differences (DiD) setups with repeated cross-sectional data and potential compositional changes across time periods. We begin our analysis by deriving the efficient influence function and the semiparametric…

Econometrics · Economics 2025-11-17 Pedro H. C. Sant'Anna , Qi Xu

It is largely taken for granted that differential abundance analysis is, by default, the best first step when analyzing genomic data. We argue that this is not necessarily the case. In this article, we identify key limitations that are…

Methodology · Statistics 2021-06-09 Thomas P Quinn , Elliott Gordon-Rodriguez , Ionas Erb

Data augmentation techniques are widely used in low-resource automatic morphological inflection to overcome data sparsity. However, the full implications of these techniques remain poorly understood. In this study, we aim to shed light on…

Computation and Language · Computer Science 2023-10-25 Farhan Samir , Miikka Silfverberg

Instrumental variables are a popular study design for the estimation of treatment effects in the presence of unobserved confounders. In the canonical instrumental variables design, the instrument is a binary variable. In many settings,…

Methodology · Statistics 2024-10-10 Prabrisha Rakshit , Alexander Levis , Luke Keele

Knowledge about existence, strength, and dominant direction of causal influences is of paramount importance for understanding complex systems. With limited amounts of realistic data, however, current methods for investigating causal links…

Data Analysis, Statistics and Probability · Physics 2020-10-20 Erik Laminski , Klaus R. Pawelzik

The last twenty years have witnessed molecular data emerge as a primary research instrument in most branches of mycology. Fungal systematics, taxonomy, and ecology have all seen tremendous progress and have undergone rapid, far-reaching…

Change point detection algorithms have numerous applications in fields of scientific and economic importance. We consider the problem of change point detection on compositional multivariate data (each sample is a probability mass function),…

Applications · Statistics 2019-01-16 Prabuchandran K. J. , Nitin Singh , Pankaj Dayama , Vinayaka Pandit

We perform differential expression analysis of high-throughput sequencing count data under a Bayesian nonparametric framework, removing sophisticated ad-hoc pre-processing steps commonly required in existing algorithms. We propose to use…

Applications · Statistics 2017-05-04 Siamak Zamani Dadaneh , Xiaoning Qian , Mingyuan Zhou

Microbial communities play important roles in the function and maintenance of various biosystems, ranging from human body to the environment. Current methods for analysis of microbial communities are typically based on taxonomic…

Genomics · Quantitative Biology 2015-12-02 Ehsaneddin Asgari , Kiavash Garakani , Mohammad R. K Mofrad

Photocount statistics are an important tool for the characterization of electromagnetic fields, especially for fields with an irrelevant phase. In the microwave domain, continuous rather than discrete measurements are the norm. Using a…

Quantum Physics · Physics 2016-04-13 Stéphane Virally , Jean Olivier Simoneau , Christian Lupien , Bertrand Reulet

Statistical data is often analyzed as a contingency table, sometimes with empty cells called zeros. Such sparse tables can be due to scarse observations classified in numerous categories, as for example in genetic association studies. Thus,…

Statistics Theory · Mathematics 2010-07-28 Audrey Finkler

Statistical data is often analyzed as a contingency table, sometimes with empty cells called zeros. Such sparse tables can be due to scarse observations classified in numerous categories, as for example in genetic association studies. Thus,…

Statistics Theory · Mathematics 2010-07-28 Audrey Finkler

Statistical matching is a technique for integrating two or more data sets when information available for matching records for individual participants across data sets is incomplete. Statistical matching can be viewed as a missing data…

Methodology · Statistics 2015-10-14 Jae-kwang Kim , Emily Berg , Taesung Park

Compositional data, also referred to as simplicial data, naturally arise in many scientific domains such as geochemistry, microbiology, and economics. In such domains, obtaining sensible lower-dimensional representations and modes of…

In a bivariate setting, we consider the problem of detecting a sparse contamination or mixture component, where the effect manifests itself as a positive dependence between the variables, which are otherwise independent in the main…

Statistics Theory · Mathematics 2020-01-13 Ery Arias-Castro , Rong Huang , Nicolas Verzelen

Questions of understanding and quantifying the representation and amount of information in organisms have become a central part of biological research, as they potentially hold the key to fundamental advances. In this paper, we demonstrate…

Genomics · Quantitative Biology 2007-10-30 H. M. Aktulga , I. Kontoyiannis , L. A. Lyznik , L. Szpankowski , A. Y. Grama , W. Szpankowski

Conformal prediction, which makes no distributional assumptions about the data, has emerged as a powerful and reliable approach to uncertainty quantification in practical applications. The nonconformity measure used in conformal prediction…

Machine Learning · Computer Science 2024-10-15 Yuko Kato , David M. J. Tax , Marco Loog

Scientists and engineers are often interested in learning the number of subpopulations (or components) present in a data set. A common suggestion is to use a finite mixture model (FMM) with a prior on the number of components. Past work has…

Statistics Theory · Mathematics 2021-07-08 Diana Cai , Trevor Campbell , Tamara Broderick