中文
相关论文

相关论文: Multi-sample estimation of centered log-ratio matr…

200 篇论文

Many diseases are considered to be closely related to the changes in the gut microbial community, including colorectal cancer (CRC), which is one of the most common cancers in the world. The diagnostic classification and etiological…

定量方法 · 定量生物学 2023-10-19 Shan Tang , Shanjun Mao , Yangyang Chen , Falong Tan , Lihua Duan , Cong Pian , Xiangxiang Zeng

A number of studies have suggested using comparisons between DNA sequences of closely related bacterial isolates to estimate the relative rate of recombination to mutation for that bacterial species. We consider such an approach which uses…

应用统计 · 统计学 2015-06-03 Paul Fearnhead , Shoukai Yu , Patrick Biggs , Barbara Holland , Nigel French

Markov Chain Monte Carlo (MCMC) requires to evaluate the full data likelihood at different parameter values iteratively and is often computationally infeasible for large data sets. In this paper, we propose to approximate the log-likelihood…

统计方法学 · 统计学 2020-05-26 Guanyu Hu , HaiYing Wang

For robust statistical inference it is crucial to obtain a good estimator of the variance of the proposed estimator of the statistical estimand. A commonly used estimator of the variance for an asymptotically linear estimator is the sample…

统计方法学 · 统计学 2025-05-19 Yunwen Ji , Mark van der Laan , Alan Hubbard

High-dimensional compositional data arise naturally in many applications such as metagenomic data analysis. The observed data lie in a high-dimensional simplex, and conventional statistical methods often fail to produce sensible results due…

统计方法学 · 统计学 2016-01-19 Yuanpei Cao , Wei Lin , Hongzhe Li

We consider a problem of clustering a sequence of multinomial observations by way of a model selection criterion. We propose a form of a penalty term for the model selection procedure. Our approach subsumes both the conventional AIC and BIC…

机器学习 · 统计学 2015-08-17 Nam H. Lee , Runze Tang , Carey E. Priebe , Michael Rosen

We consider the problem of inferring an unknown number of clusters in replicated multinomial data. Under a model based clustering point of view, this task can be treated by estimating finite mixtures of multinomial distributions with or…

统计方法学 · 统计学 2023-07-07 Panagiotis Papastamoulis

Identifying the relevant variables for a classification model with correct confidence levels is a central but difficult task in high-dimension. Despite the core role of sparse logistic regression in statistics and machine learning, it still…

机器学习 · 统计学 2022-05-31 Binh T. Nguyen , Bertrand Thirion , Sylvain Arlot

Although there is no shortage of clustering algorithms proposed in the literature, the question of the most relevant strategy for clustering compositional data (i.e., data made up of profiles, whose rows belong to the simplex) remains…

统计理论 · 数学 2018-05-17 Antoine Godichon-Baggioni , Cathy Maugis-Rabusseau , Andrea Rau

Monitoring key elements of disease dynamics (e.g., prevalence, case counts) is of great importance in infectious disease prevention and control, as emphasized during the COVID-19 pandemic. To facilitate this effort, we propose a new…

统计方法学 · 统计学 2024-10-10 Lin Ge , Yuzi Zhang , Lance A. Waller , Robert H. Lyles

Genome wide comparisons between enteric bacteria yield large sets of conserved putative regulatory sites on a gene by gene basis that need to be clustered into regulons. Using the assumption that regulatory sites can be represented as…

生物物理 · 物理学 2009-11-07 Erik van Nimwegen , Mihaela Zavolan , Nikolaus Rajewsky , Eric D. Siggia

Structural matrix-variate observations routinely arise in diverse fields such as multi-layer network analysis and brain image clustering. While data of this type have been extensively investigated with fruitful outcomes being delivered, the…

统计理论 · 数学 2022-01-25 Zhongyuan Lyu , Dong Xia

High-throughput sequencing has transformed microbiome research, but it also produces inherently compositional data that challenge standard statistical and machine learning methods. In this work, we propose a multinomial classification…

应用统计 · 统计学 2025-12-23 R. Alberich , N. A. Cruz , R. Fernández , I. García Mosquera , A. Mir , F. Rosselló

Factor analysis, a classical multivariate statistical technique is popularly used as a fundamental tool for dimensionality reduction in statistics, econometrics and data science. Estimation is often carried out via the Maximum Likelihood…

最优化与控制 · 数学 2018-01-19 Koulik Khamaru , Rahul Mazumder

Large-scale multiple testing problems require the simultaneous assessment of many p-values. This paper compares several methods to assess the evidence in multiple binomial counts of p-values: the maximum of the binomial counts after…

统计方法学 · 统计学 2014-02-26 Guenther Walther

A learning classifier must outperform a trivial solution, in case of imbalanced data, this condition usually does not hold true. To overcome this problem, we propose a novel data level resampling method - Clustering Based Oversampling for…

机器学习 · 计算机科学 2018-11-13 Naman D. Singh , Abhinav Dhall

The imbalanced data classification is one of the most crucial tasks facing modern data analysis. Especially when combined with other difficulty factors, such as the presence of noise, overlapping class distributions, and small disjuncts,…

机器学习 · 计算机科学 2020-04-08 Michał Koziarski , Michał Woźniak , Bartosz Krawczyk

The RRUM is a model that is frequently seen in language assessment studies. The objective of this research is to advance an MCMC algorithm for the Bayesian RRUM. The algorithm starts with estimating correlated attributes. Using a saturated…

应用统计 · 统计学 2017-10-25 Meng-Ta Chung , Matthew S. Johnson

The objective of many high-dimensional microarray and RNA-seq studies is to develop a classifier of cancer patients based on characteristics of their disease. The germinal center B-cell (GCB) classifier study in lymphoma and the National…

应用统计 · 统计学 2015-09-17 Sandra Safo , Xiao Song , Kevin K. Dobbin

In ordinal classification, misclassifying neighboring ranks is common, yet the consequences of these errors are not the same. For example, misclassifying benign tumor categories is less consequential, compared to an error at the…

计算机视觉与模式识别 · 计算机科学 2026-01-13 Dileepa Pitawela , Gustavo Carneiro , Hsiang-Ting Chen