中文
相关论文

相关论文: Penalized Quasi-likelihood for High-dimensional Lo…

200 篇论文

We introduce a new empirical Bayes approach for large-scale multiple linear regression. Our approach combines two key ideas: (i) the use of flexible "adaptive shrinkage" priors, which approximate the nonparametric family of scale mixture of…

统计方法学 · 统计学 2024-06-13 Youngseok Kim , Wei Wang , Peter Carbonetto , Matthew Stephens

We study model evaluation and model selection from the perspective of generalization ability (GA): the ability of a model to predict outcomes in new samples from the same population. We believe that GA is one way formally to address…

机器学习 · 统计学 2016-10-19 Ning Xu , Jian Hong , Timothy C. G. Fisher

Mixtures of matrix Gaussian distributions provide a probabilistic framework for clustering continuous matrix-variate data, which are becoming increasingly prevalent in various fields. Despite its widespread adoption and successful…

统计计算 · 统计学 2023-07-21 Andrea Cappozzo , Alessandro Casa , Michael Fop

Most existing unsupervised person re-identification (Re-ID) methods use clustering to generate pseudo labels for model training. Unfortunately, clustering sometimes mixes different true identities together or splits the same identity into…

计算机视觉与模式识别 · 计算机科学 2022-04-15 Xinyu Zhang , Dongdong Li , Zhigang Wang , Jian Wang , Errui Ding , Javen Qinfeng Shi , Zhaoxiang Zhang , Jingdong Wang

Gaussian processes (GPs) are popular as nonlinear regression models for expensive computer simulations, yet GP performance relies heavily on estimation of unknown covariance parameters. Maximum likelihood estimation (MLE) is common, but it…

统计方法学 · 统计学 2025-11-25 Ayumi Mutoh , Annie S. Booth , Jonathan W. Stallrich

We introduce resampled mutual information (ResMI), a novel measure of clustering similarity that combines insights from information theoretic and pair counting approaches to clustering and community detection. Similar to chance-corrected…

社会与信息网络 · 计算机科学 2024-12-06 Cheaheon Lim

Penalized variable selection for high dimensional longitudinal data has received much attention as accounting for the correlation among repeated measurements and providing additional and essential information for improved identification and…

统计方法学 · 统计学 2021-07-20 Fei Zhou , Xi Lu , Jie Ren , Kun Fan , Shuangge Ma , Cen Wu

Clustered multistate process data are commonly encountered in multicenter observational studies and clinical trials. A clinically important estimand with such data is the marginal probability of being in a particular transient state as a…

统计方法学 · 统计学 2022-09-05 Wenxian Zhou , Giorgos Bakoyannis , Ying Zhang , Constantin T Yiannoutsos

Clustering is considered a non-supervised learning setting, in which the goal is to partition a collection of data points into disjoint clusters. Often a bound $k$ on the number of clusters is given or assumed by the practitioner. Many…

机器学习 · 计算机科学 2012-02-01 Nir Ailon , Ron Begleiter

The clustering of bounded data presents unique challenges in statistical analysis due to the constraints imposed on the data values. This paper introduces a novel method for model-based clustering specifically designed for bounded data.…

统计方法学 · 统计学 2025-05-16 Luca Scrucca

For data with high-dimensional covariates but small to moderate sample sizes, the analysis of single datasets often generates unsatisfactory results. The integrative analysis of multiple independent datasets provides an effective way of…

统计方法学 · 统计学 2015-01-19 Yuan Huang , Qingzhao Zhang , Sanguo Zhang , Jian Huang , Shuangge Ma

Penalized generalized estimating equations with Elastic Net or L2-Smoothly Clipped Absolute Deviation penalization are proposed to simultaneously select the most important variables and estimate their effects for longitudinal Gaussian data…

统计方法学 · 统计学 2012-11-26 Adriaan Blommaert , Niel Hens , Philippe Beutels

Regression models applied to network data where node attributes are the dependent variables poses a methodological challenge. As has been well studied, naive regression neither properly accounts for community structure, nor does it account…

统计方法学 · 统计学 2024-02-16 Riddhi Pratim Ghosh , Jukka-Pekka Onnela , Ian Barnett

We study the large sample behavior of a convex clustering framework, which minimizes the sample within cluster sum of squares under an~$\ell_1$ fusion constraint on the cluster centroids. This recently proposed approach has been gaining in…

统计方法学 · 统计学 2016-12-30 Peter Radchenko , Gourab Mukherjee

Sparse regularized regression methods are now widely used in genome-wide association studies (GWAS) to address the multiple testing burden that limits discovery of potentially important predictors. Linear mixed models (LMMs) have become an…

统计方法学 · 统计学 2022-06-27 Julien St-Pierre , Karim Oualkacha , Sahir Rai Bhatnagar

We develop a penalized likelihood estimation framework to estimate the structure of Gaussian Bayesian networks from observational data. In contrast to recent methods which accelerate the learning problem by restricting the search space, our…

统计方法学 · 统计学 2015-12-24 Bryon Aragam , Qing Zhou

We propose a novel deep clustering method that integrates Variational Autoencoders (VAEs) into the Expectation-Maximization (EM) framework. Our approach models the probability distribution of each cluster with a VAE and alternates between…

机器学习 · 计算机科学 2025-01-14 Michael Adipoetra , Ségolène Martin

We study two practically important cases of model based clustering using Gaussian Mixture Models: (1) when there is misspecification and (2) on high dimensional data, in the light of recent advances in Gradient Descent (GD) based…

机器学习 · 统计学 2020-07-28 Siva Rajesh Kasa , Vaibhav Rajan

This paper proposes a penalized composite likelihood method for model selection in colored graphical Gaussian models. The method provides a sparse and symmetry-constrained estimator of the precision matrix, and thus conducts model selection…

统计方法学 · 统计学 2020-04-06 Qiong Li , Xiaoying Sun , Nanwei Wang

Computing expected information gain (EIG) from prior to posterior (equivalently, mutual information between candidate observations and model parameters or other quantities of interest) is a fundamental challenge in Bayesian optimal…

统计方法学 · 统计学 2026-01-30 Fengyi Li , Ricardo Baptista , Youssef Marzouk