中文
相关论文

相关论文: A multivariate variable selection approach for ana…

200 篇论文

In high-dimensions, many variable selection methods, such as the lasso, are often limited by excessive variability and rank deficiency of the sample covariance matrix. Covariance sparsity is a natural phenomenon in high-dimensional…

统计方法学 · 统计学 2010-06-08 X. Jessie Jeng And Z. John Daye

Variable selection in cluster analysis is important yet challenging. It can be achieved by regularization methods, which realize a trade-off between the clustering accuracy and the number of selected variables by using a lasso-type penalty.…

统计方法学 · 统计学 2016-12-23 Marbac Matthieu , Sedki Mohammed

To be fully useful for public health practice, models for epidemic response must be able to do more than predict -- it is also important to incorporate the mechanisms underlying transmission dynamics to enable policymakers and practitioners…

定量方法 · 定量生物学 2025-05-26 Jiale Tan , Marisa C. Eisenberg

Data science projects often involve various machine learning (ML) methods that depend on data, code, and models. One of the key activities in these projects is the selection of a model or algorithm that is appropriate for the data analysis…

机器学习 · 计算机科学 2023-11-27 Cristina Tavares , Nathalia Nascimento , Paulo Alencar , Donald Cowan

Alcohol consumption has been shown to influence cardiovascular mechanisms in humans, leading to observable alterations in the plasma metabolomic profile. Regression models are commonly employed to investigate these effects, treating…

统计方法学 · 统计学 2024-04-18 Yifan Yang , Chixiang Chen , Hwiyoung Lee , Ming Wang , Shuo Chen

Recommendation systems (RS) aim to provide personalized content, but they face a challenge in unbiased learning due to selection bias, where users only interact with items they prefer. This bias leads to a distorted representation of user…

机器学习 · 计算机科学 2025-06-10 Shuqiang Zhang , Yuchao Zhang , Jinkun Chen , Haochen Sui

For linear models that may have asymmetric errors, we study variable selection by cross-validation. The data are split into training and validation sets, with the number of observations in the validation set much larger than in the training…

统计方法学 · 统计学 2026-01-16 Bilel Bousselmi , Gabriela Ciuperca

Two-phase outcome dependent sampling (ODS) is widely used in many fields, especially when certain covariates are expensive and/or difficult to measure. For two-phase ODS, the conditional maximum likelihood (CML) method is very attractive…

统计方法学 · 统计学 2022-12-21 Menglu Che , Peisong Han , Jerald F. Lawless

This paper considers the problem of variable selection allowing for parameter instability. It distinguishes between signal and pseudo-signal variables that are correlated with the target variable, and noise variables that are not, and…

计量经济学 · 经济学 2024-07-17 Alexander Chudik , M. Hashem Pesaran , Mahrad Sharifvaghefi

Objective: Accurately classifying the malignancy of lesions detected in a screening scan is critical for reducing false positives. Radiomics holds great potential to differentiate malignant from benign tumors by extracting and analyzing a…

计算机视觉与模式识别 · 计算机科学 2019-02-14 Zhiguo Zhou , Shulong Li , Genggeng Qin , Michael Folkert , Steve Jiang , Jing Wang

Penalized regression models such as the Lasso have proved useful for variable selection in many fields - especially for situations with high-dimensional data where the numbers of predictors far exceeds the number of observations. These…

统计方法学 · 统计学 2014-03-19 Kasper Brink-Jensen , Claus Thorn Ekstrøm

Longitudinal data analysis is fundamental for understanding dynamic processes in biomedical and social sciences. Although varying coefficient models (VCMs) provide a flexible framework by allowing covariate effects to evolve over time,…

统计方法学 · 统计学 2026-03-10 Yu Lu , Tianni Zhang , Yuyao Wang , Mengfei Ran

Handling missing data in time series is a complex problem due to the presence of temporal dependence. General-purpose imputation methods, while widely used, often distort key statistical properties of the data, such as variance and…

统计方法学 · 统计学 2026-03-18 Guilherme Pumi , Taiane Schaedler Prass , Douglas Krauthein Verdum

Shrinkage estimators that possess the ability to produce sparse solutions have become increasingly important to the analysis of today's complex datasets. Examples include the LASSO, the Elastic-Net and their adaptive counterparts.…

统计方法学 · 统计学 2017-02-09 Hongmei Liu , J. Sunil Rao

In molecular biology, advances in high-throughput technologies have made it possible to study complex multivariate phenotypes and their simultaneous associations with high-dimensional genomic and other omics data, a problem that can be…

统计方法学 · 统计学 2021-12-02 Zhi Zhao , Marco Banterle , Leonardo Bottolo , Sylvia Richardson , Alex Lewin , Manuela Zucknick

Many problems within personalized medicine and digital health rely on the analysis of continuous-time functional biomarkers and other complex data structures emerging from high-resolution patient monitoring. In this context, this work…

机器学习 · 统计学 2025-01-14 Marcos Matabuena

We extend multi-way, multivariate ANOVA-type analysis to cases where one covariate is the view, with features of each view coming from different, high-dimensional domains. The different views are assumed to be connected by having paired…

机器学习 · 统计学 2009-12-17 Ilkka Huopaniemi , Tommi Suvitaival , Janne Nikkilä , Matej Orešič , Samuel Kaski

Heavy-tailed high-dimensional data are commonly encountered in various scientific fields and pose great challenges to modern statistical analysis. A natural procedure to address this problem is to use penalized quantile regression with…

统计理论 · 数学 2015-03-20 Jianqing Fan , Yingying Fan , Emre Barut

High-dimensional multivariate spatial-temporal data arise frequently in a wide range of applications; however, there are relatively few statistical methods that can simultaneously deal with spatial, temporal and variable-wise dependencies…

统计方法学 · 统计学 2020-02-05 Elynn Y. Chen , Xin Yun , Rong Chen , Qiwei Yao

We propose two approaches for selecting variables in latent class analysis (i.e.,mixture model assuming within component independence), which is the common model-based clustering method for mixed data. The first approach consists in…

统计计算 · 统计学 2017-03-08 Matthieu Marbac , Mohammed Sedki