English
Related papers

Related papers: Variable Selection for Stratified Sampling Designs…

200 papers

This paper studies covariate adjusted estimation of the average treatment effect in stratified experiments. We work in a general framework that includes matched tuples designs, coarse stratification, and complete randomization as special…

Econometrics · Economics 2024-07-23 Max Cytrynbaum

Variable selection naturally arises as a useful subject when faced with data with massive predictor space. In addition to the massive dimensionality, the data may be characterized by intra-subject correlation, and cure fraction, which are…

Methodology · Statistics 2025-12-24 Richard Tawiah , Shu Kay Ng , Geoffrey J. McLachlan

For many complex diseases, prognosis is of essential importance. It has been shown that, beyond the main effects of genetic (G) and environmental (E) risk factors, the gene-environment (G$\times$E) interactions also play a critical role. In…

Applications · Statistics 2015-05-15 Hao Chai , Qingzhao Zhang , Yu Jiang , Guohua Wang , Sanguo Zhang , Shuangge Ma

We consider the problem of selecting covariates in spatial linear models with Gaussian process errors. Penalized maximum likelihood estimation (PMLE) that enables simultaneous variable selection and parameter estimation is developed and,…

Methodology · Statistics 2012-02-24 Tingjin Chu , Jun Zhu , Haonan Wang

Penalized variable selection for high dimensional longitudinal data has received much attention as accounting for the correlation among repeated measurements and providing additional and essential information for improved identification and…

Methodology · Statistics 2021-07-20 Fei Zhou , Xi Lu , Jie Ren , Kun Fan , Shuangge Ma , Cen Wu

In biomedical studies it is of substantial interest to develop risk prediction scores using high-dimensional data such as gene expression data for clinical endpoints that are subject to censoring. In the presence of well-established…

Applications · Statistics 2011-11-24 Qi Long , Matthias Chung , Carlos S. Moreno , Brent A. Johnson

Semi-competing risks data arise when both non-terminal and terminal events are considered in a model. Such data with multiple events of interest are frequently encountered in medical research and clinical trials. In this framework, terminal…

Methodology · Statistics 2022-11-21 Fatemeh Mahmoudi , Xuewen Lu

In multi-state models based on high-dimensional data, effective modeling strategies are required to determine an optimal, ideally parsimonious model. In particular, linking covariate effects across transitions is needed to conduct joint…

Methodology · Statistics 2024-11-27 Kaya Miah , Jelle J. Goeman , Hein Putter , Annette Kopp-Schneider , Axel Benner

We introduce a novel method to simultaneously perform variable selection and estimation in the joint frailty model of recurrent and terminal events using the Broken Adaptive Ridge Regression penalty. The BAR penalty can be summarized as an…

Methodology · Statistics 2024-09-04 Christian Chan , Fatemeh Mahmoudi , Chel Hee Lee , Quan Long , Xuewen Lu

For data with high-dimensional covariates but small to moderate sample sizes, the analysis of single datasets often generates unsatisfactory results. The integrative analysis of multiple independent datasets provides an effective way of…

Methodology · Statistics 2015-01-19 Yuan Huang , Qingzhao Zhang , Sanguo Zhang , Jian Huang , Shuangge Ma

A novel approach for dealing with censored competing risks regression data is proposed. This is implemented by a mixture of accelerated failure time (AFT) models for a competing risks scenario within a cluster-weighted modelling (CWM)…

Methodology · Statistics 2013-12-04 Utkarsh J. Dang , Paul D. McNicholas

Estimating the generalization error (GE) of machine learning models is fundamental, with resampling methods being the most common approach. However, in non-standard settings, particularly those where observations are not independently and…

Simulation schemes for probabilistic inference in Bayesian belief networks offer many advantages over exact algorithms; for example, these schemes have a linear and thus predictable runtime while exact algorithms have exponential runtime.…

Artificial Intelligence · Computer Science 2013-02-28 Remco R. Bouckaert

Standard penalized methods of variable selection and parameter estimation rely on the magnitude of coefficient estimates to decide which variables to include in the final model. However, coefficient estimates are unreliable when the design…

Methodology · Statistics 2018-02-13 Jonathan P Williams , Jan Hannig

The accelerated failure time (AFT) model is a commonly used tool in analyzing survival data. In public health studies, data is often collected from medical service providers in different locations. Survival rates from different locations…

Applications · Statistics 2020-02-11 Guanyu Hu , Yishu Xue , Fred Huffer

Controlling the false discovery rate (FDR) in variable selection becomes challenging when predictors are correlated, as existing methods often exclude all members of correlated groups and consequently perform poorly for prediction. We…

Methodology · Statistics 2026-03-03 Sarah Organ , Toby Kenney , Hong Gu

Fault localization is to identify faulty source code. It could be done on various granularities, e.g., classes, methods, and statements. Most of the automated fault localization (AFL) approaches are coarse-grained because it is challenging…

Software Engineering · Computer Science 2021-07-21 Leping Li , Hui Liu

The semiparametric accelerated failure time (AFT) model offers a direct and interpretable alternative to the Cox proportional hazards model, yet practical diagnostic tools for this framework remain limited. We introduce afttest, an R…

Computation · Statistics 2026-03-09 Woojung Bae , Dongrak Choi , Jun Yan , Sangwook Kang

Automated feature engineering (AFE) enables AI systems to autonomously construct high-utility representations from raw tabular data. However, existing AFE methods rely on statistical heuristics, yielding brittle features that fail under…

Artificial Intelligence · Computer Science 2026-02-19 Arun Vignesh Malarkkan , Wangyang Ying , Yanjie Fu

Inference for models with recursively defined likelihoods is computationally demanding, limiting scalability to large datasets. We propose a stabilised weighted subsampling methodology for accelerated inference based on an unbiased…

Methodology · Statistics 2026-05-14 Matias Quiroz , Aishwarya Bhaskaran , Zixuan Wang , Thomas Goodwin