English
Related papers

Related papers: Variable Selection for Stratified Sampling Designs…

200 papers

An important task in survival analysis is choosing a structure for the relationship between covariates of interest and the time-to-event outcome. For example, the accelerated failure time (AFT) model structures each covariate effect as a…

Methodology · Statistics 2025-12-08 Harrison T. Reeder , Kyu Ha Lee , Sebastien Haneuse

The accelerated failure time (AFT) model is widely used to analyze relationships between variables in the presence of censored observations. However, this model relies on some assumptions such as the error distribution, which can lead to…

Methodology · Statistics 2026-02-10 Sangkon Oh , Hyunjae Lee , Sangwook Kang , Byungtae Seo

Randomized experiments are the gold standard for estimating the average treatment effect (ATE). While covariate adjustment can reduce the asymptotic variances of the unbiased Horvitz-Thompson estimators for the ATE, it suffers from…

Methodology · Statistics 2025-08-22 Xin Lu , Lei Shi , Hanzhong Liu , Peng Ding

In this paper, we present a generalized estimating equations based estimation approach and a variable selection procedure for single-index models when the observed data are clustered. Unlike the case of independent observations,…

Methodology · Statistics 2011-08-08 Peng Lai , Qihua Wang , Heng Lian

Given the limitations of using the Cox hazard ratio to summarize the magnitude of the treatment effect, alternative measures that do not have these limitations are gaining attention. One of the recently proposed alternative methods uses the…

Methodology · Statistics 2024-04-02 Zihan Qian , Lu Tian , Miki Horiguchi , Hajime Uno

Semiparametric accelerated failure time (AFT) models are a useful alternative to Cox proportional hazards models, especially when the assumption of constant hazard ratios is untenable. However, rank-based criteria for fitting AFT models are…

Methodology · Statistics 2022-01-20 Piotr M. Suder , Aaron J. Molstad

Frailty models are often the model of choice for heterogeneous survival data. A frailty model contains both random effects and fixed effects, with the random effects accommodating for the correlation in the data. Different estimation…

Methodology · Statistics 2019-09-17 Oodally Ajmal , Luc Duchateau , Estelle Kuhn

Approximate Bayesian computation (ABC) methods are standard tools for inferring parameters of complex models when the likelihood function is analytically intractable. A popular approach to improving the poor acceptance rate of the basic…

Methodology · Statistics 2025-01-27 Henri Pesonen , Jukka Corander

The paper addresses joint sparsity selection in the regression coefficient matrix and the error precision (inverse covariance) matrix for high-dimensional multivariate regression models in the Bayesian paradigm. The selected sparsity…

Methodology · Statistics 2022-01-19 Srijata Samanta , Kshitij Khare , George Michailidis

Many problems in the geophysical sciences demand the ability to calibrate the parameters and predict the time evolution of complex dynamical models using sequentially-collected data. Here we introduce a general methodology for the joint…

Computation · Statistics 2018-12-12 Sara Pérez-Vieites , Inés P. Mariño , Joaquín Míguez

We develop a constructive approach for $\ell_0$-penalized estimation in the sparse accelerated failure time (AFT) model with high-dimensional covariates. Our proposed method is based on Stute's weighted least squares criterion combined with…

Methodology · Statistics 2020-02-11 Xingdong Feng , Jian Huang , Yuling Jiao , Shuang Zhang

We study the application of the grouped fixed effects approach to binary choice models for panel data in presence of severe complete separation. Through data loss, complete separation may lead to biased estimates of Average Partial Effects…

Econometrics · Economics 2025-11-07 Claudia Pigini , Alessandro Pionati , Francesco Valentini

Feature selection is an important data preprocessing in data mining and machine learning which can be used to reduce the feature dimension without deteriorating model's performance. Since obtaining annotated data is laborious or even…

Machine Learning · Computer Science 2022-04-08 Zhenzhen Sun , Yuanlong Yu

In this paper, we consider the partially linear single-index models with longitudinal data. To deal with the variable selection problem in this context, we propose a penalized procedure combined with two bias correction methods, resulting…

Methodology · Statistics 2014-02-10 Gaorong Li , Peng Lai , Heng Lian

Clustering analysis is one of the most widely used statistical tools in many emerging areas such as microarray data analysis. For microarray and other high-dimensional data, the presence of many noise variables may mask underlying…

Machine Learning · Statistics 2008-03-26 Benhuai Xie , Wei Pan , Xiaotong Shen

This paper proposes a strategy for regularized estimation in multi-way contingency tables, which are common in meta-analyses and multi-center clinical trials. Our approach is based on data augmentation, and appeals heavily to a novel class…

Methodology · Statistics 2011-09-21 Nicholas G. Polson , James G. Scott

Longitudinal binary or count functional data are common in neuroscience, but are often too large to analyze with existing functional regression methods. We propose one-step penalized generalized estimating equations that supports…

Methodology · Statistics 2026-03-31 Gabriel Loewinger , Alex W. Levis , Erjia Cui , Francisco Pereira

This paper develops a unified framework for partial identification and inference in stratified experiments with attrition, accommodating both equal and heterogeneous treatment shares across strata. For equal-share designs, we apply recent…

Econometrics · Economics 2026-01-21 Bruno Ferman , Davi Siqueira , Vitor Possebom

For complex diseases, beyond the main effects of genetic (G) and environmental (E) factors, gene-environment (G-E) interactions also play an important role. Many of the existing G-E interaction methods conduct marginal analysis, which may…

Methodology · Statistics 2020-03-06 Qingzhao Zhang , Hao Chai , Shuangge Ma

Modeling correlated or highly stratified multiple-response data becomes a common data analysis task due to modern data monitoring facilities and methods. Generalized estimating equations (GEE) is one of the popular statistical methods for…

Methodology · Statistics 2019-03-05 Zimu Chen , Zhanfeng Wang , Yuan-chin Ivan Chang