English
Related papers

Related papers: Time-to-event prediction for grouped variables usi…

200 papers

Interval-censored data analysis is important in biomedical statistics for any type of time-to-event response where the time of response is not known exactly, but rather only known to occur between two assessment times. Many clinical trials…

Methodology · Statistics 2019-06-12 Weichi Yao , Halina Frydman , Jeffrey S. Simonoff

Machine learning models that aim to predict dementia onset usually follow the classification methodology ignoring the time until an event happens. This study presents an alternative, using survival analysis within the context of machine…

Machine Learning · Computer Science 2023-06-21 Daniel Stamate , Henry Musto , Olesya Ajnakina , Daniel Stahl

High throughput genetic sequencing arrays with thousands of measurements per sample and a great amount of related censored clinical data have increased demanding need for better measurement specific model selection. In this paper we…

Statistics Theory · Mathematics 2019-07-31 Jelena Bradic , Jianqing Fan , Jiancheng Jiang

It is well known that in a supervised classification setting when the number of features is smaller than the number of observations, Fisher's linear discriminant rule is asymptotically Bayes. However, there are numerous modern applications…

Machine Learning · Statistics 2014-09-17 Irina Gaynanova , James G. Booth , Martin T. Wells

In this paper, we develop a novel high-dimensional coefficient estimation procedure based on high-frequency data. Unlike usual high-dimensional regression procedures such as LASSO, we additionally handle the heavy-tailedness of…

Methodology · Statistics 2025-10-22 Minseok Shin , Donggyu Kim

Grouping structures arise naturally in many statistical modeling problems. Several methods have been proposed for variable selection that respect grouping structure in variables. Examples include the group LASSO and several concave group…

Statistics Theory · Mathematics 2013-01-07 Jian Huang , Patrick Breheny , Shuangge Ma

We consider the problem of estimating multiple related but distinct graphical models on the basis of a high-dimensional data set with observations that belong to distinct classes. A motivating example occurs in the analysis of gene…

Methodology · Statistics 2012-07-12 Patrick Danaher , Pei Wang , Daniela M. Witten

Network meta-analysis (NMA) allows the combination of direct and indirect evidence from a set of randomized clinical trials. Performing NMA using individual patient data (IPD) is considered as a "gold standard" approach as it provides…

Methodology · Statistics 2021-10-22 Edouard Ollier , Pierre Blanchard , Gwénaël Le Teuff , Stefan Michiels

Penalized regression is an attractive framework for variable selection problems. Often, variables possess a grouping structure, and the relevant selection problem is that of selecting groups, not individual variables. The group lasso has…

Computation · Statistics 2016-07-20 Patrick Breheny , Jian Huang

We propose a computationally intensive method, the random lasso method, for variable selection in linear models. The method consists of two major steps. In step 1, the lasso method is applied to many bootstrap samples, each using a set of…

Applications · Statistics 2011-04-19 Sijian Wang , Bin Nan , Saharon Rosset , Ji Zhu

Accurately selecting and estimating smooth functional effects in additive models with potentially many functions is a challenging task. We introduce a novel Demmler-Reinsch basis expansion to model the functional effects that allows us to…

Methodology · Statistics 2024-01-02 Paul Bach , Nadja Klein

One of the central goals in precision health is the understanding and interpretation of high-dimensional biological data to identify genes and markers associated with disease initiation, development, and outcomes. Though significant effort…

Quantitative Methods · Quantitative Biology 2020-09-18 Zhi Huang , Paul Salama , Wei Shao , Jie Zhang , Kun Huang

An approximate method for conducting resampling in Lasso, the $\ell_1$ penalized linear regression, in a semi-analytic manner is developed, whereby the average over the resampled datasets is directly computed without repeated numerical…

Machine Learning · Statistics 2018-12-11 Tomoyuki Obuchi , Yoshiyuki Kabashima

Modern technologies are producing a wealth of data with complex structures. For instance, in two-dimensional digital imaging, flow cytometry, and electroencephalography, matrix type covariates frequently arise when measurements are obtained…

Methodology · Statistics 2013-10-22 Hua Zhou , Lexin Li

Penalized (or regularized) regression, as represented by Lasso and its variants, has become a standard technique for analyzing high-dimensional data when the number of variables substantially exceeds the sample size. The performance of…

Methodology · Statistics 2019-08-13 Yunan Wu , Lan Wang

The lasso is a popular tool for sparse linear regression, especially for problems in which the number of variables p exceeds the number of observations n. But when p>n, the lasso criterion is not strictly convex, and hence it may not have a…

Statistics Theory · Mathematics 2012-11-06 Ryan J. Tibshirani

We provide a principled way for investigators to analyze randomized experiments when the number of covariates is large. Investigators often use linear multivariate regression to analyze randomized experiments instead of simply reporting the…

Statistics Theory · Mathematics 2022-06-08 Adam Bloniarz , Hanzhong Liu , Cun-Hui Zhang , Jasjeet Sekhon , Bin Yu

We consider panel data models where coefficients change smoothly over time and follow a latent group structure, being homogeneous within but heterogeneous across groups. To jointly estimate the group membership and group-specific…

Econometrics · Economics 2025-11-19 Paul Haimerl , Stephan Smeekes , Ines Wilms

Among the most popular variable selection procedures in high-dimensional regression, Lasso provides a solution path to rank the variables and determines a cut-off position on the path to select variables and estimate coefficients. In this…

Methodology · Statistics 2018-06-19 X. Jessie Jeng , Huimin Peng , Wenbin Lu

Frailty models are often the model of choice for heterogeneous survival data. A frailty model contains both random effects and fixed effects, with the random effects accommodating for the correlation in the data. Different estimation…

Methodology · Statistics 2019-09-17 Oodally Ajmal , Luc Duchateau , Estelle Kuhn