English
Related papers

Related papers: A New Covariate Selection Strategy for High Dimens…

200 papers

Inference for high-dimensional logistic regression models using penalized methods has been a challenging research problem. As an illustration, a major difficulty is the significant bias of the Lasso estimator, which limits its direct…

Methodology · Statistics 2024-10-29 Yuming Zhang , Stéphane Guerrier , Runze Li

Modern data often arises with multiple modalities. For example, covariates and a network are observed on the same subjects, and both contain useful information. Effectively integrating these modalities is important and challenging,…

Methodology · Statistics 2025-11-25 Tao Shen , Wanjie Wang

When drawing causal inferences about the effects of multiple treatments on clustered survival outcomes using observational data, we need to address implications of the multilevel data structure, multiple treatments, censoring and unmeasured…

Methodology · Statistics 2022-02-18 Liangyuan Hu , Jiayi Ji , Ronald D. Ennis , Joseph W. Hogan

Forward regression is a crucial methodology for automatically identifying important predictors from a large pool of potential covariates. In contexts with moderate predictor correlation, forward selection techniques can achieve screening…

Methodology · Statistics 2024-08-23 Xuejun Jiang , Yue Ma , Haofeng Wang

The sparse-group lasso performs both variable and group selection, simultaneously using the strengths of the lasso and group lasso. It has found widespread use in genetics, a field that regularly involves the analysis of high-dimensional…

Machine Learning · Statistics 2025-09-18 Fabio Feser , Marina Evangelou

Causal effect estimation for dynamic treatment regimes (DTRs) contributes to sequential decision making. However, censoring and time-dependent confounding under DTRs are challenging as the amount of observational data declines over time due…

Machine Learning · Statistics 2021-09-27 Adi Lin , Jie Lu , Junyu Xuan , Fujin Zhu , Guangquan Zhang

A common issue in learning decision-making policies in data-rich settings is spurious correlations in the offline dataset, which can be caused by hidden confounders. Instrumental variable (IV) regression, which utilises a key unconfounded…

Machine Learning · Computer Science 2025-06-25 Daqian Shao , Ashkan Soleymani , Francesco Quinzan , Marta Kwiatkowska

We propose generalized additive partial linear models for complex data which allow one to capture nonlinear patterns of some covariates, in the presence of linear components. The proposed method improves estimation efficiency and increases…

Statistics Theory · Mathematics 2014-05-26 Li Wang , Lan Xue , Annie Qu , Hua Liang

High-dimensional compositional data arise naturally in many applications such as metagenomic data analysis. The observed data lie in a high-dimensional simplex, and conventional statistical methods often fail to produce sensible results due…

Methodology · Statistics 2016-01-19 Yuanpei Cao , Wei Lin , Hongzhe Li

Sparse regression such as the Lasso has achieved great success in handling high-dimensional data. However, one of the biggest practical problems is that high-dimensional data often contain large amounts of missing values. Convex Conditioned…

Machine Learning · Statistics 2019-06-20 Masaaki Takada , Hironori Fujisawa , Takeichiro Nishikawa

In practical applications, one often does not know the "true" structure of the underlying conditional quantile function, especially in the ultra-high dimensional setting. To deal with ultra-high dimensionality, quantile-adaptive marginal…

Methodology · Statistics 2024-04-26 Daoji Li , Yinfei Kong , Dawit Zerom

Observational epidemiological studies commonly seek to estimate the causal effect of an exposure on an outcome. Adjustment for potential confounding bias in modern studies is challenging due to the presence of high-dimensional confounding,…

Methodology · Statistics 2025-08-29 Susan Ellul , Stijn Vansteelandt , John B. Carlin , Margarita Moreno-Betancur

Covariate-specific treatment effects (CSTEs) represent heterogeneous treatment effects across subpopulations defined by certain selected covariates. In this article, we consider marginal structural models where CSTEs are linearly…

Methodology · Statistics 2021-05-25 Peng Wu , Zhiqiang Tan , Wenjie Hu , Xiao-Hua Zhou

In this work, we developed a new Bayesian method for variable selection in function-on-scalar regression (FOSR). Our method uses a hierarchical Bayesian structure and latent variables to enable an adaptive covariate selection process for…

Methodology · Statistics 2026-03-31 Pedro Henrique T. O. Sousa , Camila P. E. de Souza , Ronaldo Dias

Consider the problem of estimating average treatment effects when a large number of covariates are used to adjust for possible confounding through outcome regression and propensity score models. The conventional approach of model building…

Statistics Theory · Mathematics 2018-01-31 Zhiqiang Tan

Adjusting for latent covariates is crucial for estimating causal effects from observational textual data. Most existing methods only account for confounding covariates that affect both treatment and outcome, potentially leading to biased…

Computation and Language · Computer Science 2023-11-27 Yuxiang Zhou , Yulan He

This paper investigates the high-dimensional linear regression with highly correlated covariates. In this setup, the traditional sparsity assumption on the regression coefficients often fails to hold, and consequently many model selection…

Methodology · Statistics 2019-03-26 Jianqing Fan , Bai Jiang , Qiang Sun

In this work, we consider causal inference in various high-dimensional treatment settings, including for single multi-valued treatments and vector treatments with binary or continuous components, when the number of treatments can be…

Statistics Theory · Mathematics 2026-02-26 Patrick Kramer , Edward H. Kennedy , Isaac M. Opper

High-dimensional classification and feature selection tasks are ubiquitous with the recent advancement in data acquisition technology. In several application areas such as biology, genomics and proteomics, the data are often functional in…

Machine Learning · Statistics 2021-09-30 W Yu , S Wade , H D Bondell , L Azizi

It is more and more frequently the case in applications that the data we observe come from one or more random variables taking values in an infinite dimensional space, e.g. curves. The need to have tools adapted to the nature of these data…

Statistics Theory · Mathematics 2023-06-01 Angelina Roche