English
Related papers

Related papers: Inference for Large Panel Data with Many Covariate…

200 papers

We introduce a new approach to variable selection, called Predictive Correlation Screening, for predictor design. Predictive Correlation Screening (PCS) implements false positive control on the selected variables, is well suited to small…

Machine Learning · Statistics 2013-04-11 Hamed Firouzi , Bala Rajaratnam , Alfred Hero

In high-dimensional data analysis, bi-level sparsity is often assumed when covariates function group-wisely and sparsity can appear either at the group level or within certain groups. In such cases, an ideal model should be able to…

Methodology · Statistics 2021-09-14 Bin Luo , Xiaoli Gao

Choosing relevant predictors is central to the analysis of biomedical time-to-event data. Classical frequentist inference, however, presumes that the set of covariates is fixed in advance and does not account for data-driven variable…

Methodology · Statistics 2026-02-10 Lena Schemet , Sarah Friedrich-Welz

Prediction-powered inference (PPI) is a recent framework for valid statistical inference with partially labeled data, combining model-based predictions on a large unlabeled set with bias correction from a smaller labeled subset. Building on…

Machine Learning · Statistics 2026-03-25 Jyotishka Datta , Nicholas G. Polson

Ultra-high dimensional longitudinal data are increasingly common and the analysis is challenging both theoretically and methodologically. We offer a new automatic procedure for finding a sparse semivarying coefficient model, which is widely…

Methodology · Statistics 2014-09-24 Ming-Yen Cheng , Toshio Honda , Jialiang Li , Heng Peng

As datasets grow larger, they are often distributed across multiple machines that compute in parallel and communicate with a central machine through short messages. In this paper, we focus on sparse regression and propose a new procedure…

Methodology · Statistics 2023-03-14 Sifan Liu , Snigdha Panigrahi

This work proposes new inference methods for a regression coefficient of interest in a (heterogeneous) quantile regression model. We consider a high-dimensional model where the number of regressors potentially exceeds the sample size but a…

Statistics Theory · Mathematics 2017-10-05 Alexandre Belloni , Victor Chernozhukov , Kengo Kato

Confounding remains one of the major challenges to causal inference with observational data. This problem is paramount in medicine, where we would like to answer causal questions from large observational datasets like electronic health…

Methodology · Statistics 2024-01-10 Linying Zhang , Yixin Wang , Martijn Schuemie , David Blei , George Hripcsak

Hierarchical inference in (generalized) regression problems is powerful for finding significant groups or even single covariates, especially in high-dimensional settings where identifiability of the entire regression parameter vector may be…

Methodology · Statistics 2021-10-22 Claude Renaux , Peter Bühlmann

While most of the convergence results in the literature on high dimensional covariance matrix are concerned about the accuracy of estimating the covariance matrix (and precision matrix), relatively less is known about the effect of…

Statistics Theory · Mathematics 2013-11-13 Jushan Bai , Yuan Liao

The model interpretation is essential in many application scenarios and to build a classification model with a ease of model interpretation may provide useful information for further studies and improvement. It is common to encounter with a…

Machine Learning · Statistics 2019-01-07 Wan-Ping Nicole Chen , Yuan-chin Ivan Chang

A new statistical procedure, based on a modified spline basis, is proposed to identify the linear components in the panel data model with fixed effects. Under some mild assumptions, the proposed procedure is shown to consistently estimate…

Econometrics · Economics 2019-11-21 Ruiqi Liu , Ben Boukai , Zuofeng Shang

We consider the problem of testing for differences in group-specific slopes between the selected groups in panel data identified via k-means clustering. In this setting, the classical Wald-type test statistic is problematic because it…

Methodology · Statistics 2025-11-07 Chuang Wan , Jiajun Sun , Xingbai Xu

Covariate-adaptive randomization is widely employed to balance baseline covariates in interventional studies such as clinical trials and experiments in development economics. Recent years have witnessed substantial progress in inference…

Methodology · Statistics 2024-05-30 Jiahui Xin , Hanzhong Liu , Wei Ma

This paper presents Sparse Partitioning, a Bayesian method for identifying predictors that either individually or in combination with others affect a response variable. The method is designed for regression problems involving binary or…

Quantitative Methods · Quantitative Biology 2011-08-31 Doug Speed , Simon Tavaré

We develop tools to do valid post-selective inference for a family of model selection procedures, including choosing a model via cross-validated Lasso. The tools apply universally when the following random vectors are jointly asymptotically…

Methodology · Statistics 2018-02-13 Jelena Markovic , Lucy Xia , Jonathan Taylor

Models with dimension more than the available sample size are now commonly used in various applications. A sensible inference is possible using a lower-dimensional structure. In regression problems with a large number of predictors, the…

Statistics Theory · Mathematics 2025-11-25 Sayantan Banerjee , Ismaël Castillo , Subhashis Ghosal

This paper presents robust inference methods for general linear hypotheses in linear panel data models with latent group structure in the coefficients. We employ a selective conditional inference approach, deriving the conditional…

Econometrics · Economics 2025-11-25 Oguzhan Akgun , Ryo Okui

Gene expression and phenotype association can be affected by potential unmeasured confounders from multiple sources, leading to biased estimates of the associations. Since genetic variants largely explain gene expression variations, they…

Methodology · Statistics 2019-10-23 Jiarui Lu , Hongzhe Li

In genetic studies, not only can the number of predictors obtained from microarray measurements be extremely large, there can also be multiple response variables. Motivated by such a situation, we consider semiparametric dimension reduction…

Methodology · Statistics 2013-09-25 Heng Lian , Shujie Ma