中文
相关论文

相关论文: Debiased Lasso After Sample Splitting for Estimati…

200 篇论文

Bayesian statistical inference for Generalized Linear Models (GLMs) with parameters lying on a constrained space is of general interest (e.g., in monotonic or convex regression), but often constructing valid prior distributions supported on…

统计方法学 · 统计学 2021-09-02 Rahul Ghosal , Sujit K. Ghosh

This work performs a non-asymptotic analysis of the generalized Lasso under the assumption of sub-exponential data. Our main results continue recent research on the benchmark case of (sub-)Gaussian sample distributions and thereby explore…

统计理论 · 数学 2023-01-18 Martin Genzel , Christian Kipp

In this paper, we consider statistical inference with generalized linear models in high dimensions under a longitudinal clustered data framework. Specifically, we propose a de-sparsified version of an initial Dantzig-type regularized…

统计方法学 · 统计学 2025-08-13 Nathan Huey

We propose a unified framework to draw inferences for regression coefficients in a generalized linear model (GLM) following Lasso-based variable selection. We adapt to non-Gaussian GLMs a recently developed parametric programming strategy…

统计方法学 · 统计学 2026-03-27 Qinyan Shen , Karl Gregory , Xianzheng Huang

We propose a general method for distributed Bayesian model choice, using the marginal likelihood, where a data set is split in non-overlapping subsets. These subsets are only accessed locally by individual workers and no data is shared…

统计计算 · 统计学 2022-10-18 Alexander Buchholz , Daniel Ahfock , Sylvia Richardson

We propose a robust inferential procedure for assessing uncertainties of parameter estimation in high-dimensional linear models, where the dimension $p$ can grow exponentially fast with the sample size $n$. Our method combines the…

机器学习 · 统计学 2015-03-19 Tianqi Zhao , Mladen Kolar , Han Liu

In this article, we develop a distributed variable screening method for generalized linear models. This method is designed to handle situations where both the sample size and the number of covariates are large. Specifically, the proposed…

统计方法学 · 统计学 2024-05-09 Tianbo Diao , Lianqiang Qu , Bo Li , Liuquan Sun

Standard high-dimensional regression methods assume that the underlying coefficient vector is sparse. This might not be true in some cases, in particular in presence of hidden, confounding variables. Such hidden confounding can be…

统计方法学 · 统计学 2020-08-19 Domagoj Ćevid , Peter Bühlmann , Nicolai Meinshausen

Debiased machine learning is a meta algorithm based on bias correction and sample splitting to calculate confidence intervals for functionals, i.e. scalar summaries, of machine learning algorithms. For example, an analyst may desire the…

机器学习 · 统计学 2022-10-25 Victor Chernozhukov , Whitney K. Newey , Rahul Singh

Completely randomized experiment is the gold standard for causal inference. When the covariate information for each experimental candidate is available, one typical way is to include them in covariate adjustments for more accurate treatment…

统计方法学 · 统计学 2025-06-10 Xin Lu , Fan Yang , Yuhao Wang

As datasets grow larger, they are often distributed across multiple machines that compute in parallel and communicate with a central machine through short messages. In this paper, we focus on sparse regression and propose a new procedure…

统计方法学 · 统计学 2023-03-14 Sifan Liu , Snigdha Panigrahi

In high-dimensional sparse regression, the \textsc{Lasso} estimator offers excellent theoretical guarantees but is well-known to produce biased estimates. To address this, \cite{Javanmard2014} introduced a method to ``debias" the…

机器学习 · 统计学 2025-02-28 Shuvayan Banerjee , James Saunderson , Radhendushka Srivastava , Ajit Rajwade

Gaussian graphical regressions have emerged as a powerful approach for regressing the precision matrix of a Gaussian graphical model on covariates, which, unlike traditional Gaussian graphical models, can help determine how graphs are…

统计方法学 · 统计学 2025-01-17 Xuran Meng , Jingfei Zhang , Yi Li

Many machine learning algorithms are trained and evaluated by splitting data from a single source into training and test sets. While such focus on in-distribution learning scenarios has led to interesting advancement, it has not been able…

计算机视觉与模式识别 · 计算机科学 2020-07-02 Hyojin Bahng , Sanghyuk Chun , Sangdoo Yun , Jaegul Choo , Seong Joon Oh

Subsampling is a computationally efficient and scalable method to draw inference in large data settings based on a subset of the data rather than needing to consider the whole dataset. When employing subsampling techniques, a crucial…

统计方法学 · 统计学 2025-10-08 Amalan Mahendran , Helen Thompson , James M. McGree

We study high-dimensional two-sample mean comparison and address the curse of dimensionality through data-adaptive projections. Leveraging the low-dimensional and localized signal structures commonly seen in single-cell genomics data, our…

统计方法学 · 统计学 2025-06-12 Tianyu Zhang , Jing Lei , Kathryn Roeder

Valid uncertainty quantification after model selection remains challenging in high-dimensional linear regression, especially within the possibilistic inferential model (PIM) framework. We develop possibilistic inferential models for…

统计方法学 · 统计学 2025-12-23 Yaohui Lin

In this study, we investigate the bias and variance properties of the debiased Lasso in linear regression when the tuning parameter of the node-wise Lasso is selected to be smaller than in previous studies. We consider the case where the…

统计理论 · 数学 2022-08-19 Akira Shinkyu , Naoya Sueishi

In this paper, we propose an abstract procedure for debiasing constrained or regularized potentially high-dimensional linear models. It is elementary to show that the proposed procedure can produce $\frac{1}{\sqrt{n}}$-confidence intervals…

统计方法学 · 统计学 2023-01-12 Yufei Yi , Matey Neykov

Classifiers are biased when trained on biased datasets. As a remedy, we propose Learning to Split (ls), an algorithm for automatic bias detection. Given a dataset with input-label pairs, ls learns to split this dataset so that predictors…

机器学习 · 计算机科学 2022-07-22 Yujia Bao , Regina Barzilay