中文
相关论文

相关论文: Two-sample testing in non-sparse high-dimensional …

200 篇论文

In this paper, we focus our attention on the high-dimensional double sparse linear regression, that is, a combination of element-wise and group-wise sparsity. To address this problem, we propose an IHT-style (iterative hard thresholding)…

统计理论 · 数学 2024-12-10 Yanhang Zhang , Zhifan Li , Shixiang Liu , Jianxin Yin

This paper uses techniques from Random Matrix Theory to find the ideal training-testing data split for a simple linear regression with m data points, each an independent n-dimensional multivariate Gaussian. It defines "ideal" as satisfying…

机器学习 · 统计学 2022-07-26 Alexander Dubbs

Simultaneous variable selection and statistical inference is challenging in high-dimensional data analysis. Most existing post-selection inference methods require explicitly specified regression models, which are often linear, as well as…

统计方法学 · 统计学 2026-03-19 Shangyuan Ye , Shauna Rakshe , Ye Liang

High-dimensional time series appear in many scientific setups, demanding a nuanced approach to model and analyze the underlying dependence structure. Theoretical advancements so far often rely on stringent assumptions regarding the sparsity…

信息论 · 计算机科学 2025-03-20 Daria Tieplova , Samriddha Lahiry , Jean Barbier

This paper studies a tensor-structured linear regression model with a scalar response variable and tensor-structured predictors, such that the regression parameters form a tensor of order $d$ (i.e., a $d$-fold multiway array) in…

机器学习 · 计算机科学 2020-11-26 Talal Ahmed , Haroon Raja , Waheed U. Bajwa

Fitting high-dimensional statistical models often requires the use of non-linear parameter estimation procedures. As a consequence, it is generally impossible to obtain an exact characterization of the probability distribution of the…

统计方法学 · 统计学 2014-04-03 Adel Javanmard , Andrea Montanari

This paper considers the problem of testing temporal homogeneity of $p$-dimensional population mean vectors from the repeated measurements of $n$ subjects over $T$ times. To cope with the challenges brought by high-dimensional longitudinal…

统计方法学 · 统计学 2016-08-29 Ping-Shou Zhong , Jun Li

We formulate nonparametric and semiparametric hypothesis testing of multivariate stationary linear time series in a unified fashion and propose new test statistics based on estimators of the spectral density matrix. The limiting…

统计理论 · 数学 2009-09-03 Yoshihiro Yajima , Yasumasa Matsuda

Deep models trained with noisy labels are prone to over-fitting and struggle in generalization. Most existing solutions are based on an ideal assumption that the label noise is class-conditional, i.e., instances of the same class share the…

计算机视觉与模式识别 · 计算机科学 2022-08-01 Ganlong Zhao , Guanbin Li , Yipeng Qin , Feng Liu , Yizhou Yu

This paper studies model checking for general parametric regression models having no dimension reduction structures on the predictor vector. Using any U-statistic type test as an initial test, this paper combines the sample-splitting and…

统计方法学 · 统计学 2023-08-21 Feng Liang , Chuhan Wang , jiaqi Huang , Lixing Zhu

In this article, we propose some two-sample tests based on ball divergence and investigate their high dimensional behavior. First, we study their behavior for High Dimension, Low Sample Size (HDLSS) data, and under appropriate regularity…

统计理论 · 数学 2024-10-08 Bilol Banerjee , Anil K. Ghosh

A significant hurdle for analyzing large sample data is the lack of effective statistical computing and inference methods. An emerging powerful approach for analyzing large sample data is subsampling, by which one takes a random subsample…

统计方法学 · 统计学 2015-11-24 Rong Zhu , Ping Ma , Michael W. Mahoney , Bin Yu

The ability to predict individualized treatment effects (ITEs) based on a given patient's profile is essential for personalized medicine. We propose a hypothesis testing approach to choosing between two potential treatments for a given…

统计方法学 · 统计学 2020-08-11 Tianxi Cai , Tony Cai , Zijian Guo

This paper presents a novel method to make statistical inferences for both the model support and regression coefficients in a high-dimensional logistic regression model. Our method is based on the repro samples framework, in which we…

统计方法学 · 统计学 2024-03-18 Xiaotian Hou , Linjun Zhang , Peng Wang , Min-ge Xie

We consider the problem of model selection and estimation in sparse high dimensional linear regression models with strongly correlated variables. First, we study the theoretical properties of the dual Lasso solution, and we show that joint…

应用统计 · 统计学 2017-03-21 Niharika Gauraha

Neoteric works have shown that modern deep learning models can exhibit a sparse double descent phenomenon. Indeed, as the sparsity of the model increases, the test performance first worsens since the model is overfitting the training data;…

机器学习 · 计算机科学 2024-02-09 Victor Quétu , Enzo Tartaglione

This paper considers the problem of testing whether there exists a solution satisfying certain non-negativity constraints to a linear system of equations. Importantly and in contrast to some prior work, we allow all parameters in the system…

Sliced inverse regression (SIR) is a popular sufficient dimension reduction method that identifies a few linear transformations of the covariates without losing regression information with the response. In high-dimensional settings, SIR can…

统计方法学 · 统计学 2025-12-04 Linh H. Nghiem , Francis. K. C. Hui , Samuel Muller , A. H. Welsh

High dimensional Vector Autoregressions (VAR) have received a lot of interest recently due to novel applications in health, engineering, finance and the social sciences. Three issues arise when analyzing VAR's: (a) The high dimensional…

统计理论 · 数学 2022-11-15 Sagnik Halder , George Michailidis

In this paper, we consider the classic measurement error regression scenario in which our independent, or design, variables are observed with several sources of additive noise. We will show that our motivating example's replicated…

应用统计 · 统计学 2012-07-10 David J. Biagioni , Ryan Elmore , Wesley Jones