English
Related papers

Related papers: Model Selection for Unit-root Time Series with Man…

200 papers

In this article, we first establish the joint central limit theorem (CLT) for the extreme eigenvalues of the sample correlation matrix of high-dimensional random walks with cross-sectional dependence. We further investigate the asymptotic…

Methodology · Statistics 2025-08-05 Ruihan Liu , Chen Wang

We propose the holdout randomization test (HRT), an approach to feature selection using black box predictive models. The HRT is a specialized version of the conditional randomization test (CRT; Candes et al., 2018) that uses data splitting…

Methodology · Statistics 2021-03-23 Wesley Tansey , Victor Veitch , Haoran Zhang , Raul Rabadan , David M. Blei

In this paper, we set up the theoretical foundations for a high-dimensional functional factor model approach in the analysis of large cross-sections (panels) of functional time series (FTS). We first establish a representation result…

Statistics Theory · Mathematics 2021-04-14 Shahin Tavakoli , Gilles Nisol , Marc Hallin

Feature selection is frequently used as a pre-processing step to machine learning. It is a process of choosing a subset of original features so that the feature space is optimally reduced according to a certain evaluation criterion. The…

Computer Vision and Pattern Recognition · Computer Science 2014-01-07 Vijendra Singh , Shivani Pathak

Subsampled Randomized Hadamard Transform (SRHT), a popular random projection method that can efficiently project a $d$-dimensional data into $r$-dimensional space ($r \ll d$) in $O(dlog(d))$ time, has been widely used to address the…

Machine Learning · Computer Science 2020-10-07 Zijian Lei , Liang Lan

This paper examines LASSO, a widely-used $L_{1}$-penalized regression method, in high dimensional linear predictive regressions, particularly when the number of potential predictors exceeds the sample size and numerous unit root regressors…

Econometrics · Economics 2024-01-17 Ziwei Mei , Zhentao Shi

Time Series Classification (TSC) is essential in fields like medicine, environmental science, and finance, enabling tasks such as disease diagnosis, anomaly detection, and stock price analysis. While machine learning models like Recurrent…

Machine Learning · Computer Science 2024-06-25 Gonzalo Uribarri , Federico Barone , Alessio Ansuini , Erik Fransén

This paper studies model selection consistency for high dimensional sparse regression when data exhibits both cross-sectional and serial dependency. Most commonly-used model selection methods fail to consistently recover the true model when…

Methodology · Statistics 2018-09-12 Jianqing Fan , Yuan Ke , Kaizheng Wang

In the context of high-dimensional Gaussian linear regression for ordered variables, we study the variable selection procedure via the minimization of the penalized least-squares criterion. We focus on model selection where the penalty…

Statistics Theory · Mathematics 2024-07-01 Perrine Lacroix , Marie-Laure Martin

The problem of selecting a handful of truly relevant variables in supervised machine learning algorithms is a challenging problem in terms of untestable assumptions that must hold and unavailability of theoretical assurances that selection…

Methodology · Statistics 2023-11-10 Mehdi Rostami , Olli Saarela

Variable selection on the large-scale networks has been extensively studied in the literature. While most of the existing methods are limited to the local functionals especially the graph edges, this paper focuses on selecting the discrete…

Methodology · Statistics 2023-09-18 Lu Zhang , Junwei Lu

The Median of Medians (also known as BFPRT) algorithm, although a landmark theoretical achievement, is seldom used in practice because it and its variants are slower than simple approaches based on sampling. The main contribution of this…

Data Structures and Algorithms · Computer Science 2016-08-05 Andrei Alexandrescu

We propose a new variable selection procedure for a functional linear model with multiple scalar responses and multiple functional predictors. This method is based on basis expansions of the involved functional predictors and coefficients…

Statistics Theory · Mathematics 2023-11-03 Alban Mina Mbina , Guy Martial Nkiet

Several recent randomized linear algebra algorithms rely upon fast dimension reduction methods. A popular choice is the Subsampled Randomized Hadamard Transform (SRHT). In this article, we address the efficacy, in the Frobenius and spectral…

Data Structures and Algorithms · Computer Science 2015-03-20 Christos Boutsidis , Alex Gittens

We address the multiple testing problem under the assumption that the true/false hypotheses are driven by a Hidden Markov Model (HMM), which is recognized as a fundamental setting to model multiple testing under dependence since the seminal…

Methodology · Statistics 2021-05-04 Marie Perrot-Dockès , Gilles Blanchard , Pierre Neuvial , Etienne Roquain

Addressing the simultaneous identification of contributory variables while controlling the false discovery rate (FDR) in high-dimensional data is a crucial statistical challenge. In this paper, we propose a novel model-free variable…

Methodology · Statistics 2024-04-23 Yixin Han , Xu Guo , Changliang Zou

This paper presents a novel data-driven, direct filtering approach for unknown linear time-invariant systems affected by unknown-but-bounded measurement noise. The proposed technique combines independent multistep prediction models,…

Optimization and Control · Mathematics 2020-08-28 Marco Lauricella , Lorenzo Fagiano

Feature selection is important for modeling high-dimensional data, where the number of variables can be much larger than the sample size. In this paper, we develop a support detection and root finding procedure to learn the high dimensional…

Machine Learning · Statistics 2020-01-17 Jian Huang , Yuling Jiao , Lican Kang , Jin Liu , Yanyan Liu , Xiliang Lu

High-dimensional regression specification and analysis is a complex and active area of research in statistics, machine learning, and econometrics. This paper proposes a new approach, Boosting with Multiple Testing (BMT), which combines…

Econometrics · Economics 2026-02-24 George Kapetanios , Vasilis Sarafidis , Alexia Ventouri

There has been recent interest in extending the ideas of False Discovery Rates (FDR) to variable selection in regression settings. Traditionally the FDR in these settings has been defined in terms of the coefficients of the full regression…

Methodology · Statistics 2013-02-12 Max Grazier G'Sell , Trevor Hastie , Robert Tibshirani