Related papers: The Spurious Factor Dilemma: Robust Inference in H…
The accurate specification of the number of factors is critical to the validity of factor models and the topic almost occupies the central position in factor analysis. Plenty of estimators are available under the restrictive condition that…
Determining the number of factors in high-dimensional factor modeling is essential but challenging, especially when the data are heavy-tailed. In this paper, we introduce a new estimator based on the spectral properties of Spearman sample…
In high-dimensional statistics, the Lasso is a cornerstone method for simultaneous variable selection and parameter estimation. However, its reliance on the squared loss function renders it highly sensitive to outliers and heavy-tailed…
Factor models are a very efficient way to describe high dimensional vectors of data in terms of a small number of common relevant factors. This problem, which is of fundamental importance in many disciplines, is usually reformulated in…
Recent works have proposed incorporating heavy-tailed (HT) noise into diffusion- and flow-based generative models, with the goals of better recovering the tails of target distributions and improving generative diversity. This motivation is…
Estimating eigenvectors and low-dimensional subspaces is of central importance for numerous problems in statistics, computer science, and applied mathematics. This paper characterizes the behavior of perturbed eigenvectors for a range of…
In this paper, we develop connections between two seemingly disparate, but central, models in robust statistics: Huber's epsilon-contamination model and the heavy-tailed noise model. We provide conditions under which this connection…
In this paper, we investigate and develop a new approach to the numerical analysis and characterization of random fluctuations with heavy-tailed probability distribution function (PDF), such as turbulent heat flow and solar flare…
Though very popular, it is well known that the EM for GMM algorithm suffers from non-Gaussian distribution shapes, outliers and high-dimensionality. In this paper, we design a new robust clustering algorithm that can efficiently deal with…
Large-scale multiple testing with correlated and heavy-tailed data arises in a wide range of research areas from genomics, medical imaging to finance. Conventional methods for estimating the false discovery proportion (FDP) often ignore the…
Feature selection has remained a daunting challenge in machine learning and artificial intelligence, where increasingly complex, high-dimensional datasets demand principled strategies for isolating the most informative predictors. Despite…
Ecess-over-Threshold method is a crucial technique in extreme value analysis, which approximately models larger observations over a threshold using a Generalized Pareto Distribution. This paper presents a comprehensive framework for…
We study the problem of factor modelling vector- and tensor-valued time series in the presence of heavy tails in the data, which produce extreme observations with non-negligible probability. We propose to combine a two-step procedure for…
Factor analysis is over a century old, but it is still problematic to choose the number of factors for a given data set. The scree test is popular but subjective. The best performing objective methods are recommended on the basis of…
We address a parametric joint detection-estimation problem for discrete signals of the form $x(t) = \sum_{n=1}^{N} \alpha_n e^{-i \lambda_n t } + \epsilon_t$, $t \in \mathbb{N}$, with an additive noise represented by independent centered…
Identifying the number of factors in a high-dimensional factor model has attracted much attention in recent years and a general solution to the problem is still lacking. A promising ratio estimator based on the singular values of the lagged…
This paper proposes new estimators of the number of factors for a generalised factor model with more relaxed assumptions than the strict factor model. Under the framework of large cross-sections $N$ and large time dimensions $T$, we first…
We study large deviation probabilities for a sum of dependent random variables from a heavy-tailed factor model, assuming that the components are regularly varying. We identify conditions where both the factor and the idiosyncratic terms…
Despite the successes of probabilistic models based on passing noise through neural networks, recent work has identified that such methods often fail to capture tail behavior accurately, unless the tails of the base distribution are…
We analyze quantitatively the effect of spurious multifractality induced by the presence of fat-tailed symmetric and asymmetric probability distributions of fluctuations in time series. In the presented approach different kinds of symmetric…