Related papers: A central limit theorem for the Benjamini-Hochberg…
A stylized feature of high-dimensional data is that many variables have heavy tails, and robust statistical inference is critical for valid large-scale statistical inference. Yet, the existing developments such as Winsorization,…
Simultaneously performing variable selection and inference in high-dimensional models is an open challenge in statistics and machine learning. The increasing availability of vast amounts of variables requires the adoption of specific…
We provide an abstract multivariate central limit theorem with the Lindeberg-type error bounded in terms of Lipschitz functions (Wasserstein 1-distance) or functions with bounded second or third derivatives. The result is proved by means of…
Effectively controlling the false discovery rate (FDR) in high-dimensional variable selection is a fundamental statistical problem that has garnered significant research interest. In this paper, we propose a novel, user-friendly, and…
The likelihood ratio test (LRT) and the related $F$ test, do not (even asymptotically) adhere to their nominal $\chi^2$ and $F$ distributions in many statistical tests common in astrophysics, thereby casting many marginal line or source…
Statistical methods are frequently built upon assumptions that limit their applicability to certain problems and conditions. Failure to recognize these limitations can lead to conclusions that may be inaccurate or biased. An example of such…
We consider eigenvalues of generalized Wishart processes as well as particle systems, of which the empirical measures converge to deterministic measures as the dimension goes to infinity. In this paper, we obtain central limit theorems to…
In the literature, compact binary coalescences (CBCs) have been proposed as one of the main scenarios to explain the origin of some non-repeating fast radio bursts (FRBs). The large discrepancy between the FRB and CBC event rate densities…
Multivariate statistics are often available as well as necessary in hypothesis tests. We study how to use such statistics to control not only false discovery rate (FDR) but also positive FDR (pFDR) with good power. We show that FDR can be…
We consider large-scale studies in which thousands of significance tests are performed simultaneously. In some of these studies, the multiple testing procedure can be severely biased by latent confounding factors such as batch effects and…
Results on the false discovery rate (FDR) and the false nondiscovery rate (FNR) are developed for single-step multiple testing procedures. In addition to verifying desirable properties of FDR and FNR as measures of error rates, these…
The False Discovery Rate (FDR) is a new statistical procedure to control the number of mistakes made when performing multiple hypothesis tests, i.e. when comparing many data against a given model hypothesis. The key advantage of FDR is that…
Anomalous diffusion is an established phenomenon but still a theoretical challenge in non-equilibrium statistical mechanics. Physical models are built incrementally, and the most recent and most general family is based on the fractional…
This article considers the problem of multiple hypothesis testing using $t$-tests. The observed data are assumed to be independently generated conditional on an underlying and unknown two-state hidden model. We propose an asymptotically…
Belief Propagation (BP) is a powerful algorithm for distributed inference in probabilistic graphical models, however it quickly becomes infeasible for practical compute and memory budgets. Many efficient, non-parametric forms of BP have…
Controlling the false discovery rate (FDR) in variable selection becomes challenging when predictors are correlated, as existing methods often exclude all members of correlated groups and consequently perform poorly for prediction. We…
This paper designs methods for decentralized multiple hypothesis testing on graphs that are equipped with provable guarantees on the false discovery rate (FDR). We consider the setting where distinct agents reside on the nodes of an…
A block covariance structure is widely observed across large-scale and high-dimensional datasets in diverse fields such as biology, medicine, engineering, economics, and finance. This pattern entails partitioning a covariance matrix into…
We analyze control of the familywise error rate (FWER) in a multiple testing scenario with a great many null hypotheses about the distribution of a high-dimensional random variable among which only a very small fraction are false, or…
Empirical research in the social and medical sciences frequently involves testing multiple hypotheses simultaneously, increasing the risk of false positives due to chance. Classical multiple testing procedures, such as the Bonferroni…