Related papers: On Identity Tests for High Dimensional Data Using …
With large language models (LLMs) like GPT-4 appearing to behave increasingly human-like in text-based interactions, it has become popular to attempt to evaluate personality traits of LLMs using questionnaires originally developed for…
We develop a class of tests for time series models such as multiple regression with growing dimension, infinite-order autoregression and nonparametric sieve regression. Examples include the Chow test and general linear restriction tests of…
We study sample covariance matrices arising from multi-level components of variance. Thus, let $ B_n=\frac{1}{N}\sum_{j=1}^NT_{j}^{1/2}x_jx_j^TT_{j}^{1/2}$, where $x_j\in R^n$ are i.i.d. standard Gaussian, and…
We consider large non-Hermitian random matrices $X$ with complex, independent, identically distributed centred entries and show that the linear statistics of their eigenvalues are asymptotically Gaussian for test functions having…
Originating from a system theory and an input/output point of view, I introduce a new class of generalized distributions. A parametric nonlinear transformation converts a random variable $X$ into a so-called Lambert $W$ random variable $Y$,…
For a high-dimensional parameter of interest, tests based on quadratic statistics are known to have low power against subsets of the parameter space (henceforth, parameter subspaces). In addition, they typically involve an inverse…
Every student in statistics or data science learns early on that when the sample size largely exceeds the number of variables, fitting a logistic model produces estimates that are approximately unbiased. Every student also learns that there…
Large language models (LLMs) have shown strong performance on clinical de-identification, the task of identifying sensitive identifiers to protect privacy. However, previous work has not examined their generalizability between formats,…
Large language models (LLMs) are increasingly used in statistical research and applications. However,they are also notorious for unreliable or biased information. Here, we explore whether LLMs can be used to improve the precision of…
Researchers in the behavioral and social sciences use linear discriminant analysis (LDA) for predictions of group membership (classification) and for identifying the variables most relevant to group separation among a set of continuous…
The recent success of generative adversarial networks and variational learning suggests training a classifier network may work well in addressing the classical two-sample problem. Network-based tests have the computational advantage that…
Most of the existing methods for estimating the local intrinsic dimension of a data distribution do not scale well to high-dimensional data. Many of them rely on a non-parametric nearest neighbors approach which suffers from the curse of…
A fundamental assumption underling any Hypothesis Testing (HT) problem is that the available data follow the parametric model assumed to derive the test statistic. Nevertheless, a perfect match between the true and the assumed data models…
The test of independence is a crucial component of modern data analysis. However, traditional methods often struggle with the complex dependency structures found in high-dimensional data. To overcome this challenge, we introduce a novel…
High-dimensional tests are applied to find relevant sets of variables and relevant models. If variables are selected by analyzing the sums of products matrices and a corresponding mean-value test is performed, there is the danger that the…
For general repeated measures designs the Wald-type statistic (WTS) is an asymptotically valid procedure allowing for unequal covariance matrices and possibly non-normal multivariate observations. The drawback of this procedure is the poor…
This paper presents a selective survey of recent developments in statistical inference and multiple testing for high-dimensional regression models, including linear and logistic regression. We examine the construction of confidence…
We investigate the likelihood ratio test for a large block-diagonal covariance matrix with an increasing number of blocks under the null hypothesis. While so far the likelihood ratio statistic has only been studied for normal populations,…
A fundamental tenet of pattern recognition is that overlap between training and testing sets causes an optimistic accuracy estimate. Deep CNNs for face recognition are trained for N-way classification of the identities in the training set.…
Under a multinormal distribution with an arbitrary unknown covariance matrix, the main purpose of this paper is to propose a framework to achieve the goal of reconciliation of Bayesian, frequentist, and Fisher's reporting $p$-values,…