Related papers: How the instability of ranks under long memory aff…
While LLMs have seen substantial improvement in reasoning capabilities, they also sometimes overthink, generating unnecessary reasoning steps, particularly under uncertainty, given ill-posed or ambiguous queries. We introduce statistically…
This is the second part of a study of the limiting distributions of the top eigenvalues of a Hermitian matrix model with spiked external source under a general external potential. The case when the external source is of rank one was…
So-called linear rank statistics provide a means for distribution-free (even in finite samples), yet highly flexible, two-sample testing in the setting of univariate random variables. Their flexibility derives from a choice of weights that…
Large audio-language models (LALMs) are often used in tasks that involve reasoning over ordered options. An open question is whether their predictions are influenced by the order of answer choices, which would indicate a form of position…
We consider functionals of long-range dependent Gaussian sequences with infinite variance and obtain nonstandard limit theorems. When the long-range dependence is strong enough, the limit is a Hermite process, while for weaker long-range…
For some variants of regression models, including partial, measurement error or error-in-variables, latent effects, semi-parametric and otherwise corrupted linear models, the classical parametric tests generally do not perform well. Various…
We consider inference in linear regression models that is robust to heteroskedasticity and the presence of many control variables. When the number of control variables increases at the same rate as the sample size the usual…
We find that the performance of state-of-the-art models on Natural Language Inference (NLI) and Reading Comprehension (RC) analysis/stress sets can be highly unstable. This raises three questions: (1) How will the instability affect the…
We study the asymptotic behavior of the rank statistic for unimodal sequences. We use analytic techniques involving asymptotic expansions in order to prove asymptotic formulas for the moments of the rank. Furthermore, when appropriately…
This paper studies inference in linear models with a high-dimensional parameter matrix that can be well-approximated by a ``spiked low-rank matrix.'' A spiked low-rank matrix has rank that grows slowly compared to its dimensions and nonzero…
Large language models (LLMs) are increasingly used as decision-support tools in data-constrained scientific workflows, where correctness and validity are critical. However, evaluation practices often emphasize stability or reproducibility…
We revisit the problem of perturbing a large, i.i.d. random matrix by a finite rank error. It is known that when elements of the i.i.d. matrix have finite fourth moment, then the outlier eigenvalues of the perturbed matrix are close to the…
In complex scale-free networks, ranking the individual nodes based upon their importance has useful applications, such as the identification of hubs for epidemic control, or bottlenecks for controlling traffic congestion. However, in most…
Tests based on sample mean vectors and sample spatial signs have been studied in the recent literature for high dimensional data with the dimension larger than the sample size. For suitable sequences of alternatives, we show that the powers…
We use the martingale-theoretic approach of game-theoretic probability to incorporate imprecision into the study of randomness. In particular, we define several notions of randomness associated with interval, rather than precise,…
The Non-Linear Sigma Model (NLSM) is an example of a field theory on a target space exhibiting intricate geometry. One remarkable characteristic of the NLSM is asymptotic freedom, which triggers interest in perturbative calculations. In the…
Many statistical estimators are defined as the fixed point of a data-dependent operator, with estimators based on minimizing a cost function being an important special case. The limiting performance of such estimators depends on the…
Multivariate extreme value theory is concerned with modeling the joint tail behavior of several random variables. Existing work mostly focuses on asymptotic dependence, where the probability of observing a large value in one of the…
Many testing problems are readily amenable to randomised tests such as those employing data splitting. However despite their usefulness in principle, randomised tests have obvious drawbacks. Firstly, two analyses of the same dataset may…
Inspired by the importance of inhibitory and excitatory couplings in the brain, we analyze the largest eigenvalue statistics of random networks incorporating such features. We find that the largest real part of eigenvalues of a network,…