Related papers: Concentration inequalities for the sample correlat…
The random coefficients model is an extension of the linear regression model that allows for unobserved heterogeneity in the population by modeling the regression coefficients as random variables. Given data from this model, the statistical…
Subsampling is an efficient method to deal with massive data. In this paper, we investigate the optimal subsampling for linear quantile regression when the covariates are functions. The asymptotic distribution of the subsampling estimator…
For many important problems the quantity of interest is an unknown function of the parameters, which is a random vector with known statistics. Since the dependence of the output on this random vector is unknown, the challenge is to identify…
Composite endpoints are widely used in cardiovascular clinical trials to improve statistical efficiency while preserving clinical relevance. The Win Ratio (WR) measure and more general frameworks of Win Statistics have emerged as…
A common goal in observational research is to estimate marginal causal effects in the presence of confounding variables. One solution to this problem is to use the covariate distribution to weight the outcomes such that the data appear…
Computing ratios of normalizing constants plays an important role in statistical modeling. Two important examples are hypothesis testing in latent variables models, and model comparison in Bayesian statistics. In both examples, the…
Cochran's $Q$ statistic is routinely used for testing heterogeneity in meta-analysis. Its expected value is also used for estimation of between-study variance $\tau^2$. Cochran's $Q$, or $Q_{IV}$, uses estimated inverse-variance weights…
We consider the standard non-parametric regression model with Gaussian errors but where the data consist of different samples. The question to be answered is whether the samples can be adequately represented by the same regression function.…
As context windows in large language models continue to expand, it is essential to characterize how attention behaves at extreme sequence lengths. We introduce token-sample complexity: the rate at which attention computed on $n$ tokens…
Testing differences between a treatment and control group is common practice in biomedical research like randomized controlled trials (RCT). The standard two-sample t-test relies on null hypothesis significance testing (NHST) via p-values,…
Let $Y$ be a nonnegative random variable with mean $\mu$ and finite positive variance $\sigma^2$, and let $Y^s$, defined on the same space as $Y$, have the $Y$ size biased distribution, that is, the distribution characterized by…
We derive a Gaussian approximation result for the maximum of a sum of random vectors under $(2+\iota)$-th moments. Our main theorem is abstract and nonasymptotic, and can be applied to a variety of statistical learning problems. The proof…
In a range of genomic applications, it is of interest to quantify the evidence that the signal at site~$i$ is active given conditionally independent replicate observations summarized by the sample mean and variance $(\bar Y, s^2)$ at each…
This paper is about vector autoregressive-moving average (VARMA) models with time-dependent coefficients to represent non-stationary time series. Contrarily to other papers in the univariate case, the coefficients depend on time but not on…
We propose an empirically stable and asymptotically efficient covariate-balancing approach to the problem of estimating survival causal effects in data with conditionally-independent censoring. This addresses a challenge often encountered…
In comparative studies, such as in causal inference and clinical trials, balancing important covariates is often one of the most important concerns for both efficient and credible comparison. However, chance imbalance still exists in many…
We consider a random variable $X$ that takes values in a (possibly infinite-dimensional) topological vector space $\mathcal{X}$. We show that, with respect to an appropriate "normal distance" on $\mathcal{X}$, concentration inequalities for…
Following [1], the aim of this paper is to analyze the relative weighted entropy involving the central moments weight functions. We compare the standard relative entropy with the weighted case in two particular forms of Gaussian…
The Gaussian product inequality is an important conjecture concerning the moments of Gaussian random vectors. While all attempts to prove the Gaussian product inequality in full generality have been unsuccessful to date, numerous partial…
Covariance matrix estimation concerns the problem of estimating the covariance matrix from a collection of samples, which is of extreme importance in many applications. Classical results have shown that $O(n)$ samples are sufficient to…