Related papers: Maximum value of the standardized log of odds rati…
The ordered weighted $\ell_1$ (OWL) norm is a newly developed generalization of the Octogonal Shrinkage and Clustering Algorithm for Regression (OSCAR) norm. This norm has desirable statistical properties and can be used to perform…
The Oort constants describe the local variations of the stellar streaming field. Classically, they are determined from stellar proper motions. We discuss problems arising in this procedure. A large, hitherto overlooked, source of systematic…
We undertake a critical evaluation of recent observational information on $\Omega_m$ and $\Omega_\Lambda$ in order to identify possible sources of systematic errors and effects of simplified statistical analyses. We combine observations for…
If the log likelihood is approximately quadratic with constant Hessian, then the maximum likelihood estimator (MLE) is approximately normally distributed. No other assumptions are required. We do not need independent and identically…
In this paper, we address high-dimensional parametric estimation of the drift function in diffusion models, specifically focusing on a $d$-dimensional ergodic diffusion process observed at discrete time points. We consider both a general…
Motivated by multiple statistical hypothesis testing, we obtain the limit of likelihood ratio of large deviations for self-normalized random variables, specifically, the ratio of $P(\sqrt{n}(\bar X +d/n) \ge x_n V)$ to $P(\sqrt{n}\bar X \ge…
Linear mixed models (LMMs) are used as an important tool in the data analysis of repeated measures and longitudinal studies. The most common form of LMMs utilize a normal distribution to model the random effects. Such assumptions can often…
Ordinal classification models assign higher penalties to predictions further away from the true class. As a result, they are appropriate for relevant diagnostic tasks like disease progression prediction or medical image grading. The…
The behavior of maximum likelihood estimates (MLEs) and the likelihood ratio statistic in a family of problems involving pointwise nonparametric estimation of a monotone function is studied. This class of problems differs radically from the…
Conditional effects are commonly used measures for understanding how treatment effects vary across different groups, and are often used to target treatments/interventions to groups who benefit most. In this work we review existing methods…
Double blind randomized controlled trials are traditionally seen as the gold standard for causal inferences as the difference-in-means estimator is an unbiased estimator of the average treatment effect in the experiment. The fact that this…
The asymptotic normality of the Maximum Likelihood Estimator (MLE) is a long established result. Explicit bounds for the distributional distance between the distribution of the MLE and the normal distribution have recently been obtained for…
Logistic regression is commonly used for modeling dichotomous outcomes. In the classical setting, where the number of observations is much larger than the number of parameters, properties of the maximum likelihood estimator in logistic…
Let $I(n,l)$ denote the maximum possible number of incidences between $n$ points and $l$ lines. It is well known that $I(n,l) = \Theta(n^{2/3}l^{2/3} + n + l)$. Let $c_{\mathrm{SzTr}}$ denote the lower bound on the constant of…
System identification is a fundamental problem in control and learning, particularly in high-stakes applications where data efficiency is critical. Classical approaches, such as the ordinary least squares estimator (OLS), achieve an…
Selberg's central limit theorem states that the values of $\log|\zeta(1/2+i \tau)|$, where $\tau$ is a uniform random variable on $[T,2T]$, is distributed like a Gaussian random variable of mean $0$ and standard deviation…
Score matching is an alternative to maximum likelihood (ML) for estimating a probability distribution parametrized up to a constant of proportionality. By fitting the ''score'' of the distribution, it sidesteps the need to compute this…
The paper considers the problem of out-of-sample risk estimation under the high dimensional settings where standard techniques such as $K$-fold cross validation suffer from large biases. Motivated by the low bias of the leave-one-out cross…
Leave-one-out (LOO) prediction provides a principled, data-dependent measure of generalization, yet guarantees in fully transductive settings remain poorly understood beyond specialized models. We introduce Median of Level-Set Aggregation…
The normalized maximum likelihood (NML) is a recent penalized likelihood that has properties that justify defining the amount of discrimination information (DI) in the data supporting an alternative hypothesis over a null hypothesis as the…