Related papers: A Grid-Rate Condition for Valid Uniform Inference
Many functionals of interest in statistics and machine learning can be written as minimizers of expected loss functions. Such functionals are called $M$-estimands, and can be estimated by $M$-estimators -- minimizers of empirical average…
We introduce the space of grid functions, a space of generalized functions of nonstandard analysis that provides a coherent generalization both of the space of distributions and of the space of Young measures. We will show that in the space…
This paper considers an empirical likelihood inference for parameters defined by general estimating equations, when data are missing at random. The efficiency of existing estimators depends critically on correctly specifying the conditional…
Grokking is the intriguing phenomenon where a model learns to generalize long after it has fit the training data. We show both analytically and numerically that grokking can surprisingly occur in linear networks performing linear tasks in a…
Discrete-time affine processes are widely used in finance and economics and encompass count, positive, and nonnegative-valued processes. This paper develops near-unit-root asymptotic theory for this class of models. Unlike linear AR(1)…
Nonparametric regression problems with qualitative constraints such as monotonicity or convexity are ubiquitous in applications. For example, in predicting the yield of a factory in terms of the number of labor hours, the monotonicity of…
We provide a simple convergence proof for AdaGrad optimizing non-convex objectives under only affine noise variance and bounded smoothness assumptions. The proof is essentially based on a novel auxiliary function $\xi$ that helps eliminate…
Consider additive functionals of a Markov chain $W_k$, with stationary (marginal) distribution and transition function denoted by $\pi$ and $Q$, say $S_n=g(W_1)+...+g(W_n)$, where $g$ is square integrable and has mean 0 with respect to…
When are asymptotic approximations using the delta-method uniformly valid? We provide sufficient conditions as well as closely related necessary conditions for uniform negligibility of the remainder of such approximations. These conditions…
It is well known that approximation of functions on $[0,1]$ whose periodic extension is not continuous fail to converge uniformly due to rapid Gibbs oscillations near the boundary. Among several approaches that have been proposed toward the…
Linear models are foundational tools in statistics and ubiquitous across the applied sciences. However, conventional statistical inference -- such as $t$-tests and $F$-tests -- are only valid at fixed sample sizes, making them unsuitable…
A reinforcement algorithm introduced by H.A. Simon \cite{Simon} produces a sequence of uniform random variables with memory as follows. At each step, with a fixed probability $p\in(0,1)$, $\hat U_{n+1}$ is sampled uniformly from $\hat U_1,…
An important problem in network analysis is predicting a node attribute using both network covariates, such as graph embedding coordinates or local subgraph counts, and conventional node covariates, such as demographic characteristics.…
We consider the estimation of two-sample integral functionals, of the type that occur naturally, for example, when the object of interest is a divergence between unknown probability densities. Our first main result is that, in wide…
In this paper, the two settings we are concerned with are $\Gamma < \operatorname{SO}(n, 1)$ a Zariski dense Schottky semigroup and $\Gamma < \operatorname{SL}_2(\mathbb C)$ a Zariski dense continued fractions semigroup. In both settings,…
We consider the minimization of integral functionals in one dimension and their approximation by $r$-adaptive finite elements. Including the grid of the FEM approximation as a variable in the minimization, we are able to show that the…
Stochastic convex optimization is one of the most well-studied models for learning in modern machine learning. Nevertheless, a central fundamental question in this setup remained unresolved: "How many data points must be observed so that…
We consider the problem of testing equality of functions $f_j:[0,1]\to \mathbb{R}$ for $j=1,2,...,J$ the basis of $J$ independent samples from possibly different distributions under the assumption that the functions are monotone. We provide…
Many causal estimands, such as average treatment effects under unconfoundedness, can be written as continuous linear functionals of an unknown regression function. We study a weighting estimator that sets weights by a minimax procedure:…
Linear mixed models (LMMs) are suitable for clustered data and are common in biometrics, medicine, survey statistics and many other fields. In those applications, it is essential to carry out valid inference after selecting a subset of the…