Related papers: Efficiency of Z-estimators indexed by the objectiv…
A linear functional of an object from a convex symmetric set can be optimally estimated, in a worst-case sense, by a linear functional of observations made on the object. This well-known fact is extended here to a nonlinear setting: other…
We consider first order expansions of convex penalized estimators in high-dimensional regression problems with random designs. Our setting includes linear regression and logistic regression as special cases. For a given penalty function $h$…
We introduce a deficiency-based representation and approximation framework for values of the Riemann zeta function. The method is based on comparing two nonlinear accumulation mechanisms: global transformation of a base partial sum and…
To evaluate Riemann's zeta function is important for many investigations related to the area of number theory, and to have quickly converging series at hand in particular. We investigate a class of summation formulae and find, as a special…
We consider a stochastic version of the proximal point algorithm for optimization problems posed on a Hilbert space. A typical application of this is supervised learning. While the method is not new, it has not been extensively analyzed in…
The following theorem is the main result of this note. Theorem 1. Let $(E, \|\cdot\|_E) $ be a rearrangement invariant Banach function space on the interval $[0, 1]$. If $E$ is isometric to $\L_p [0, 1]$ for some $1\le p<\infty$, then $E$…
In this paper, we study the stability of the orthogonal equation,which is closely related to the results by Wlodzimierz Fechner and Justyna Sikorska in 2010. There are some differences that we consider the target space with the…
Given a large set $U$ where each item $a\in U$ has weight $w(a)$, we want to estimate the total weight $W=\sum_{a\in U} w(a)$ to within factor of $1\pm\varepsilon$ with some constant probability $>1/2$. Since $n=|U|$ is large, we want to do…
For expectation functions on metric spaces, we provide sufficient conditions for epi-convergence under varying probability measures and integrands, and examine applications in the area of sieve estimators, mollifier smoothing,…
An algorithm is given for determining an optimal $b$-step approximation of weighted data, where the error is measured with respect to the $L_\infty$ norm. For data presorted by the independent variable the algorithm takes $\Theta(n + \log n…
The aim of this paper is to prove the exponential convergence, local and global, of Adam algorithm under precise conditions on the parameters, when the objective function lacks differentiability. More precisely, we require Lipschitz…
Importance weighting is a standard tool for correcting distribution shift, but its statistical behavior under target shift -- where the label distribution changes between training and testing while the conditional distribution of inputs…
We investigate the consequence of two Lip$(\gamma)$ functions, in the sense of Stein, being close throughout a subset of their domain. A particular consequence of our results is the following. Given $K_0 > \varepsilon > 0$ and $\gamma >…
We study the values taken by the Riemann zeta-function $\zeta$ on discrete sets. We show that infinite vertical arithmetic progressions are uniquely determined by the values of $\zeta$ taken on this set. Moreover, we prove a joint discrete…
We consider a robust linear regression model $y=X\beta^* + \eta$, where an adversary oblivious to the design $X\in \mathbb{R}^{n\times d}$ may choose $\eta$ to corrupt all but an $\alpha$ fraction of the observations $y$ in an arbitrary…
Assuming that $(X_t)_{t\in\Z}$ is a vector valued time series with a common marginal distribution admitting a density $f$, our aim is to provide a wide range of consistent estimators of $f$. We consider different methods of estimation of…
Evaluation of treatment effects and more general estimands is typically achieved via parametric modelling, which is unsatisfactory since model misspecification is likely. Data-adaptive model building (e.g. statistical/machine learning) is…
Test-time adaptation (TTA) seeks to tackle potential distribution shifts between training and testing data by adapting a given model w.r.t. any testing sample. This task is particularly important for deep models when the test environment…
Estimating value functions is a core component of reinforcement learning algorithms. Temporal difference (TD) learning algorithms use bootstrapping, i.e. they update the value function toward a learning target using value estimates at…
Convex and penalized robust regression methods often suffer from a persistent bias induced by large outliers, limiting their effectiveness in adversarial or heavy-tailed settings. In this work, we study a smooth redescending non-convex…