Related papers: Sharper bounds for online learning of smooth funct…
Let $M_{G}$ be the centered Hardy-Littlewood maximal operator on a finite graph $G$. We find $\underset{p\to \infty}{\lim}\|M_{G}\|_{p}^{p }$ when $G$ is the start graph ($S_n$) and the complete graph ($K_n$), and we fully describe…
Q-functions are widely used in discrete-time learning and control to model future costs arising from a given control policy, when the initial state and input are given. Although some of their properties are understood, Q-functions…
For each $p \in (0,2]$, we present a randomized algorithm that returns an $\epsilon$-approximation of the $p$th frequency moment of a data stream $F_p = \sum_{i = 1}^n \abs{f_i}^p$. The algorithm requires space $O(\epsilon^{-2} \log…
Let $f: \mathbb{N}^2 \mapsto \mathbb{C}$ be an arithmetic function of two variables. We study the existence of the limit: \[\displaystyle \lim_{x \to \infty} \frac{1}{x^2 (\log x)^{k-1}} \sum_{n_1 , n_2 \le x} f (n_1, n_2) \] where $k$ is a…
We consider learning in an adversarial Markov Decision Process (MDP) where the loss functions can change arbitrarily over $K$ episodes and the state space can be arbitrarily large. We assume that the Q-function of any policy is linear in…
Unconstrained Online Linear Optimization (OLO) is a practical problem setting to study the training of machine learning models. Existing works proposed a number of potential-based algorithms, but in general the design of these potential…
In many iterative optimization methods, fixed-point theory enables the analysis of the convergence rate via the contraction factor associated with the linear approximation of the fixed-point operator. While this factor characterizes the…
Error bounds and complexity bounds in numerical analysis and information-based complexity are often proved for functions that are defined on very simple domains, such as a cube, a torus, or a sphere. We study optimal error bounds for the…
We establish optimal Statistical Query (SQ) lower bounds for robustly learning certain families of discrete high-dimensional distributions. In particular, we show that no efficient SQ algorithm with access to an $\epsilon$-corrupted binary…
We investigate the problem of minimizing the excess generalization error with respect to the best expert prediction in a finite family in the stochastic setting, under limited access to information. We assume that the learner only has…
This work focuses on the setting of dynamic regret in the context of online learning with full information. In particular, we analyze regret bounds with respect to the temporal variability of the loss functions. By assuming that the…
Much of modern learning theory has been split between two regimes: the classical offline setting, where data arrive independently, and the online setting, where data arrive adversarially. While the former model is often both computationally…
The regret bound of dynamic online learning algorithms is often expressed in terms of the variation in the function sequence ($V_T$) and/or the path-length of the minimizer sequence after $T$ rounds. For strongly convex and smooth…
In many quantum tasks, there is an unknown quantum object that one wishes to learn. An online strategy for this task involves adaptively refining a hypothesis to reproduce such an object or its measurement statistics. A common evaluation…
We study algorithms for online linear optimization in Hilbert spaces, focusing on the case where the player is unconstrained. We develop a novel characterization of a large class of minimax algorithms, recovering, and even improving,…
It is well known that if a function $f$ satisfies $$\|f(x) e^{\pi \alpha |x|^2}\|_p + \| \widehat{f}(\xi) e^{\pi \alpha |\xi|^2} \|_q<\infty \qquad\qquad\qquad(*)$$ with $\alpha=1$ and $1\le p,q<\infty$, then $f\equiv 0.$ We prove that if…
We extend classical analytic tools for finite-state statistical models to allow zero probabilities. Using methods from algebraic statistics and information geometry, we develop a framework in which a smooth statistical model could hit the…
We propose an active-learning method for nonlinear minimax regression. Given a nonlinear function that can be arbitrarily evaluated over a compact set, we fit a surrogate model, such as a feedforward neural network, by minimizing the…
The purpose of this paper is to prove pointwise inequalities and to establish the boundedness on weighted $L^{p}$ spaces for pseudo-differential operators $T_{a}$ defined by the symbol $a\in S^{m}_{\varrho,\delta}$ with $0\leq\varrho\leq1,$…
We consider the stochastic optimization problem with smooth but not necessarily convex objectives in the heavy-tailed noise regime, where the stochastic gradient's noise is assumed to have bounded $p$th moment ($p\in(1,2]$). Zhang et al.…