Related papers: Gradient and Lipschitz estimates for tug-of-war ty…
We describe algorithms for finding the regression of t, a sequence of values, to the closest sequence s by mean squared error, so that s is always increasing (isotonicity) and so the values of two consecutive points do not increase by too…
We formulate, for continuous-time dynamical systems, a sufficient condition to be a gradient-like system, i.e. that all bounded trajectories approach stationary points and therefore that periodic orbits, chaotic attractors, etc. do not…
For two-person dynamic zero-sum games (both discrete and continuous settings), we investigate the limit of value functions of finite horizon games with long run average cost as the time horizon tends to infinity and the limit of value…
We provide sufficient conditions for instability of the subgradient method with constant step size around a local minimum of a locally Lipschitz semi-algebraic function. They are satisfied by several spurious local minima arising in robust…
We consider Calder\'on-Zygmund type estimates for the non-homogeneous $p(\cdot)$-Laplacian system $ -\text{div}(|D u|^{p(\cdot)-2} Du) = -\text{div}(|G|^{p(\cdot)-2} G),$ where $p$ is a variable exponent. We show that $|G|^{p(\cdot)} \in…
We examine the impact of learning Lipschitz continuous models in the context of model-based reinforcement learning. We provide a novel bound on multi-step prediction error of Lipschitz models where we quantify the error using the…
We prove interpolating estimates providing a bound for the oscillation of a function in terms of two $L^p$ norms of its gradient. They are based on a pointwise bound of a function on cones in terms of the Riesz potential of its gradient.…
We study a class of zero-sum stochastic games between a stopper and a singular-controller, previously considered in [Bovo and De Angelis (2025)]. The underlying singularly-controlled dynamics takes values in…
Here we study what we call bounded rough Riemannian metrics $(M,g)$, which are positive definite, symmetric tensors on each tangent space, $T_pM$, which are bounded and measurable as functions in coordinates. This is enough structure to…
We obtain Gr\"onwall type estimates for the gradient of the harmonic functions for a L\'evy operator with order strictly larger than 1 and minimal assumptions of its L\'evy measure.
In reinforcement learning, temporal difference (TD) is the most direct algorithm to learn the value function of a policy. For large or infinite state spaces, exact representations of the value function are usually not available, and it must…
Rademacher's Theorem can be interpreted as an almost-everywhere \emph{little-$o$ improvement principle}: if a function admits a uniform pointwise first-order Lipschitz control at every point, then this control improves to a vanishing one at…
Reparameterization (RP) and likelihood ratio (LR) gradient estimators are used to estimate gradients of expectations throughout machine learning and reinforcement learning; however, they are usually explained as simple mathematical tricks,…
Gradient-variation online learning aims to achieve regret guarantees that scale with variations in the gradients of online functions, which has been shown to be crucial for attaining fast convergence in games and robustness in stochastic…
Gradient clipping is a popular modification to standard (stochastic) gradient descent, at every iteration limiting the gradient norm to a certain value $c >0$. It is widely used for example for stabilizing the training of deep learning…
We establish gradient estimates of solutions to a class of nonlinear elliptic equations with measure data under Orlicz-type growth conditions. The growth is governed by the structural condition \[ 0<i_a\le t g'(t)/g(t)\le s_a<1. \] We…
The aim of this short paper is to show that some assumptions in [10] can be relaxed and even dropped when looking for weak solutions instead of strong ones. This improvement is a consequence of two results concerning gradient terms: an…
We consider the linear elliptic systems or equations in divergence form with periodically oscillating coefficients. We prove the large-scale boundary Lipschitz estimate for the weak solutions in domains satisfying the so-called…
Off-policy evaluation provides an essential tool for evaluating the effects of different policies or treatments using only observed data. When applied to high-stakes scenarios such as medical diagnosis or financial decision-making, it is…
Understanding and analyzing markets is crucial, yet analytical equilibrium solutions remain largely infeasible. Recent breakthroughs in equilibrium computation rely on zeroth-order policy gradient estimation. These approaches commonly…