Related papers: Implicit Bias of Mirror Flow on Separable Data
First-order optimization methods tend to inherently favor certain solutions over others when minimizing an underdetermined training objective that has multiple global optima. This phenomenon, known as implicit bias, plays a critical role in…
Balancing policy expressiveness with the exploration-exploitation trade-off is a core challenge in online Reinforcement Learning (RL). While Stochastic Differential Equation (SDE)-based diffusion policies can represent complex, multimodal…
Real-world object classes appear in imbalanced ratios. This poses a significant challenge for classifiers which get biased towards frequent classes. We hypothesize that improving the generalization capability of a classifier should improve…
This letter presents an almost sure convergence of the zeroth-order mirror descent algorithm. The algorithm admits non-smooth convex functions and a biased oracle which only provides noisy function value at any desired point. We approximate…
Gradient descent, or negative gradient flow, is a standard technique in optimization to find minima of functions. Many implementations of gradient descent rely on discretized versions, i.e., moving in the gradient direction for a set step…
Recently there has been renewed interests in derivative free approaches to stochastic optimization. In this paper, we examine the rates of convergence for the Kiefer-Wolfowitz algorithm and the mirror descent algorithm, under various…
We first define Pseudo-Calabi flow, as {equation*} {{aligned}{{\partial \varphi}\over {\partial t}}&= -f(\varphi), \triangle_varphi f(\varphi) &= S(\varphi) - \ul S.{aligned}. \end{equation*} Then we prove the well-posedness of this flow…
Let $F:[0,T]\times\R^n\mapsto 2^{\R^n}$ be a continuous multifunction with compact, not necessarily convex values. In this paper, we prove that, if $F$ satisfies the following Lipschitz Selection Property: \begin{itemize} \item[{(LSP)}]…
While first-order optimization methods are usually designed to efficiently reduce the function value $f(x)$, there has been recent interest in methods efficiently reducing the magnitude of $\nabla f(x)$, and the findings show that the two…
On a polarized manifold $(X,L)$, the Bergman iteration $\phi_k^{(m)}$ is defined as a sequence of Bergman metrics on $L$ with two integer parameters $k, m$. We study the relation between the K\"ahler-Ricci flow $\phi_t$ at any time $t \geq…
We propose primal-dual stochastic mirror descent for the convex optimization problems with functional constraints. We obtain the rate of convergence in terms of probability of large deviations.
Distributed optimization often requires finding the minimum of a global objective function written as a sum of local functions. A group of agents work collectively to minimize the global function. We study a continuous-time decentralized…
We study the implicit bias of gradient flow (i.e., gradient descent with infinitesimal step size) on linear neural network training. We propose a tensor formulation of neural networks that includes fully-connected, diagonal, and…
This paper is devoted to the variational inequality problems. We consider two classes of problems, the first is classical constrained variational inequality and the second is the same problem with functional (inequality type) constraints.…
Mirror descent is an elegant optimization technique that leverages a dual space of parametric models to perform gradient descent. While originally developed for convex optimization, it has increasingly been applied in the field of machine…
The steady, asymmetric and two-dimensional flow of viscous, incompressible micropolar fluid through a rectangular channel with a splitter (parallel to walls) was formulated and simulated numerically. The plane Poiseuille flow was considered…
In this work we review the coarse-to-fine spatial feature pyramid concept, which is used in state-of-the-art optical flow estimation networks to make exploration of the pixel flow search space computationally tractable and efficient. Within…
Entropy regularization has been widely used in policy optimization algorithms to enhance exploration and the robustness of the optimal control; however it also introduces an additional regularization bias. This work quantifies the impact of…
We provide several applications of Optimistic Mirror Descent, an online learning algorithm based on the idea of predictable sequences. First, we recover the Mirror Prox algorithm for offline optimization, prove an extension to Holder-smooth…
We reduce the solution of the scattering problem defined on the half-line $[0,\infty)$ by a real or complex potential $v(x)$ and a general homogenous boundary condition at $x=0$ to that of the extension of $v(x)$ to the full line that…