Related papers: Towards Quantifying the Preconditioning Effect of …
We consider the problem of removal of ordering ambiguity in position dependent mass quantum systems characterized by a generalized position dependent mass Hamiltonian which generalizes a number of Hermitian as well as non-Hermitian ordered…
Canonical Polyadic Decomposition (CPD) of a third-order tensor is a minimal decomposition into a sum of rank-$1$ tensors. We find new mild deterministic conditions for the uniqueness of individual rank-$1$ tensors in CPD and present an…
The success of deep learning can be attributed to various factors such as increase in computational power, large datasets, deep convolutional neural networks, optimizers etc. Particularly, the choice of optimizer affects the generalization,…
Many problems encountered in science and engineering can be formulated as estimating a low-rank object (e.g., matrices and tensors) from incomplete, and possibly corrupted, linear measurements. Through the lens of matrix and tensor…
Following the previous work [1], we investigate the impact of damping on the oscillation of smooth solutions to some kind of quasilinear wave equations with Robin and Dirichlet boundary condition. By using generalized Riccati transformation…
We consider coordinate descent (CD) methods with exact line search on convex quadratic problems. Our main focus is to study the performance of the CD method that use random permutations in each epoch and compare it to the performance of the…
The remarkable success of the Adam in training neural networks has naturally led to the widespread use of its descent-ascent counterpart, Adam-DA, for solving zero-sum games. Despite its popularity in practice, a rigorous theoretical…
We propose an operator preconditioner for general elliptic pseudodifferential equations in a domain $\Omega$, where $\Omega$ is either in $\mathbb{R}^n$ or in a Riemannian manifold. For linear systems of equations arising from low-order…
In this paper, we prove that an Adam-type algorithm with smooth clipping approaches the global minimizer of the regularized non-convex loss function. Adding smooth clipping and taking the state space as the set of all trajectories, we can…
Research into optimisation for deep learning is characterised by a tension between the computational efficiency of first-order, gradient-based methods (such as SGD and Adam) and the theoretical efficiency of second-order, curvature-based…
Adaptive optimization algorithms, such as Adam and RMSprop, have shown better optimization performance than stochastic gradient descent (SGD) in some scenarios. However, recent studies show that they often lead to worse generalization…
The computation of \(\operatorname{tr}(AB)\) is essential in quantum science and artificial intelligence, yet classical methods for \( d \)-dimensional matrices \( A \) and \( B \) require \( O(d^2) \) complexity, which becomes infeasible…
In this paper, we investigate the convergence properties of a wide class of Adam-family methods for minimizing quadratically regularized nonsmooth nonconvex optimization problems, especially in the context of training nonsmooth neural…
Our main result is an abstract good-$\lambda$ inequality that allows us to consider three self-improving properties related to oscillation estimates in a very general context. The novelty of our approach is that there is one principle…
Canonical quantization relies on Cartesian, canonical, phase-space coordinates to promote to Hermitian operators, which also become the principal ingredients in the quantum Hamiltonian. While generally appropriate, this procedure can also…
In solving a system of $n$ linear equations in $d$ variables $Ax=b$, the condition number of the $n,d$ matrix $A$ measures how much errors in the data $b$ affect the solution $x$. Estimates of this type are important in many inverse…
We study Hessian estimators for functions defined over an $n$-dimensional complete analytic Riemannian manifold. We introduce new stochastic zeroth-order Hessian estimators using $O (1)$ function evaluations. We show that, for an analytic…
We investigate the connection between local minima in the problem Hamiltonian and first order quantum phase transitions during an adiabatic quantum computation. We demonstrate how some properties of the local minima can lead to an extremely…
This paper studies a class of adaptive gradient based momentum algorithms that update the search directions and learning rates simultaneously using past gradients. This class, which we refer to as the "Adam-type", includes the popular…
We here adapt an extended version of the adaptive cubic regularisation method with dynamic inexact Hessian information for nonconvex optimisation in [3] to the stochastic optimisation setting. While exact function evaluations are still…