Related papers: Gradient descent in higher codimension
The effective field theory for hydrodynamics allows to write the action functional for fluid. In this paper, some simplest possible higher derivative terms in the fluid action and the cosmological consequences of their presence are…
Understanding the behavior of stochastic gradient descent (SGD) in the context of deep neural networks has raised lots of concerns recently. Along this line, we study a general form of gradient based optimization dynamics with unbiased…
Modern neural networks are usually highly over-parameterized. Behind the wide usage of over-parameterized networks is the belief that, if the data are simple, then the trained network will be automatically equivalent to a simple predictor.…
This paper develops the theoretical foundations for the ability of a control field to cooperate with noise in the manipulation of quantum dynamics. The noise enters as run-to-run variations in the control amplitudes, phases and frequencies…
Several key questions remain unanswered regarding overparameterized learning models. It is unclear how (stochastic) gradient descent finds solutions that generalize well, and in particular the role of small random initializations. Matrix…
The predictability of discrete-time processes is studied in a deterministic setting. A family of one-step-ahead predictors is suggested for processes of which the energy decays at higher frequencies. For such processes, the prediction error…
Fluctuations in a fluid are strongly affected by the presence of a macroscopic gradient making them long-ranged and enhancing their amplitude. While small-scale fluctuations exhibit diffusive lifetimes, larger-scale fluctuations live…
Assuming a-priori a smooth generating vector field, we introduce a generally covariant measure of the flow geometry called the referential gradient of the flow. The main result is the explicit relation between the referential gradient and…
We present a systematic study of moment evolution in multidimensional stochastic difference systems, focusing on characterizing systems whose low-order moments diverge in the neighborhood of a stable fixed point. We consider systems with a…
Understanding the behavior of stochastic gradient methods is a central problem in modern machine learning. Recent work has highlighted diagonal linear networks as a simplified yet expressive setting for analyzing the optimization and…
Gradient descent (GD) on logistic regression has many fascinating properties. When the dataset is linearly separable, it is known that the iterates converge in direction to the maximum-margin separator regardless of how large the step size…
We present a study of the effects of decoherence in the operation of a discrete quantum walk on a line, cycle and hypercube. We find high sensitivity to decoherence, increasing with the number of steps in the walk, as the particle is…
We present a detailed analysis of random motions moving in higher spaces with a natural number of velocities. In the case of the so-called minimal random dynamics, under some wide assumptions, we show the joint distribution of the position…
Gradient descent algorithms have been used in countless applications since the inception of Newton's method. The explosion in the number of applications of neural networks has re-energized efforts in recent years to improve the standard…
The fundamental processes of biological development are governed by multiple signaling molecules that create non-uniform concentration profiles known as morphogen gradients. It is widely believed that the establishment of morphogen…
Interpreting gradient methods as fixed-point iterations, we provide a detailed analysis of those methods for minimizing convex objective functions. Due to their conceptual and algorithmic simplicity, gradient methods are widely used in…
The dynamical properties of a quantum system can be profoundly influenced by its environment. Usually, the environment provokes decoherence and its action on the system can often be schematized by adding a noise term in the Hamiltonian.…
Distributed gradient descent algorithms have come to the fore in modern machine learning, especially in parallelizing the handling of large datasets that are distributed across several workers. However, scant attention has been paid to…
In this paper, we propose new structured second-order methods and structured adaptive-gradient methods obtained by performing natural-gradient descent on structured parameter spaces. Natural-gradient descent is an attractive approach to…
We prove quantitative convergence rates at which discrete Langevin-like processes converge to the invariant distribution of a related stochastic differential equation. We study the setup where the additive noise can be non-Gaussian and…