Related papers: Fermat descent
This brief survey deals with multi-dimensional Diophantine approximations in sense of linear form and with simultaneous Diophantine approximations. We discuss the phenomenon of degenerate dimension of linear subspaces generated by the best…
Composite optimization problems, where the sum of a smooth and a merely lower semicontinuous function has to be minimized, are often tackled numerically by means of proximal gradient methods as soon as the lower semicontinuous part of the…
A wide range of optimization problems arising in machine learning can be solved by gradient descent algorithms, and a central question in this area is how to efficiently compress a large-scale dataset so as to reduce the computational…
This is an overview of higher structural constructions in physics. The main motivations of our current attempt are as follows: (i) to provide a brief introduction to derived algebraic geometry, (ii) to understand how derived objects…
The notion of descent set, for permutations as well as for standard Young tableaux (SYT), is classical. Cellini introduced a natural notion of {\em cyclic descent set} for permutations, and Rhoades introduced such a notion for SYT --- but…
Fractional gradient descent has been studied extensively, with a focus on its ability to extend traditional gradient descent methods by incorporating fractional-order derivatives. This approach allows for more flexibility in navigating…
In deep learning, it is common to use more network parameters than training points. In such scenarioof over-parameterization, there are usually multiple networks that achieve zero training error so that thetraining algorithm induces an…
A perplexing problem in understanding physical reality is why the universe seems comprehensible, and correspondingly why there should exist physical systems capable of comprehending it. In this essay I explore the possibility that rather…
This second part of the paper strengthens the descent theory described in the first part to rational maps, arbitrary base fields, and dynamics given by correspondences. We obtain in particular a decomposition of any difference field…
We obtain additional Diophantine applications of the methods surrounding Darmon's program for the generalized Fermat equation developed in the first part of this series of papers. As a first application, we use a multi-Frey approach…
We develop a theory of descent and forms of tensor categories over arbitrary fields. We describe the general scheme of classification of such forms using algebraic and homotopical language, and give examples of explicit classification of…
The decentralized gradient descent (DGD) algorithm, and its sibling, diffusion, are workhorses in decentralized machine learning, distributed inference and estimation, and multi-agent coordination. We propose a novel, principled framework…
In this paper, we establish new convergence results for the quantized distributed gradient descent and suggest a novel strategy of choosing the stepsizes for the high-performance of the algorithm. Under the strongly convexity assumption on…
It seems that in the current age, computers, computation, and data have an increasingly important role to play in scientific research and discovery. This is reflected in part by the rise of machine learning and artificial intelligence,…
We develop a formalism for studying descent and codescent in the context of Iwasawa theory. The main result essentially states that to control descent or codescent amounts to the same. Arithmetic applications are given.
Optimization problems occurring in a wide variety of physical design problems, including but not limited to optical engineering, quantum control, structural engineering, involve minimization of a simple cost function of the state of the…
In gradient descent, changing how we parametrize the model can lead to drastically different optimization trajectories, giving rise to a surprising range of meaningful inductive biases: identifying sparse classifiers or reconstructing…
Gradient Descent (GD) is a ubiquitous algorithm for finding the optimal solution to an optimization problem. For reduced computational complexity, the optimal solution $\mathrm{x^*}$ of the optimization problem must be attained in a minimum…
Gradient descent method, as one of the major methods in numerical optimization, is the key ingredient in many machine learning algorithms. As one of the most fundamental way to solve the optimization problems, it promises the function value…
This paper considers the problem of solving systems of quadratic equations, namely, recovering an object of interest $\mathbf{x}^{\natural}\in\mathbb{R}^{n}$ from $m$ quadratic equations/samples…