English

The Implicit Bias of AdaGrad on Separable Data

Machine Learning 2019-06-11 v1 Machine Learning

Abstract

We study the implicit bias of AdaGrad on separable linear classification problems. We show that AdaGrad converges to a direction that can be characterized as the solution of a quadratic optimization problem with the same feasible set as the hard SVM problem. We also give a discussion about how different choices of the hyperparameters of AdaGrad might impact this direction. This provides a deeper understanding of why adaptive methods do not seem to have the generalization ability as good as gradient descent does in practice.

Keywords

Cite

@article{arxiv.1906.03559,
  title  = {The Implicit Bias of AdaGrad on Separable Data},
  author = {Qian Qian and Xiaoyuan Qian},
  journal= {arXiv preprint arXiv:1906.03559},
  year   = {2019}
}
R2 v1 2026-06-23T09:47:57.747Z