We prove that the evolution of weight vectors in online gradient descent can encode arbitrary polynomial-space computations, even in very simple learning settings. Our results imply that, under weak complexity-theoretic assumptions, it is impossible to reason efficiently about the fine-grained behavior of online gradient descent.
@article{arxiv.1807.01280,
title = {On the Computational Power of Online Gradient Descent},
author = {Vaggos Chatziafratis and Tim Roughgarden and Joshua R. Wang},
journal= {arXiv preprint arXiv:1807.01280},
year = {2019}
}
Comments
Added results, linear regression, neural nets. Fixed typos