On Convergence and Generalization of Dropout Training
Machine Learning
2020-10-27 v1 Machine Learning
Abstract
We study dropout in two-layer neural networks with rectified linear unit (ReLU) activations. Under mild overparametrization and assuming that the limiting kernel can separate the data distribution with a positive margin, we show that dropout training with logistic loss achieves -suboptimality in test error in iterations.
Cite
@article{arxiv.2010.12711,
title = {On Convergence and Generalization of Dropout Training},
author = {Poorya Mianjy and Raman Arora},
journal= {arXiv preprint arXiv:2010.12711},
year = {2020}
}