English

On Convergence and Generalization of Dropout Training

Machine Learning 2020-10-27 v1 Machine Learning

Abstract

We study dropout in two-layer neural networks with rectified linear unit (ReLU) activations. Under mild overparametrization and assuming that the limiting kernel can separate the data distribution with a positive margin, we show that dropout training with logistic loss achieves ϵ\epsilon-suboptimality in test error in O(1/ϵ)O(1/\epsilon) iterations.

Keywords

Cite

@article{arxiv.2010.12711,
  title  = {On Convergence and Generalization of Dropout Training},
  author = {Poorya Mianjy and Raman Arora},
  journal= {arXiv preprint arXiv:2010.12711},
  year   = {2020}
}
R2 v1 2026-06-23T19:36:30.164Z