English

Orthogonalized SGD and Nested Architectures for Anytime Neural Networks

Machine Learning 2020-08-18 v1 Machine Learning

Abstract

We propose a novel variant of SGD customized for training network architectures that support anytime behavior: such networks produce a series of increasingly accurate outputs over time. Efficient architectural designs for these networks focus on re-using internal state; subnetworks must produce representations relevant for both immediate prediction as well as refinement by subsequent network stages. We consider traditional branched networks as well as a new class of recursively nested networks. Our new optimizer, Orthogonalized SGD, dynamically re-balances task-specific gradients when training a multitask network. In the context of anytime architectures, this optimizer projects gradients from later outputs onto a parameter subspace that does not interfere with those from earlier outputs. Experiments demonstrate that training with Orthogonalized SGD significantly improves generalization accuracy of anytime networks.

Keywords

Cite

@article{arxiv.2008.06635,
  title  = {Orthogonalized SGD and Nested Architectures for Anytime Neural Networks},
  author = {Chengcheng Wan and Henry Hoffmann and Shan Lu and Michael Maire},
  journal= {arXiv preprint arXiv:2008.06635},
  year   = {2020}
}

Comments

ICML 2020

R2 v1 2026-06-23T17:52:30.259Z