Joslim: Joint Widths and Weights Optimization for Slimmable Neural Networks
Abstract
Slimmable neural networks provide a flexible trade-off front between prediction error and computational requirement (such as the number of floating-point operations or FLOPs) with the same storage requirement as a single model. They are useful for reducing maintenance overhead for deploying models to devices with different memory constraints and are useful for optimizing the efficiency of a system with many CNNs. However, existing slimmable network approaches either do not optimize layer-wise widths or optimize the shared-weights and layer-wise widths independently, thereby leaving significant room for improvement by joint width and weight optimization. In this work, we propose a general framework to enable joint optimization for both width configurations and weights of slimmable networks. Our framework subsumes conventional and NAS-based slimmable methods as special cases and provides flexibility to improve over existing methods. From a practical standpoint, we propose Joslim, an algorithm that jointly optimizes both the widths and weights for slimmable nets, which outperforms existing methods for optimizing slimmable networks across various networks, datasets, and objectives. Quantitatively, improvements up to 1.7% and 8% in top-1 accuracy on the ImageNet dataset can be attained for MobileNetV2 considering FLOPs and memory footprint, respectively. Our results highlight the potential of optimizing the channel counts for different layers jointly with the weights for slimmable networks. Code available at https://github.com/cmu-enyac/Joslim.
Cite
@article{arxiv.2007.11752,
title = {Joslim: Joint Widths and Weights Optimization for Slimmable Neural Networks},
author = {Ting-Wu Chin and Ari S. Morcos and Diana Marculescu},
journal= {arXiv preprint arXiv:2007.11752},
year = {2021}
}
Comments
Accepted at ECML-PKDD 2021 (Research Track), 4-page abridged versions have been accepted at non-archival venues including RealML and DMMLSys workshops at ICML'20 and DLP-KDD and AdvML workshops at KDD'20