English

Exploiting Elasticity in Tensor Ranks for Compressing Neural Networks

Machine Learning 2021-05-11 v1

Abstract

Elasticities in depth, width, kernel size and resolution have been explored in compressing deep neural networks (DNNs). Recognizing that the kernels in a convolutional neural network (CNN) are 4-way tensors, we further exploit a new elasticity dimension along the input-output channels. Specifically, a novel nuclear-norm rank minimization factorization (NRMF) approach is proposed to dynamically and globally search for the reduced tensor ranks during training. Correlation between tensor ranks across multiple layers is revealed, and a graceful tradeoff between model size and accuracy is obtained. Experiments then show the superiority of NRMF over the previous non-elastic variational Bayesian matrix factorization (VBMF) scheme.

Keywords

Cite

@article{arxiv.2105.04218,
  title  = {Exploiting Elasticity in Tensor Ranks for Compressing Neural Networks},
  author = {Jie Ran and Rui Lin and Hayden K. H. So and Graziano Chesi and Ngai Wong},
  journal= {arXiv preprint arXiv:2105.04218},
  year   = {2021}
}

Comments

8 pages, 5 figures

R2 v1 2026-06-24T01:56:11.955Z