English

TrimTuner: Efficient Optimization of Machine Learning Jobs in the Cloud via Sub-Sampling

Machine Learning 2020-11-11 v1 Distributed, Parallel, and Cluster Computing

Abstract

This work introduces TrimTuner, the first system for optimizing machine learning jobs in the cloud to exploit sub-sampling techniques to reduce the cost of the optimization process while keeping into account user-specified constraints. TrimTuner jointly optimizes the cloud and application-specific parameters and, unlike state of the art works for cloud optimization, eschews the need to train the model with the full training set every time a new configuration is sampled. Indeed, by leveraging sub-sampling techniques and data-sets that are up to 60x smaller than the original one, we show that TrimTuner can reduce the cost of the optimization process by up to 50x. Further, TrimTuner speeds-up the recommendation process by 65x with respect to state of the art techniques for hyper-parameter optimization that use sub-sampling techniques. The reasons for this improvement are twofold: i) a novel domain specific heuristic that reduces the number of configurations for which the acquisition function has to be evaluated; ii) the adoption of an ensemble of decision trees that enables boosting the speed of the recommendation process by one additional order of magnitude.

Keywords

Cite

@article{arxiv.2011.04726,
  title  = {TrimTuner: Efficient Optimization of Machine Learning Jobs in the Cloud via Sub-Sampling},
  author = {Pedro Mendes and Maria Casimiro and Paolo Romano and David Garlan},
  journal= {arXiv preprint arXiv:2011.04726},
  year   = {2020}
}

Comments

Mascots 2020

R2 v1 2026-06-23T20:01:44.508Z