We address the problem of predicting whether sufficient memory and CPU resources have been requested for jobs at submission time. For this purpose, we examine the task of training a supervised machine learning system to predict the outcome - whether the job will fail specifically due to insufficient resources - as a classification task. Sufficiently high accuracy, precision, and recall at this task facilitates more anticipatory decision support applications in the domain of HPC resource allocation. Our preliminary results using a new test bed show that the probability of failed jobs is associated with information freely available at job submission time and may thus be usable by a learning system for user modeling that gives personalized feedback to users.
@article{arxiv.1806.01116,
title = {Machine Learning for Predictive Analytics of Compute Cluster Jobs},
author = {Dan Andresen and William Hsu and Huichen Yang and Adedolapo Okanlawon},
journal= {arXiv preprint arXiv:1806.01116},
year = {2018}
}
Comments
7 pages, CSC'18 - The 16th Int'l Conf on Scientific Computing