ML-for-ML
Abstract
AI training workloads are growing rapidly, making their time, energy, and infrastructure costs increasingly important. In shared cloud clusters, training and fine-tuning jobs compete with co-running workloads for network resources, while network mechanisms and ML training choices are typically optimized separately: networking controls how bytes move, whereas ML systems control when and how much communication occurs. We argue that this separation leaves end-to-end performance on the table. We present ML-for-ML, a cross-layer perspective in which network-side and ML-side knobs are selected jointly under a shared time-to-target-loss objective. Our preliminary prototype shows that by co-optimizing the ML and network parameters, we reach the target loss up to 42% faster.
Cite
@article{arxiv.2608.06046,
title = {ML-for-ML},
author = {Yutong Zhao and Noga H. Rotman and Gianni Antichi and Ran Ben Basat},
journal= {arXiv preprint arXiv:2608.06046},
year = {2026}
}
Comments
8 pages, 3 figures