English

ML-for-ML

Networking and Internet Architecture 2026-08-06 v1 Distributed, Parallel, and Cluster Computing Machine Learning

Abstract

AI training workloads are growing rapidly, making their time, energy, and infrastructure costs increasingly important. In shared cloud clusters, training and fine-tuning jobs compete with co-running workloads for network resources, while network mechanisms and ML training choices are typically optimized separately: networking controls how bytes move, whereas ML systems control when and how much communication occurs. We argue that this separation leaves end-to-end performance on the table. We present ML-for-ML, a cross-layer perspective in which network-side and ML-side knobs are selected jointly under a shared time-to-target-loss objective. Our preliminary prototype shows that by co-optimizing the ML and network parameters, we reach the target loss up to 42% faster.

Cite

@article{arxiv.2608.06046,
  title  = {ML-for-ML},
  author = {Yutong Zhao and Noga H. Rotman and Gianni Antichi and Ran Ben Basat},
  journal= {arXiv preprint arXiv:2608.06046},
  year   = {2026}
}

Comments

8 pages, 3 figures