English

An Engineering Journey Training Large Language Models at Scale on Alps: The Apertus Experience

Distributed, Parallel, and Cluster Computing 2026-04-16 v2

Abstract

Large Language Models (LLMs) have surged as a transformative technology for science and society, prompting governments worldwide to pursue sovereign AI capabilities that ensure data compliance and cultural representation. However, the associated capital costs and engineering complexity required to train these models have largely restricted such capabilities to the private sector, leaving a significant gap for public institutions. This paper details the engineering journey behind training Apertus, a fully open multilingual foundation model, on the Alps supercomputer. Representing a first-of-its-kind achievement for academia at the 70B parameter scale, we successfully deployed a massive pre-training campaign on one of Europe's largest systems for open science, powered by NVIDIA GH200 Grace Hopper Superchips. We detail the challenges encountered in readying HPC infrastructure for training AI models, from overcoming storage bottlenecks to stabilizing large-scale interconnects, and the lessons learned in transforming a supercomputer into a resilient software-defined Machine Learning Platform. Finally, we discuss the post-training requirements and evolution of our Machine Learning platform, outlining how this initial release lays the groundwork for a sustained, iterative operational capability, in particular for fine tuning foundation models, that extends well beyond a single model training run.

Keywords

Cite

@article{arxiv.2604.12973,
  title  = {An Engineering Journey Training Large Language Models at Scale on Alps: The Apertus Experience},
  author = {Jonathan Coles and Stefano Schuppli and Lukas Drescher and Fawzi Roberto Mohamed and Elia Palme and Henrique Mendonça and Miguel Gila and Mark Klein and Maxime Martinasso and Joost VandeVondele and Torsten Hoefler and Thomas Schulthess and Josh Romero and Igor Gorodetsky and Ryan Hankins and Isa Wazirzada and Martin Jaggi and Antoine Bosselut and Imanol Schlag and Antoni-Joan Solergibert i Llaquet and Alejandro Hernández Cano and Theofilos Ioannis Manitaras and Nicholas John Browning},
  journal= {arXiv preprint arXiv:2604.12973},
  year   = {2026}
}
R2 v1 2026-07-01T12:09:14.896Z