MinT: Managed Infrastructure for Training and Serving Millions of LLMs
Abstract
We present MindLab Toolkit (MinT), a managed infrastructure system for Low-Rank Adaptation (LoRA) post-training and online serving. MinT targets a setting where many trained policies are produced over a small number of expensive base-model deployments. Instead of materializing each policy as a merged full checkpoint, MinT keeps the base model resident and moves exported LoRA adapter revisions through rollout, update, export, evaluation, serving, and rollback, hiding distributed training, serving, scheduling, and data movement behind a service interface. MinT scales this path along three axes. Scale Up extends LoRA RL to frontier-scale dense and MoE architectures, including MLA and DSA attention paths, with training and serving validated beyond 1T total parameters. Scale Down moves only the exported LoRA adapter, which can be under 1% of base-model size in rank-1 settings; adapter-only handoff reduces the measured step by 18.3x on a 4B dense model and 2.85x on a 30B MoE, while concurrent multi-policy GRPO shortens wall time by 1.77x and 1.45x without raising peak memory. Scale Out separates durable policy addressability from CPU/GPU working sets: a tensor-parallel deployment supports 10^6-scale addressable catalogs (measured single-engine sweeps through 100K) and thousand-adapter active waves at cluster scale, with cold loading treated as scheduled service work and packed MoE LoRA tensors improving live engine loading by 8.5-8.7x. MinT thus manages million-scale LoRA policy catalogs while training and serving selected adapter revisions over shared 1T-class base models.
Cite
@article{arxiv.2605.13779,
title = {MinT: Managed Infrastructure for Training and Serving Millions of LLMs},
author = {Mind Lab and : and Song Cao and Vic Cao and Andrew Chen and Kaijie Chen and Cleon Cheng and Steven Chiang and Kaixuan Fan and Hera Feng and Huan Feng and Arthur Fu and Jun Gao and Hongquan Gu and Aaron Guan and Nolan Ho and Mutian Hong and Hailee Hou and Peixuan Hua and Charles Huang and Miles Jiang and Nora Jiang and Yuyi Jiang and Qiuyu Jin and Fancy Kong and Andrew Lei and Kyrie Lei and Alexy Li and Lucian Li and Ray Li and Theo Li and Zhihui Li and Jiayi Lin and Kairus Liu and Kieran Liu and Logan Liu and Xiang Liu and Irvine Lu and Maeve Luo and Runze Lv and Pony Ma and Verity Niu and Anson Qiu and Vincent Wang and Rio Yang and Maxwell Yao and Carrie Ye and Regis Ye and Wenlin Ye and Josh Ying and Danney Zeng and Yuhan Zhan and Anya Zhang and Di Zhang and Ruijia Zhang and Sueky Zhang and Ya Zhang and Wei Zhao and Ada Zhou and Changhai Zhou and Yuhua Zhou and Xinyue Zhu and Murphy Zhuang},
journal= {arXiv preprint arXiv:2605.13779},
year = {2026}
}
Comments
30 pages, technical report