Full-Pipeline Inference Optimization for MiMo-V2.5 Series: Pushing Hybrid SWA Efficiency to the Limit
Abstract
We present a full-pipeline inference optimization for the MiMo-V2.5 model family, which combines Hybrid Sliding Window Attention (Hybrid SWA), sparse Mixture-of-Experts (MoE), and multimodal encoders. While Hybrid SWA can ideally reduce both attention compute and KVCache storage significantly compared to Full Attention, realizing these gains in production requires substantial engineering effort. We systematically optimize the KVCache system with layerwise prefetch, SWA-aware prefix cache trees, and specialized placement strategies, achieving strict SWA storage and high cache hit rates. We further build GCache, a high-performance distributed cache infrastructure with RDMA-optimized networking, and develop a KVCache-affinity router to reduce computation while preserving load balancing. We also optimize for multimodal inputs, including GPU image preprocessing, parallel video decoding, and multimodal cache sharing. Together, these optimizations constitute the first large-scale LLM serving system in production that efficiently covers the Hybrid SWA + MoE + multimodal composite architecture.
Keywords
Cite
@article{arxiv.2607.13095,
title = {Full-Pipeline Inference Optimization for MiMo-V2.5 Series: Pushing Hybrid SWA Efficiency to the Limit},
author = {Xiaomi MiMo Team and Anqi Liu and Aoxin Ma and Bo Chen and Bo Yang and Chen Wang and Chen Zhang and Chengda Tang and Chengwei Wang and Chiheng Lou and Depeng Yan and Fuli Luo and Gang Wang and Hailin Zhang and Jiale Sun and Kang Zhou and Rui Huang and Shaohui Liu and Shen Huang and Shijie Cao and Shuaishuai Fan and Tianling Zhou and Xiangwei Deng and Xueyang Xie and Xuli Wang and Yingchun Lai and Yu Yang and Yuan Zhang and Zhen Tang and Zhonghua Deng and Zihan Jiang},
journal= {arXiv preprint arXiv:2607.13095},
year = {2026}
}
Comments
technical report