RelayGR: Scaling Long-Sequence Generative Recommendation via Cross-Stage Relay-Race Inference
Abstract
Real-time recommender systems execute multi-stage cascades (retrieval, pre-processing, fine-grained ranking) under strict tail-latency SLOs, leaving only tens of milliseconds for ranking. Generative recommendation (GR) models can improve quality by consuming long user-behavior sequences, but in production their online sequence length is tightly capped by the ranking-stage P99 budget. We observe that the majority of GR tokens encode user behaviors that are independent of the item candidates, suggesting an opportunity to pre-infer a user-behavior prefix once and reuse it during ranking rather than recomputing it on the critical path. Realizing this idea at industrial scale is non-trivial: the prefix cache must survive across multiple pipeline stages before the final ranking instance is determined, the user population implies cache footprints far beyond a single device, and indiscriminate pre-inference would overload shared resources under high QPS. We present RelayGR, a production system that enables in-HBM relay-race inference for GR. RelayGR selectively pre-infers long-term user prefixes, keeps their KV caches resident in HBM over the request lifecycle, and ensures the subsequent ranking can consume them without remote fetches. RelayGR combines three techniques: 1) a sequence-aware trigger that admits only at-risk requests under a bounded cache footprint and pre-inference load, 2) an affinity-aware router that co-locates cache production and consumption by routing both the auxiliary pre-infer signal and the ranking request to the same instance, and 3) a memory-aware expander that uses server-local DRAM to capture short-term cross-request reuse while avoiding redundant reloads. We implement RelayGR on Huawei Ascend NPUs and evaluate it with real queries. Under a fixed P99 SLO, RelayGR supports up to 1.5 longer sequences and improves SLO-compliant throughput by up to 3.6.
Keywords
Cite
@article{arxiv.2601.01712,
title = {RelayGR: Scaling Long-Sequence Generative Recommendation via Cross-Stage Relay-Race Inference},
author = {Jiarui Wang and Huichao Chai and Yuanhang Zhang and Zongjin Zhou and Wei Guo and Xingkun Yang and Qiang Tang and Bo Pan and Jiawei Zhu and Ke Cheng and Yuting Yan and Shulan Wang and Yingjie Zhu and Zhengfan Yuan and Jiaqi Huang and Yuhan Zhang and Xiaosong Sun and Zhinan Zhang and Hong Zhu and Yongsheng Zhang and Tiantian Dong and Zhong Xiao and Deliang Liu and Chengzhou Lu and Yuan Sun and Zhiyuan Chen and Xinming Han and Zaizhu Liu and Yaoyuan Wang and Ziyang Zhang and Yong Liu and Jinxin Xu and Yajing Sun and Zhoujun Yu and Wenting Zhou and Qidong Zhang and Zhengyong Zhang and Zhonghai Gu and Yibo Jin and Yongxiang Feng and Pengfei Zuo},
journal= {arXiv preprint arXiv:2601.01712},
year = {2026}
}