ROCS: Request-Oriented Compute Sharing for Efficient Large-Scale Recommendation
Abstract
Modern recommendation models gain prediction quality by scaling feature-interaction and sequence modules, but production cost constraints cap how far systems can scale. In this work, we propose Request-Oriented Compute Sharing (ROCS), a modeling and inference paradigm that exploits a unique property of recommendation inference: each user request is evaluated against many candidates, while request-side features are shared across candidates. ROCS defers request-candidate interactions as late as possible, isolates candidate-dependent representations, and evaluates substantial portions of the model once per request rather than once per candidate, significantly improving inference efficiency while maintaining or improving prediction quality. To realize this paradigm, we develop Generalized Layer Masking (GLM) to enforce candidate isolation in feature-interaction architectures, and Deep Cross Attention (DCA) to extend request-oriented sharing to sequence architectures. To support efficient GPU deployment, we co-design In-Kernel Broadcast Optimization (IKBO) that significantly accelerates ROCS model execution. Experiments on public benchmarks show that ROCS consistently improves the quality-efficiency tradeoff across recommendation backbones. On production-scale workloads, ROCS achieves up to a 3x QPS improvement on retrieval models without quality degradation and a 0.5% relative LogLoss improvement with a 50% QPS gain on a short-form video ranking model. ROCS has been deployed across large-scale recommendation systems spanning ads and organic surfaces, retrieval and ranking stages, and more than two orders of magnitude in inference complexity, delivering significant online gains at reduced infrastructure cost.
Cite
@article{arxiv.2607.27744,
title = {ROCS: Request-Oriented Compute Sharing for Efficient Large-Scale Recommendation},
author = {Yuxin Chen and Liang Luo and Buyun Zhang and Jian Jiao and Boda Li and Haoyu Wang and Tongyi Tang and Ao Cai and Zijian Shen and Zhengkai Zhang and Wenyi Xie and Ryan Dick and Han Liu and Neng Shi and Bin Yu and Jianbo Xiao and Shuyao Bi and Hongtao Yu and Yuanwei Fang and Zhuoran Zhao and Sijia Chen and Yang Chen and Shuqi Yang and Qianru Li and Zikun Liu and Wei Ling and Sihan Zeng and Longhao Jin and Jiaxin Lu and Yinbin Ma and Jiawei Li and Yichen Ruan and Yong Ler Lee and Birmingham Guan and Zijian Li and Jianbo Sun and Zhengyu Zhang and Zeliang Chen and Xiaohan Wei and Yuchen Hao and GP Musumeci and Venkatesh Ranganathan and Yantao Yao and Chunqiang Tang and Wenlin Chen and Santanu Kolay and Ellie Dingqiao Wen},
journal= {arXiv preprint arXiv:2607.27744},
year = {2026}
}