Mixture of Horizons in Action Chunking
Abstract
Vision-language-action (VLA) models have shown remarkable capabilities in robotic manipulation, but their performance is sensitive to the used during training, termed . Our empirical study reveals an inherent trade-off: longer horizons provide stronger global foresight but degrade fine-grained accuracy, while shorter ones sharpen local control yet struggle on long-term tasks, implying fixed choice of single horizons being suboptimal. To mitigate the trade-off, we propose a strategy. MoH rearranges the action chunk into several segments with different horizons, processes them in parallel with a shared action transformer, and fuses outputs with a light linear gate. It has three appealing benefits. 1) MoH exploits long-term foresight and short-term precision jointly within a single model, improving both performance and generalizability to complex tasks. 2) MoH is plug-and-play for full-attention action modules with minimal training or inference overhead. 3) MoH enables dynamic inference with adaptive horizons, which selects stable actions through cross-horizon consensus, achieving 2.5 higher throughput than baselines while preserving superior performance. Extensive experiments over flow-based policies , , and one-step regression policy demonstrate that MoH yields consistent and significant gains on both simulations and real-world tasks. Notably, under mixed-task setting, with MoH reaches a new state-of-the-art with 99 average success rate on LIBERO after only training iterations. Project page: https://github.com/Timsty1/MixtureOfHorizons
Keywords
Cite
@article{arxiv.2511.19433,
title = {Mixture of Horizons in Action Chunking},
author = {Dong Jing and Gang Wang and Jiaqi Liu and Weiliang Tang and Zelong Sun and Yunchao Yao and Zhenyu Wei and Yunhui Liu and Zhiwu Lu and Mingyu Ding},
journal= {arXiv preprint arXiv:2511.19433},
year = {2025}
}
Comments
15 pages, 14 figures