English

IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation

Robotics 2026-05-15 v1 Artificial Intelligence Computation and Language Computer Vision and Pattern Recognition

Abstract

Robot imitation data are often multimodal: similar visual-language observations may be followed by different action chunks because human demonstrators act with different short-horizon intents, task phases, or recent context. Existing frame-conditioned VLA policies infer each chunk from the current observation and instruction alone, so under partial observability they may resample different intents across adjacent replanning steps, leading to inter-chunk conflict and unstable execution. We introduce IntentVLA, a history-conditioned VLA framework that encodes recent visual observations into a compact short-horizon intent representation and uses it to condition chunk generation. We further introduce AliasBench, a 12-task ambiguity-aware benchmark on RoboTwin2 with matched training data and evaluation environments that isolate short-horizon observation aliasing. Across AliasBench, SimplerEnv, LIBERO, and RoboCasa, IntentVLA improves rollout stability and outperforms strong VLA baselines

Keywords

Cite

@article{arxiv.2605.14712,
  title  = {IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation},
  author = {Shijie Lian and Bin Yu and Xiaopeng Lin and Zhaolong Shen and Laurence Tianruo Yang and Yurun Jin and Haishan Liu and Changti Wu and Hang Yuan and Cong Huang and Kai Chen},
  journal= {arXiv preprint arXiv:2605.14712},
  year   = {2026}
}

Comments

Code can be found in https://github.com/ZGC-EmbodyAI/IntentVLA

R2 v1 2026-07-22T07:12:10.974Z