English

Afford-VLA: Action-Aligned Visual Planning via Internalized Affordance

Robotics 2026-05-26 v1

Abstract

Vision-language-action (VLA) models have shown strong potential for generalist robot manipulation, yet they remain limited by insufficient spatial reasoning, particularly in determining where to interact in complex visual scenes. While recent efforts introduce various forms of visual planning to address this issue, existing approaches either rely on global geometric cues, symbolic intermediate representations, or externally generated visual signals, which are often weakly coupled with downstream action prediction. In this work, we revisit visual planning in VLA systems and argue that effective planning should be local, visually grounded, internally generated, and directly aligned with action. Based on this insight, we propose Afford-VLA, a unified framework that internalizes task-conditioned affordance as an explicit visual planning interface within VLA models. Concretely, we introduce learnable <AFF> tokens to query task-relevant interaction regions, decode affordance masks from multimodal features, and convert them into compact embeddings that directly condition action generation. This design enables affordance to be both generated and utilized within the VLA, forming a tightly coupled perception-action pathway. To further support this integration, we adopt a training strategy that allows the affordance pathway to be jointly optimized with action prediction, improving its effectiveness for downstream control. We evaluate our method on multiple simulation benchmarks, including LIBERO, LIBERO-Plus, and SimplerEnv, achieving consistent state-of-the-art performance, along with strong real-world results. These findings demonstrate that internalizing affordance as action-aligned visual planning provides a powerful paradigm for improving VLA systems.

Keywords

Cite

@article{arxiv.2605.24203,
  title  = {Afford-VLA: Action-Aligned Visual Planning via Internalized Affordance},
  author = {Runze Wang and Yuqian Fu and Yu Li and Tao Lin and Tianwen Qian and Mohamed Elhoseiny and Bo Zhao and Yanwei Fu and Yu-Gang Jiang and Xiangyang Xue},
  journal= {arXiv preprint arXiv:2605.24203},
  year   = {2026}
}

Comments

20 pages

R2 v1 2026-07-22T07:29:26.055Z