English

Before the Body Moves: Learning Anticipatory Joint Intent for Language-Conditioned Humanoid Control

Robotics 2026-05-21 v2 Computer Vision and Pattern Recognition

Abstract

Natural language is an intuitive interface for humanoid robots, yet streaming whole-body control requires control representations that are executable now and anticipatory of future physical transitions. Existing language-conditioned humanoid systems typically generate kinematic references that a low-level tracker must repair reactively, or use latent/action policies whose outputs do not explicitly encode upcoming contact changes, support transfers, and balance preparation. We propose \textbf{DAJI} (\emph{Dynamics-Aligned Joint Intent}), a hierarchical framework that learns an anticipatory joint-intent interface between language generation and closed-loop control. DAJI-Act distills a future-aware teacher into a deployable diffusion action policy through student-driven rollouts, while DAJI-Flow autoregressively generates future intent chunks from language and intent history. Experiments show that DAJI achieves strong results in anticipatory latent learning, single-instruction generation, and streaming instruction following, reaching 94.42\% rollout success on HumanML3D-style generation and 0.152 subsequence FID on BABEL.

Keywords

Cite

@article{arxiv.2605.14417,
  title  = {Before the Body Moves: Learning Anticipatory Joint Intent for Language-Conditioned Humanoid Control},
  author = {Haozhe Jia and Honglei Jin and Yuan Zhang and Youcheng Fan and Shaofeng Liang and Lei Wang and Shuxu Jin and Kuimou Yu and Zinuo Zhang and Jianfei Song and Wenshuo Chen and Yutao Yue},
  journal= {arXiv preprint arXiv:2605.14417},
  year   = {2026}
}
R2 v1 2026-07-22T07:11:41.757Z