English

HOI-PAGE: Zero-Shot Human-Object Interaction Generation with Part Affordance Guidance

Graphics 2026-05-20 v2 Computer Vision and Pattern Recognition

Abstract

We present HOI-PAGE, a new approach that prioritizes part-level affordance reasoning to generate high-fidelity 4D human-object interactions (HOIs) from text prompts in a zero-shot fashion. In contrast to prior works that focus on global, whole body-object motion synthesis, our approach explicitly reasons about the underlying part-level mechanics of interactions using large language models (LLMs). We capture this reasoning in a structured part affordance graph (PAG) representation, serving as a high-level interaction scaffolding to guide a three-stage synthesis: first, decomposing input 3D objects into semantic parts; then, generating reference HOI videos from text prompts to extract part-based motion constraints; and finally, optimizing for 4D HOI motion sequences that mimic the reference dynamics while satisfying part-level contact constraints. Extensive experiments show that our approach is flexible and capable of generating complex multi-object or multi-person interaction sequences, with significantly improved realism and text alignment for zero-shot 4D HOI generation.

Keywords

Cite

@article{arxiv.2506.07209,
  title  = {HOI-PAGE: Zero-Shot Human-Object Interaction Generation with Part Affordance Guidance},
  author = {Lei Li and Angela Dai},
  journal= {arXiv preprint arXiv:2506.07209},
  year   = {2026}
}

Comments

ICML 2026. Project page: https://craigleili.github.io/projects/hoipage/ Video: https://www.youtube.com/watch?v=gwXjOffCFyk