English

Surface Keypoint Representation for Multi-Object and Articulated Human-Object Interaction Generation

Computer Vision and Pattern Recognition 2026-08-04 v1

Abstract

Daily activities require humans to coordinate whole-body motion with the motion of surrounding objects. Despite recent progress in human-object interaction (HOI) generation, most existing methods assume interactions with a single rigid object and do not extend well to scenarios involving a variable number of objects or articulated objects with diverse joint mechanisms. We propose surface keypoint trajectories as an object motion representation: for each rigid component, whether a standalone object or one part of an articulated assembly, we track a small set of non-collinear surface points over time. This representation handles multi-object coordination and diverse articulation mechanisms directly from point dynamics without requiring explicit joint-type specification. To model when and where each body region contacts each object, we introduce a spatio-temporal contact distance field that extends distance-based contact modeling to whole-body, multi-object, and articulated settings. We factorize HOI generation into three stages: generating object motions from text or waypoints, predicting the contact distance field, and synthesizing whole-body motion with contact-guided optimization. Experiments on ParaHome, HIMO, ARCTIC, and OMOMO demonstrate better or comparable performance to existing methods across single-object, multi-object, and articulated interaction settings.

Cite

@article{arxiv.2608.03158,
  title  = {Surface Keypoint Representation for Multi-Object and Articulated Human-Object Interaction Generation},
  author = {Xiaogang Peng and Zeyu Han and Zichong Meng and Yiming Xie and Jihua Zhu and Gang Hua and Huaizu Jiang},
  journal= {arXiv preprint arXiv:2608.03158},
  year   = {2026}
}

Comments

Project page: https://neu-vi.github.io/SK-HOI