English

Discovering A Variety of Objects in Spatio-Temporal Human-Object Interactions

Computer Vision and Pattern Recognition 2022-11-21 v2 Artificial Intelligence Machine Learning

Abstract

Spatio-temporal Human-Object Interaction (ST-HOI) detection aims at detecting HOIs from videos, which is crucial for activity understanding. In daily HOIs, humans often interact with a variety of objects, e.g., holding and touching dozens of household items in cleaning. However, existing whole body-object interaction video benchmarks usually provide limited object classes. Here, we introduce a new benchmark based on AVA: Discovering Interacted Objects (DIO) including 51 interactions and 1,000+ objects. Accordingly, an ST-HOI learning task is proposed expecting vision systems to track human actors, detect interactions and simultaneously discover interacted objects. Even though today's detectors/trackers excel in object detection/tracking tasks, they perform unsatisfied to localize diverse/unseen objects in DIO. This profoundly reveals the limitation of current vision systems and poses a great challenge. Thus, how to leverage spatio-temporal cues to address object discovery is explored, and a Hierarchical Probe Network (HPN) is devised to discover interacted objects utilizing hierarchical spatio-temporal human/context cues. In extensive experiments, HPN demonstrates impressive performance. Data and code are available at https://github.com/DirtyHarryLYL/HAKE-AVA.

Keywords

Cite

@article{arxiv.2211.07501,
  title  = {Discovering A Variety of Objects in Spatio-Temporal Human-Object Interactions},
  author = {Yong-Lu Li and Hongwei Fan and Zuoyu Qiu and Yiming Dou and Liang Xu and Hao-Shu Fang and Peiyang Guo and Haisheng Su and Dongliang Wang and Wei Wu and Cewu Lu},
  journal= {arXiv preprint arXiv:2211.07501},
  year   = {2022}
}

Comments

Techniqual report. A part of the HAKE project. Project: https://github.com/DirtyHarryLYL/HAKE-AVA

R2 v1 2026-06-28T05:49:23.131Z