English

What to look at and where: Semantic and Spatial Refined Transformer for detecting human-object interactions

Computer Vision and Pattern Recognition 2022-05-27 v2

Abstract

We propose a novel one-stage Transformer-based semantic and spatial refined transformer (SSRT) to solve the Human-Object Interaction detection task, which requires to localize humans and objects, and predicts their interactions. Differently from previous Transformer-based HOI approaches, which mostly focus at improving the design of the decoder outputs for the final detection, SSRT introduces two new modules to help select the most relevant object-action pairs within an image and refine the queries' representation using rich semantic and spatial features. These enhancements lead to state-of-the-art results on the two most popular HOI benchmarks: V-COCO and HICO-DET.

Keywords

Cite

@article{arxiv.2204.00746,
  title  = {What to look at and where: Semantic and Spatial Refined Transformer for detecting human-object interactions},
  author = {A S M Iftekhar and Hao Chen and Kaustav Kundu and Xinyu Li and Joseph Tighe and Davide Modolo},
  journal= {arXiv preprint arXiv:2204.00746},
  year   = {2022}
}

Comments

CVPR 2022 Oral