English
Related papers

Related papers: ViPLO: Vision Transformer based Pose-Conditioned S…

200 papers

Human-Object Interaction (HOI) detection plays a core role in activity understanding. Though recent two/one-stage methods have achieved impressive results, as an essential step, discovering interactive human-object pairs remains…

Computer Vision and Pattern Recognition · Computer Science 2022-04-19 Xinpeng Liu , Yong-Lu Li , Xiaoqian Wu , Yu-Wing Tai , Cewu Lu , Chi-Keung Tang

Recent state-of-the-art methods for HOI detection typically build on transformer architectures with two decoder branches, one for human-object pair detection and the other for interaction classification. Such disentangled transformers,…

Computer Vision and Pattern Recognition · Computer Science 2023-04-12 Sanghyun Kim , Deunsol Jung , Minsu Cho

Human-Object Interaction (HOI) aims to identify the pairs of humans and objects in images and to recognize their relationships, ultimately forming $\langle human, object, verb \rangle$ triplets. Under default settings, HOI performance is…

Computer Vision and Pattern Recognition · Computer Science 2024-08-13 Chaoyi Ai

We propose a single-stage Human-Object Interaction (HOI) detection method that has outperformed all existing methods on HICO-DET dataset at 37 fps on a single Titan XP GPU. It is the first real-time HOI detection method. Conventional HOI…

Computer Vision and Pattern Recognition · Computer Science 2020-03-26 Yue Liao , Si Liu , Fei Wang , Yanjie Chen , Chen Qian , Jiashi Feng

Understanding and recognizing human-object interaction (HOI) is a pivotal application in AR/VR and robotics. Recent open-vocabulary HOI detection approaches depend exclusively on large language models for richer textual prompts, neglecting…

Computer Vision and Pattern Recognition · Computer Science 2025-08-18 Zhenhao Zhang , Hanqing Wang , Xiangyu Zeng , Ziyu Cheng , Jiaxin Liu , Haoyu Yan , Zhirui Liu , Kaiyang Ji , Tianxiang Gui , Ke Hu , Kangyi Chen , Yahao Fan , Mokai Pan

Monocular vertex-level human-scene contact prediction is a fundamental capability for interactive systems such as assistive monitoring, embodied AI, and rehabilitation analysis. In this work, we study this task jointly with single-image 3D…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Xiaojian Lin , Yaomin Shen , Junyuan Ma , Yujie Sun , Chengqing Bu , Wenxin Zhang , Zongzheng Zhang , Hao Fei , Lei Jin , Hao Zhao

Human-Object Interaction (HOI) detection aims to localize human-object pairs and classify their interactions from a single image, a task that demands strong visual understanding and nuanced contextual reasoning. Recent approaches have…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Soo Won Seo , KyungChae Lee , Hyungchan Cho , Taein Son , Nam Ik Cho , Jun Won Choi

Synthesizing realistic human-object interaction (HOI) is essential for 3D computer vision and robotics, underpinning animation and embodied control. Existing approaches often require manually specified intermediate waypoints and place all…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Hwanhee Jung , Seunggwan Lee , Jeongyoon Yoon , SeungHyeon Kim , Giljoo Nam , Qixing Huang , Sangpil Kim

In this paper, we propose a new instance-level human-object interaction detection task on videos called ST-HOID, which aims to distinguish fine-grained human-object interactions (HOIs) and the trajectories of subjects and objects. It is…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Xu Sun , Yunqing He , Tongwei Ren , Gangshan Wu

Joint forecasting of human trajectory and pose dynamics is a fundamental building block of various applications ranging from robotics and autonomous driving to surveillance systems. Predicting body dynamics requires capturing subtle…

Computer Vision and Pattern Recognition · Computer Science 2022-07-07 Vida Adeli , Mahsa Ehsanpour , Ian Reid , Juan Carlos Niebles , Silvio Savarese , Ehsan Adeli , Hamid Rezatofighi

Human pose estimation aims to figure out the keypoints of all people in different scenes. Current approaches still face some challenges despite promising results. Existing top-down methods deal with a single person individually, without the…

Computer Vision and Pattern Recognition · Computer Science 2022-09-07 Shuaitao Zhao , Kun Liu , Yuhang Huang , Qian Bao , Dan Zeng , Wu Liu

Video-based human-object interaction (HOI) understanding requires both detecting ongoing interactions and anticipating their future evolution. However, existing methods usually treat anticipation as a downstream forecasting task built on…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Yuanhao Luo , Di Wen , Kunyu Peng , Ruiping Liu , Junwei Zheng , Yufan Chen , Jiale Wei , Rainer Stiefelhage

Local Transformer-based classification models have recently achieved promising results with relatively low computational costs. However, the effect of aggregating spatial global information of local Transformer-based architecture is not…

Computer Vision and Pattern Recognition · Computer Science 2022-02-01 Krushi Patel , Andres M. Bur , Fengjun Li , Guanghui Wang

Human-Object Interaction (HOI) detection devotes to learn how humans interact with surrounding objects. Latest end-to-end HOI detectors are short of relation reasoning, which leads to inability to learn HOI-specific interactive semantics…

Computer Vision and Pattern Recognition · Computer Science 2021-05-03 Dongming Yang , Yuexian Zou , Can Zhang , Meng Cao , Jie Chen

GRAP-MOT is a new approach for solving the person MOT problem dedicated to videos of closed areas with overlapping multi-camera views, where person occlusion frequently occurs. Our novel graph-weighted solution updates a person's…

Computer Vision and Pattern Recognition · Computer Science 2025-10-27 Marek Socha , Michał Marczyk , Aleksander Kempski , Michał Cogiel , Paweł Foszner , Radosław Zawiski , Michał Staniszewski

Human-Object Interaction (HOI) detection lies at the core of action understanding. Besides 2D information such as human/object appearance and locations, 3D pose is also usually utilized in HOI learning since its view-independence. However,…

Computer Vision and Pattern Recognition · Computer Science 2020-05-22 Yong-Lu Li , Xinpeng Liu , Han Lu , Shiyi Wang , Junqi Liu , Jiefeng Li , Cewu Lu

Hand-Object Interaction (HOI) generation has significant application potential. However, current 3D HOI motion generation approaches heavily rely on predefined 3D object models and lab-captured motion data, limiting generalization…

Computer Vision and Pattern Recognition · Computer Science 2025-06-06 Lingwei Dang , Ruizhi Shao , Hongwen Zhang , Wei Min , Yebin Liu , Qingyao Wu

We study the problem of precisely swapping objects in videos, with a focus on those interacted with by hands, given one user-provided reference object image. Despite the great advancements that diffusion models have made in video editing…

Computer Vision and Pattern Recognition · Computer Science 2024-11-12 Zihui Xue , Mi Luo , Changan Chen , Kristen Grauman

Human-Object Interaction (HOI) recognition in videos requires understanding both visual patterns and geometric relationships as they evolve over time. Visual and geometric features offer complementary strengths. Visual features capture…

Computer Vision and Pattern Recognition · Computer Science 2025-06-06 Tanqiu Qiao , Ruochen Li , Frederick W. B. Li , Yoshiki Kubotani , Shigeo Morishima , Hubert P. H. Shum

Human-Object Interaction (HOI) detection is a task to localize humans and objects in an image and predict the interactions in human-object pairs. In real-world scenarios, HOI detection models need systematic generalization, i.e.,…

Computer Vision and Pattern Recognition · Computer Science 2024-04-15 Kentaro Takemoto , Moyuru Yamada , Tomotake Sasaki , Hisanao Akima