English
Related papers

Related papers: RR-Net: Injecting Interactive Semantics in Human-O…

200 papers

Large Vision Language Models (VLMs) are now the de facto state-of-the-art for a number of tasks including visual question answering, recognising objects, and spatial referral. In this work, we propose the HOI-Ref task for egocentric images…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Siddhant Bansal , Michael Wray , Dima Damen

Real-world scenes often feature multiple humans interacting with multiple objects in ways that are causal, goal-oriented, or cooperative. Yet existing 3D human-object interaction (HOI) benchmarks consider only a fraction of these complex…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Kaen Kogashi , Anoop Cherian , Meng-Yu Jennifer Kuo

This paper revisits human-object interaction (HOI) recognition at image level without using supervisions of object location and human pose. We name it detection-free HOI recognition, in contrast to the existing detection-supervised…

Computer Vision and Pattern Recognition · Computer Science 2021-07-29 Ying Jin , Yinpeng Chen , Lijuan Wang , Jianfeng Wang , Pei Yu , Zicheng Liu , Jenq-Neng Hwang

The task of Human-Object conTact (HOT) detection involves identifying the specific areas of the human body that are touching objects. Nevertheless, current models are restricted to just one type of image, often leading to too much…

Computer Vision and Pattern Recognition · Computer Science 2025-07-24 Yuxiao Wang , Yu Lei , Zhenao Wei , Weiying Xue , Xinyu Jiang , Nan Zhuang , Qi Liu

Video-based Human-Object Interaction (HOI) recognition explores the intricate dynamics between humans and objects, which are essential for a comprehensive understanding of human behavior and intentions. While previous work has made…

Computer Vision and Pattern Recognition · Computer Science 2024-07-24 Tanqiu Qiao , Ruochen Li , Frederick W. B. Li , Hubert P. H. Shum

Zero-shot human-object interaction (HOI) detection remains a challenging task, particularly in generalizing to unseen actions. Existing methods address this challenge by tapping Vision-Language Models (VLMs) to access knowledge beyond the…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Qinqian Lei , Bo Wang , Robby T. Tan

Synthesizing human--object interaction (HOI) videos has broad practical value in e-commerce, digital advertising, and virtual marketing. However, current diffusion models, despite their photorealistic rendering capability, still frequently…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Xiangyang Luo , Xiaozhe Xin , Tao Feng , Xu Guo , Meiguang Jin , Junfeng Ma

We propose CG-HOI, the first method to address the task of generating dynamic 3D human-object interactions (HOIs) from text. We model the motion of both human and object in an interdependent fashion, as semantically rich human motion rarely…

Computer Vision and Pattern Recognition · Computer Science 2024-05-20 Christian Diller , Angela Dai

Two-stage methods have dominated Human-Object Interaction (HOI) detection for several years. Recently, one-stage HOI detection methods have become popular. In this paper, we aim to explore the essential pros and cons of two-stage and…

Computer Vision and Pattern Recognition · Computer Science 2021-10-14 Aixi Zhang , Yue Liao , Si Liu , Miao Lu , Yongliang Wang , Chen Gao , Xiaobo Li

We present HOIDiNi, a text-driven diffusion framework for synthesizing realistic and plausible human-object interaction (HOI). HOI generation is extremely challenging since it induces strict contact accuracies alongside a diverse motion…

Computer Vision and Pattern Recognition · Computer Science 2025-10-22 Roey Ron , Guy Tevet , Haim Sawdayee , Amit H. Bermano

Objects are entities we act upon, where the functionality of an object is determined by how we interact with it. In this work we propose a Dual Attention Network model which reasons about human-object interactions. The dual-attentional…

Computer Vision and Pattern Recognition · Computer Science 2019-09-12 Tete Xiao , Quanfu Fan , Dan Gutfreund , Mathew Monfort , Aude Oliva , Bolei Zhou

We propose ChainHOI, a novel approach for text-driven human-object interaction (HOI) generation that explicitly models interactions at both the joint and kinetic chain levels. Unlike existing methods that implicitly model interactions using…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Ling-An Zeng , Guohong Huang , Yi-Lin Wei , Shengbo Gu , Yu-Ming Tang , Jingke Meng , Wei-Shi Zheng

Human-object contact (HOT) is designed to accurately identify the areas where humans and objects come into contact. Current methods frequently fail to account for scenarios where objects are frequently blocking the view, resulting in…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Yuxiao Wang , Wenpeng Neng , Zhenao Wei , Yu Lei , Weiying Xue , Nan Zhuang , Yanwu Xu , Xinyu Jiang , Qi Liu

The Human-Object Interaction (HOI) task explores the dynamic interactions between humans and objects in physical environments, providing essential biomechanical and cognitive-behavioral foundations for fields such as robotics, virtual…

Computer Vision and Pattern Recognition · Computer Science 2025-08-11 Ruiyan Wang , Lin Zuo , Zonghao Lin , Qiang Wang , Zhengxue Cheng , Rong Xie , Jun Ling , Li Song

Recent developments in transformer models for visual data have led to significant improvements in recognition and detection tasks. In particular, using learnable queries in place of region proposals has given rise to a new class of…

Computer Vision and Pattern Recognition · Computer Science 2022-03-29 Frederic Z. Zhang , Dylan Campbell , Stephen Gould

Humans naturally interact with both others and the surrounding multiple objects, engaging in various social activities. However, recent advances in modeling human-object interactions mostly focus on perceiving isolated individuals and…

Computer Vision and Pattern Recognition · Computer Science 2024-04-03 Juze Zhang , Jingyan Zhang , Zining Song , Zhanhe Shi , Chengfeng Zhao , Ye Shi , Jingyi Yu , Lan Xu , Jingya Wang

Amodal completion, which is the process of inferring the full appearance of objects despite partial occlusions, is crucial for understanding complex human-object interactions (HOI) in computer vision and robotics. Existing methods, such as…

Computer Vision and Pattern Recognition · Computer Science 2025-08-04 Seunggeun Chi , Enna Sachdeva , Pin-Hao Huang , Kwonjoon Lee

Human actions often involve complex interactions across several inter-related objects in the scene. However, existing approaches to fine-grained video understanding or visual relationship detection often rely on single object representation…

Computer Vision and Pattern Recognition · Computer Science 2018-03-22 Chih-Yao Ma , Asim Kadav , Iain Melvin , Zsolt Kira , Ghassan AlRegib , Hans Peter Graf

Although it is well believed for years that modeling relations between objects would help object recognition, there has not been evidence that the idea is working in the deep learning era. All state-of-the-art object detection systems still…

Computer Vision and Pattern Recognition · Computer Science 2018-06-15 Han Hu , Jiayuan Gu , Zheng Zhang , Jifeng Dai , Yichen Wei

Interaction is one of the core abilities of humanoid robots. However, most existing frameworks focus on non-interactive whole-body control, which limits their practical applicability. In this work, we develop InterReal, a unified…

Robotics · Computer Science 2026-03-10 Dayang Liang , Yuhang Lin , Xinzhe Liu , Jiyuan Shi , Yunlong Liu , Chenjia Bai
‹ Prev 1 8 9 10 Next ›