English
Related papers

Related papers: HOI-Swap: Swapping Objects in Videos with Hand-Obj…

200 papers

While large-scale human motion capture datasets have advanced human motion generation, modeling and generating dynamic 3D human-object interactions (HOIs) remain challenging due to dataset limitations. Existing datasets often lack…

Computer Vision and Pattern Recognition · Computer Science 2025-09-12 Sirui Xu , Dongting Li , Yucheng Zhang , Xiyan Xu , Qi Long , Ziyin Wang , Yunzhi Lu , Shuchang Dong , Hezi Jiang , Akshat Gupta , Yu-Xiong Wang , Liang-Yan Gui

When humans grasp an object, they naturally form trajectories in their minds to manipulate it for specific tasks. Modeling hand-object interaction priors holds significant potential to advance robotic and embodied AI systems in learning to…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Alexey Gavryushin , Alexandros Delitzas , Luc Van Gool , Marc Pollefeys , Kaichun Mo , Xi Wang

Human-Object Interaction (HOI) detection has seen substantial advances in recent years. However, existing works focus on the standard setting with ideal images and natural distribution, far from practical scenarios with inevitable…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Chi Xie , Shuang Liang , Jie Li , Feng Zhu , Rui Zhao , Yichen Wei , Shengjie Zhao

We present a diffusion-based video editing framework, namely DiffusionAtlas, which can achieve both frame consistency and high fidelity in editing video object appearance. Despite the success in image editing, diffusion models still…

Computer Vision and Pattern Recognition · Computer Science 2023-12-08 Shao-Yu Chang , Hwann-Tzong Chen , Tyng-Luh Liu

Learning how to interact with objects is an important step towards embodied visual intelligence, but existing techniques suffer from heavy supervision or sensing requirements. We propose an approach to learn human-object interaction…

Computer Vision and Pattern Recognition · Computer Science 2019-06-06 Tushar Nagarajan , Christoph Feichtenhofer , Kristen Grauman

Diffusion models revolutionize image generation by leveraging natural language to guide the creation of multimedia content. Despite significant advancements in such generative models, challenges persist in depicting detailed human-object…

Generative image editing has recently witnessed extremely fast-paced growth. Some works use high-level conditioning such as text, while others use low-level conditioning. Nevertheless, most of them lack fine-grained control over the…

Computer Vision and Pattern Recognition · Computer Science 2024-04-10 Vidit Goel , Elia Peruzzo , Yifan Jiang , Dejia Xu , Xingqian Xu , Nicu Sebe , Trevor Darrell , Zhangyang Wang , Humphrey Shi

Recent advances in 3D human-aware generation have made significant progress. However, existing methods still struggle with generating novel Human Object Interaction (HOI) from text, particularly for open-set objects. We identify three main…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Jinlu Zhang , Yixin Chen , Zan Wang , Jie Yang , Yizhou Wang , Siyuan Huang

Recent years have seen significant progress in human image generation, particularly with the advancements in diffusion models. However, existing diffusion methods encounter challenges when producing consistent hand anatomy and the generated…

Computer Vision and Pattern Recognition · Computer Science 2024-05-01 Anton Pelykh , Ozge Mercanoglu Sincan , Richard Bowden

Open-vocabulary human-object interaction (HOI) detection requires recognizing interaction phrases that may not appear as annotated categories during training. Recent vision-language HOI detectors improve semantic transfer by matching…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Minh Anh Nguyen , Quang Huy Tran , Bao Ngoc Le , SuiYang Guang , Tuan Kiet Pham , Linh Chi Vo

Current Virtual Reality systems are designed for interaction under visual control. Using built-in cameras, headsets track the user's hands or hand-held controllers while they are inside the field of view. Current systems thus ignore the…

Human-Computer Interaction · Computer Science 2023-05-02 Paul Streli , Rayan Armani , Yi Fei Cheng , Christian Holz

This paper proposes an interaction reasoning network for modelling spatio-temporal relationships between hands and objects in video. The proposed interaction unit utilises a Transformer module to reason about each acting hand, and its…

Computer Vision and Pattern Recognition · Computer Science 2022-01-14 Jian Ma , Dima Damen

We address the task of multi-view image editing from sparse input views, where the inputs can be seen as a mix of images capturing the scene from different viewpoints. The goal is to modify the scene according to a textual instruction while…

Computer Vision and Pattern Recognition · Computer Science 2025-11-20 Daniel Gilo , Or Litany

Human-Object Interaction (HOI) detection devotes to learn how humans interact with surrounding objects via inferring triplets of < human, verb, object >. However, recent HOI detection methods mostly rely on additional annotations (e.g.,…

Computer Vision and Pattern Recognition · Computer Science 2020-07-15 Dongming Yang , Yuexian Zou

To address key limitations in human-object interaction (HOI) video generation -- specifically the reliance on curated motion data, limited generalization to novel objects/scenarios, and restricted accessibility -- we introduce…

Computer Vision and Pattern Recognition · Computer Science 2025-06-11 Ziyao Huang , Zixiang Zhou , Juan Cao , Yifeng Ma , Yi Chen , Zejing Rao , Zhiyong Xu , Hongmei Wang , Qin Lin , Yuan Zhou , Qinglin Lu , Fan Tang

Image-conditioned Video diffusion models achieve impressive visual realism but often suffer from weakened motion fidelity, e.g., reduced motion dynamics or degraded long-term temporal coherence, especially after fine-tuning. We study the…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Xi Ye , Wenjia Yang , Yangyang Xu , Xiaoyang Liu , Duo Su , Mengfei Xia , Jun Zhu

Recent progress in diffusion-based video editing has shown remarkable potential for practical applications. However, these methods remain prohibitively expensive and challenging to deploy on mobile devices. In this study, we introduce a…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Adil Karjauv , Noor Fathima , Ioannis Lelekas , Fatih Porikli , Amir Ghodrati , Amirhossein Habibian

Human-object interaction (HOI) detection has seen advancements with Vision Language Models (VLMs), but these methods often depend on extensive manual annotations. Vision Large Language Models (VLLMs) can inherently recognize and reason…

Computer Vision and Pattern Recognition · Computer Science 2024-10-22 Jianjun Gao , Chen Cai , Ruoyu Wang , Wenyang Liu , Kim-Hui Yap , Kratika Garg , Boon-Siew Han

Human-Object Interaction (HOI) detection is the task of identifying a set of <human, object, interaction> triplets from an image. Recent work proposed transformer encoder-decoder architectures that successfully eliminated the need for many…

Computer Vision and Pattern Recognition · Computer Science 2022-03-29 Bumsoo Kim , Jonghwan Mun , Kyoung-Woon On , Minchul Shin , Junhyun Lee , Eun-Sol Kim

Human-Object Interaction (HOI) detection is a fundamental visual task aiming at localizing and recognizing interactions between humans and objects. Existing works focus on the visual and linguistic features of humans and objects. However,…

Computer Vision and Pattern Recognition · Computer Science 2021-08-20 Tao He , Lianli Gao , Jingkuan Song , Yuan-Fang Li