中文
相关论文

相关论文: GenHOI: Towards Object-Consistent Hand-Object Inte…

200 篇论文

Video Object Grounding (VOG) is the problem of associating spatial object regions in the video to a descriptive natural language query. This is a challenging vision-language task that necessitates constructing the correct cross-modal…

多媒体 · 计算机科学 2022-08-12 Mengze Li , Tianbao Wang , Haoyu Zhang , Shengyu Zhang , Zhou Zhao , Wenqiao Zhang , Jiaxu Miao , Shiliang Pu , Fei Wu

We propose HOI Transformer to tackle human object interaction (HOI) detection in an end-to-end manner. Current approaches either decouple HOI task into separated stages of object detection and interaction classification or introduce…

计算机视觉与模式识别 · 计算机科学 2021-03-09 Cheng Zou , Bohan Wang , Yue Hu , Junqi Liu , Qian Wu , Yu Zhao , Boxun Li , Chenguang Zhang , Chi Zhang , Yichen Wei , Jian Sun

Large-scale text-to-image (T2I) diffusion models have showcased incredible capabilities in generating coherent images based on textual descriptions, enabling vast applications in content generation. While recent advancements have introduced…

计算机视觉与模式识别 · 计算机科学 2024-02-28 Jiun Tian Hoe , Xudong Jiang , Chee Seng Chan , Yap-Peng Tan , Weipeng Hu

3D hand-object interaction data is scarce due to the hardware constraints in scaling up the data collection process. In this paper, we propose HOIDiffusion for generating realistic and diverse 3D hand-object interaction data. Our model is a…

计算机视觉与模式识别 · 计算机科学 2024-03-19 Mengqi Zhang , Yang Fu , Zheng Ding , Sifei Liu , Zhuowen Tu , Xiaolong Wang

We propose G-HOP, a denoising diffusion based generative prior for hand-object interactions that allows modeling both the 3D object and a human hand, conditioned on the object category. To learn a 3D spatial diffusion model that can capture…

计算机视觉与模式识别 · 计算机科学 2024-04-19 Yufei Ye , Abhinav Gupta , Kris Kitani , Shubham Tulsiani

Hand-Object Interaction (HOI) is gaining significant attention, particularly with the creation of numerous egocentric datasets driven by AR/VR applications. However, third-person view HOI has received less attention, especially in terms of…

计算机视觉与模式识别 · 计算机科学 2024-09-17 Arya Farkhondeh , Samy Tafasca , Jean-Marc Odobez

We introduce D3D-HOI: a dataset of monocular videos with ground truth annotations of 3D object pose, shape and part motion during human-object interactions. Our dataset consists of several common articulated objects captured from diverse…

计算机视觉与模式识别 · 计算机科学 2021-08-20 Xiang Xu , Hanbyul Joo , Greg Mori , Manolis Savva

Video-based Human-Object Interaction (HOI) recognition explores the intricate dynamics between humans and objects, which are essential for a comprehensive understanding of human behavior and intentions. While previous work has made…

计算机视觉与模式识别 · 计算机科学 2024-07-24 Tanqiu Qiao , Ruochen Li , Frederick W. B. Li , Hubert P. H. Shum

We study the problem of detecting human-object interactions (HOI) in static images, defined as predicting a human and an object bounding box with an interaction class label that connects them. HOI detection is a fundamental problem in…

计算机视觉与模式识别 · 计算机科学 2018-03-02 Yu-Wei Chao , Yunfan Liu , Xieyang Liu , Huayi Zeng , Jia Deng

Livestreaming often involves interactions between streamers and objects, which is critical for understanding and regulating web content. While human-object interaction (HOI) detection has made some progress in general-purpose video…

计算机视觉与模式识别 · 计算机科学 2025-08-05 Menghui Zhang , Jing Zhang , Lin Chen , Li Zhuo

Scene graph generation (SGG) and human-object interaction (HOI) detection are two important visual tasks aiming at localising and recognising relationships between objects, and interactions between humans and objects, respectively.…

计算机视觉与模式识别 · 计算机科学 2023-11-06 Tao He , Lianli Gao , Jingkuan Song , Yuan-Fang Li

A significant gap remains between today's visual pattern recognition models and human-level visual cognition especially when it comes to few-shot learning and compositional reasoning of novel concepts. We introduce Bongard-HOI, a new visual…

计算机视觉与模式识别 · 计算机科学 2023-04-14 Huaizu Jiang , Xiaojian Ma , Weili Nie , Zhiding Yu , Yuke Zhu , Song-Chun Zhu , Anima Anandkumar

Reconstructing 3D human-object interaction (HOI) from single-view RGB images is challenging due to the absence of depth information and potential occlusions. Existing methods simply predict the body poses merely rely on network training on…

计算机视觉与模式识别 · 计算机科学 2024-07-22 Yuhang Chen , Chenxing Wang

Recent advancements in diffusion models have led to significant improvements in the generation and animation of 4D full-body human-object interactions (HOI). Nevertheless, existing methods primarily focus on SMPL-based motion generation,…

计算机视觉与模式识别 · 计算机科学 2024-10-10 Yukang Cao , Liang Pan , Kai Han , Kwan-Yee K. Wong , Ziwei Liu

Open-vocabulary human-object interaction (HOI) detection, which is concerned with the problem of detecting novel HOIs guided by natural language, is crucial for understanding human-centric scenes. However, prior zero-shot HOI detectors…

计算机视觉与模式识别 · 计算机科学 2024-04-11 Ting Lei , Shaofeng Yin , Yang Liu

The generation of anchor-style product promotion videos presents promising opportunities in e-commerce, advertising, and consumer engagement. Despite advancements in pose-guided human video generation, creating product promotion videos…

计算机视觉与模式识别 · 计算机科学 2025-06-24 Ziyi Xu , Ziyao Huang , Juan Cao , Yong Zhang , Xiaodong Cun , Qing Shuai , Yuchen Wang , Linchao Bao , Jintao Li , Fan Tang

We introduce Robowheel, a data engine that converts human hand object interaction (HOI) videos into training-ready supervision for cross morphology robotic learning. From monocular RGB or RGB-D inputs, we perform high precision HOI…

Human-Object Interaction (HOI) detection is crucial for robot-human assistance, enabling context-aware support. However, models trained on clean datasets degrade in real-world conditions due to unforeseen corruptions, leading to inaccurate…

计算机视觉与模式识别 · 计算机科学 2025-10-14 Di Wen , Kunyu Peng , Kailun Yang , Yufan Chen , Ruiping Liu , Junwei Zheng , Alina Roitberg , Danda Pani Paudel , Luc Van Gool , Rainer Stiefelhagen

We introduce HOIGPT, a token-based generative method that unifies 3D hand-object interactions (HOI) perception and generation, offering the first comprehensive solution for captioning and generating high-quality 3D HOI sequences from a…

计算机视觉与模式识别 · 计算机科学 2025-03-26 Mingzhen Huang , Fu-Jen Chu , Bugra Tekin , Kevin J Liang , Haoyu Ma , Weiyao Wang , Xingyu Chen , Pierre Gleize , Hongfei Xue , Siwei Lyu , Kris Kitani , Matt Feiszli , Hao Tang

We present HOI-PAGE, a new approach that prioritizes part-level affordance reasoning to generate high-fidelity 4D human-object interactions (HOIs) from text prompts in a zero-shot fashion. In contrast to prior works that focus on global,…

图形学 · 计算机科学 2026-05-20 Lei Li , Angela Dai