English
Related papers

Related papers: FlowHOI: Flow-based Semantics-Grounded Generation …

200 papers

Vision-Language-Action (VLA) models have emerged as a unified paradigm for robotic perception and control, enabling emergent generalization and long-horizon task execution. However, their deployment in dynamic, real-world environments is…

Artificial Intelligence · Computer Science 2025-12-24 Yuntao Dai , Hang Gu , Teng Wang , Qianyu Cheng , Yifei Zheng , Zhiyong Qiu , Lei Gong , Wenqi Lou , Xuehai Zhou

Recent advancements in diffusion models have led to significant improvements in the generation and animation of 4D full-body human-object interactions (HOI). Nevertheless, existing methods primarily focus on SMPL-based motion generation,…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Yukang Cao , Liang Pan , Kai Han , Kwan-Yee K. Wong , Ziwei Liu

The capability of performing long-horizon, language-guided robotic manipulation tasks critically relies on leveraging historical information and generating coherent action sequences. However, such capabilities are often overlooked by…

Robotics · Computer Science 2025-12-24 Xiaofan Wang , Xingyu Gao , Jianlong Fu , Zuolei Li , Dean Fortier , Galen Mullins , Andrey Kolobov , Baining Guo

Robust robotic manipulation requires not only predicting how the scene evolves over time, but also recognizing task-relevant objects in complex scenes. However, existing VLA models face two limitations. They typically act only on the…

Robotics · Computer Science 2026-04-21 Kuanning Wang , Ke Fan , Chenhao Qiu , Zeyu Shangguan , Yuqian Fu , Yanwei Fu , Daniel Seita , Xiangyang Xue

We introduce PersonaHOI, a training- and tuning-free framework that fuses a general StableDiffusion model with a personalized face diffusion (PFD) model to generate identity-consistent human-object interaction (HOI) images. While existing…

Computer Vision and Pattern Recognition · Computer Science 2025-01-13 Xinting Hu , Haoran Wang , Jan Eric Lenssen , Bernt Schiele

Generalized robots must learn from diverse, large-scale human-object interactions (HOI) to operate robustly in the real world. Monocular internet videos offer a nearly limitless and readily available source of data, capturing an…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Boran Wen , Ye Lu , Sirui Wang , Keyan Wan , Jiahong Zhou , Junxuan Liang , Xinpeng Liu , Bang Xiao , Ruiyang Liu , Yong-Lu Li

Whole-body Humanoid-Object Interaction (HOI) is bottlenecked by the scarcity of high-fidelity 3D data. While video generative priors offer a promising alternative, existing methods suffer from \textit{Representation Misalignment} due to…

Prevalent human-object interaction (HOI) detection approaches typically leverage large-scale visual-linguistic models to help recognize events involving humans and objects. Though promising, models trained via contrastive learning on…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Liulei Li , Wenguan Wang , Yi Yang

Text-to-motion generation aims to generate 3D human motions that are tightly aligned with the input text while remaining physically plausible and rich in fine-grained detail. Although recent approaches can produce complex and natural…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Heng Li , Xiaotong Lin , Ling-An Zeng , Yulei Kang , Shuai Li , Jian-Fang Hu

Human-object interaction (HOI) detection plays a key role in high-level visual understanding, facilitating a deep comprehension of human activities. Specifically, HOI detection aims to locate the humans and objects involved in interactions…

Computer Vision and Pattern Recognition · Computer Science 2025-03-19 Yuxiao Wang , Yu Lei , Li Cui , Weiying Xue , Qi Liu , Zhenao Wei

One of the central challenges preventing robots from acquiring complex manipulation skills is the prohibitive cost of collecting large-scale robot demonstrations. In contrast, humans are able to learn efficiently by watching others interact…

Robotics · Computer Science 2025-11-13 Changhe Chen , Quantao Yang , Xiaohao Xu , Nima Fazeli , Olov Andersson

Interactive humanoid video generation aims to synthesize lifelike visual agents that can engage with humans through continuous and responsive video. Despite recent advances in video synthesis, existing methods often grapple with the…

Diffusion models revolutionize image generation by leveraging natural language to guide the creation of multimedia content. Despite significant advancements in such generative models, challenges persist in depicting detailed human-object…

Generative video modeling has emerged as a compelling tool to zero-shot reason about plausible physical interactions for open-world manipulation. Yet, it remains a challenge to translate such human-led motions into the low-level actions…

Robotics · Computer Science 2026-01-01 Karthik Dharmarajan , Wenlong Huang , Jiajun Wu , Li Fei-Fei , Ruohan Zhang

We present a comprehensive framework for egocentric interaction recognition using markerless 3D annotations of two hands manipulating objects. To this end, we propose a method to create a unified dataset for egocentric 3D interaction…

Computer Vision and Pattern Recognition · Computer Science 2021-08-25 Taein Kwon , Bugra Tekin , Jan Stuhmer , Federica Bogo , Marc Pollefeys

Human-Object Interaction Detection (HOI-DET) aims to localize human-object pairs and identify their interactive relationships. To aggregate contextual cues, existing methods typically propagate information across all detected entities via…

Computer Vision and Pattern Recognition · Computer Science 2025-10-22 Jiajun Hong , Jianan Wei , Wenguan Wang

Text-conditioned human motion generation has experienced significant advancements with diffusion models trained on extensive motion capture data and corresponding textual annotations. However, extending such success to 3D dynamic…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Sirui Xu , Ziyin Wang , Yu-Xiong Wang , Liang-Yan Gui

In this work, we are dedicated to a new task, i.e., hand-object interaction image generation, which aims to conditionally generate the hand-object image under the given hand, object and their interaction status. This task is challenging and…

Computer Vision and Pattern Recognition · Computer Science 2022-11-29 Hezhen Hu , Weilun Wang , Wengang Zhou , Houqiang Li

Many Vision-Language-Action (VLA) models are built upon an internal world model trained via next-frame prediction ``$v_t \rightarrow v_{t+1}$''. However, this paradigm attempts to predict the future frame's appearance directly, without…

Understanding the human-object interactions (HOIs) from a video is essential to fully comprehend a visual scene. This line of research has been addressed by detecting HOIs from images and lately from videos. However, the video-based HOI…

Computer Vision and Pattern Recognition · Computer Science 2023-06-07 Zhifan Ni , Esteve Valls Mascaró , Hyemin Ahn , Dongheui Lee