English
Related papers

Related papers: ZeroHSI: Zero-Shot 4D Human-Scene Interaction by V…

200 papers

Humans are in constant contact with the world as they move through it and interact with it. This contact is a vital source of information for understanding 3D humans, 3D scenes, and the interactions between them. In fact, we demonstrate…

Computer Vision and Pattern Recognition · Computer Science 2022-03-30 Hongwei Yi , Chun-Hao P. Huang , Dimitrios Tzionas , Muhammed Kocabas , Mohamed Hassan , Siyu Tang , Justus Thies , Michael J. Black

We present a fully automatic system that takes a 3D scene and generates plausible 3D human bodies that are posed naturally in that 3D scene. Given a 3D scene without people, humans can easily imagine how people could interact with the scene…

Computer Vision and Pattern Recognition · Computer Science 2020-04-21 Yan Zhang , Mohamed Hassan , Heiko Neumann , Michael J. Black , Siyu Tang

Understanding how humans interact with each other is key to building realistic multi-human virtual reality systems. This area remains relatively unexplored due to the lack of large-scale datasets. Recent datasets focusing on this issue…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Rawal Khirodkar , Jyun-Ting Song , Jinkun Cao , Zhengyi Luo , Kris Kitani

Developing autonomous physical human-robot interaction (pHRI) systems is limited by the scarcity of large-scale training data to learn robust robot behaviors for real-world applications. In this paper, we introduce a zero-shot…

Robotics · Computer Science 2026-04-13 Junxiang Wang , Xinwen Xu , Tiancheng Wu , Julian Millan , Nir Pechuk , Zackory Erickson

Synthesizing natural interactions between virtual humans and their 3D environments is critical for numerous applications, such as computer games and AR/VR experiences. Our goal is to synthesize humans interacting with a given 3D scene…

Computer Vision and Pattern Recognition · Computer Science 2022-07-27 Kaifeng Zhao , Shaofei Wang , Yan Zhang , Thabo Beeler , Siyu Tang

Synthesizing human--object interaction (HOI) videos has broad practical value in e-commerce, digital advertising, and virtual marketing. However, current diffusion models, despite their photorealistic rendering capability, still frequently…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Xiangyang Luo , Xiaozhe Xin , Tao Feng , Xu Guo , Meiguang Jin , Junfeng Ma

Synthesizing human motions in 3D environments, particularly those with complex activities such as locomotion, hand-reaching, and human-object interaction, presents substantial demands for user-defined waypoints and stage transitions. These…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Nan Jiang , Zimo He , Zi Wang , Hongjie Li , Yixin Chen , Siyuan Huang , Yixin Zhu

In this study, we tackle the complex task of generating 3D human-object interactions (HOI) from textual descriptions in a zero-shot text-to-3D manner. We identify and address two key challenges: the unsatisfactory outcomes of direct…

Computer Vision and Pattern Recognition · Computer Science 2024-07-17 Sisi Dai , Wenhao Li , Haowen Sun , Haibin Huang , Chongyang Ma , Hui Huang , Kai Xu , Ruizhen Hu

The way humans interact with each other, including interpersonal distances, spatial configuration, and motion, varies significantly across different situations. To enable machines to understand such complex, context-dependent behaviors, it…

Computer Vision and Pattern Recognition · Computer Science 2025-12-25 Jeonghyeon Na , Sangwon Baik , Inhee Lee , Junyoung Lee , Hanbyul Joo

This paper tackles the problem of physics-aware human motion synthesis in a dynamic scene. Unlike existing works which mainly tend to generate physically unrealistic motions due to limited contact modeling, typically restricted to hands, in…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Chaoyue Xing , Wei Mao , Miaomiao Liu

Creating scenes for captured motions that achieve realistic human-scene interaction is crucial for 3D animation in movies or video games. As character motion is often captured in a blue-screened studio without real furniture or objects in…

Computer Vision and Pattern Recognition · Computer Science 2024-05-22 Jianan Li , Tao Huang , Qingxu Zhu , Tien-Tsin Wong

We present DreamHOI, a novel method for zero-shot synthesis of human-object interactions (HOIs), enabling a 3D human model to realistically interact with any given object based on a textual description. This task is complicated by the…

Computer Vision and Pattern Recognition · Computer Science 2024-09-13 Thomas Hanwen Zhu , Ruining Li , Tomas Jakab

Modeling human-scene interactions (HSI) is essential for understanding and simulating everyday human behaviors. Recent approaches utilizing generative modeling have made progress in this domain; however, they are limited in controllability…

Computer Vision and Pattern Recognition · Computer Science 2025-07-28 Inwoo Hwang , Bing Zhou , Young Min Kim , Jian Wang , Chuan Guo

Recent advances in video generative models enable the synthesis of realistic human-object interaction videos across a wide range of scenarios and object categories, including complex dexterous manipulations that are difficult to capture…

Computer Vision and Pattern Recognition · Computer Science 2026-04-23 Hyeonwoo Kim , Jeonghwan Kim , Kyungwon Cho , Hanbyul Joo

Text-conditioned motion synthesis has made remarkable progress with the emergence of diffusion models. However, the majority of these motion diffusion models are primarily designed for a single character and overlook multi-human…

Computer Vision and Pattern Recognition · Computer Science 2024-11-22 Zhenzhi Wang , Jingbo Wang , Yixuan Li , Dahua Lin , Bo Dai

Reconstructing metrically accurate humans and their surrounding scenes from a single image is crucial for virtual reality, robotics, and comprehensive 3D scene understanding. However, existing methods struggle with depth ambiguity,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Pradyumna Yalandur Muralidhar , Yuxuan Xue , Xianghui Xie , Margaret Kostyrko , Gerard Pons-Moll

Existing video generation models predominantly emphasize appearance fidelity while exhibiting limited ability to synthesize complex human motions, such as whole-body movements, long-range dynamics, and fine-grained human-environment…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Haoyu Wang , Hao Tang , Donglin Di , Zhilu Zhang , Wangmeng Zuo , Feng Gao , Siwei Ma , Shiliang Zhang

We present HSImul3R, a unified framework for simulation-ready 3D reconstruction of human-scene interactions (HSI) from casual captures, including sparse-view images and monocular videos. Existing methods suffer from a perception-simulation…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Yukang Cao , Haozhe Xie , Fangzhou Hong , Long Zhuo , Zhaoxi Chen , Liang Pan , Ziwei Liu

Human video synthesis aims to create lifelike characters in various environments, with wide applications in VR, storytelling, and content creation. While 2D diffusion-based methods have made significant progress, they struggle to generalize…

Computer Vision and Pattern Recognition · Computer Science 2024-12-19 Liyuan Cui , Xiaogang Xu , Wenqi Dong , Zesong Yang , Hujun Bao , Zhaopeng Cui

We present HOI-PAGE, a new approach that prioritizes part-level affordance reasoning to generate high-fidelity 4D human-object interactions (HOIs) from text prompts in a zero-shot fashion. In contrast to prior works that focus on global,…

Graphics · Computer Science 2026-05-20 Lei Li , Angela Dai