English
Related papers

Related papers: Generative Proxemics: A Prior for 3D Social Intera…

200 papers

Humanoid robots can benefit from their similarity to the human shape by learning from humans. When humans teach other humans how to perform actions, they often demonstrate the actions, and the learning human imitates the demonstration to…

Robotics · Computer Science 2024-10-07 Josua Spisak , Matthias Kerzel , Stefan Wermter

Understanding 3d human interactions is fundamental for fine-grained scene analysis and behavioural modeling. However, most of the existing models predict incorrect, lifeless 3d estimates, that miss the subtle human contact aspects--the…

Computer Vision and Pattern Recognition · Computer Science 2023-08-07 Mihai Fieraru , Mihai Zanfir , Elisabeta Oneata , Alin-Ionut Popa , Vlad Olaru , Cristian Sminchisescu

While 3D hand reconstruction from monocular images has made significant progress, generating accurate and temporally coherent motion estimates from videos remains challenging, particularly during hand-object interactions. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Yufei Zhang , Zijun Cui , Jeffrey O. Kephart , Qiang Ji

The task of three-dimensional (3D) human pose estimation from a single image can be divided into two parts: (1) Two-dimensional (2D) human joint detection from the image and (2) estimating a 3D pose from the 2D joints. Herein, we focus on…

Computer Vision and Pattern Recognition · Computer Science 2018-03-23 Yasunori Kudo , Keisuke Ogaki , Yusuke Matsui , Yuri Odagiri

This paper presents PolyDiffuse, a novel structured reconstruction algorithm that transforms visual sensor data into polygonal shapes with Diffusion Models (DM), an emerging machinery amid exploding generative AI, while formulating…

Computer Vision and Pattern Recognition · Computer Science 2023-12-27 Jiacheng Chen , Ruizhi Deng , Yasutaka Furukawa

We present a generative approach to forecast long-term future human behavior in 3D, requiring only weak supervision from readily available 2D human action data. This is a fundamental task enabling many downstream applications. The required…

Computer Vision and Pattern Recognition · Computer Science 2024-05-20 Christian Diller , Thomas Funkhouser , Angela Dai

We present a method for learning 3D spatial relationships between object pairs, referred to as object-object spatial relationships (OOR), by leveraging synthetically generated 3D samples from pre-trained 2D diffusion models. We hypothesize…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Sangwon Baik , Hyeonwoo Kim , Hanbyul Joo

Generating high-quality whole-body human object interaction motion sequences is becoming increasingly important in various fields such as animation, VR/AR, and robotics. The main challenge of this task lies in determining the level of…

Computer Vision and Pattern Recognition · Computer Science 2024-12-31 Yonghao Zhang , Qiang He , Yanguang Wan , Yinda Zhang , Xiaoming Deng , Cuixia Ma , Hongan Wang

Image generative models, particularly diffusion-based models, have surged in popularity due to their remarkable ability to synthesize highly realistic images. However, since these models are data-driven, they inherit biases from the…

Machine Learning · Computer Science 2025-03-18 Lin-Chun Huang , Ching Chieh Tsao , Fang-Yi Su , Jung-Hsien Chiang

Recovering expressive humans from images is essential for understanding human behavior. Methods that estimate 3D bodies, faces, or hands have progressed significantly, yet separately. Face methods recover accurate 3D shape and geometric…

Computer Vision and Pattern Recognition · Computer Science 2021-10-19 Yao Feng , Vasileios Choutas , Timo Bolkart , Dimitrios Tzionas , Michael J. Black

The estimation of 3D human body pose and shape from a single image has been extensively studied in recent years. However, the texture generation problem has not been fully discussed. In this paper, we propose an end-to-end learning strategy…

Computer Vision and Pattern Recognition · Computer Science 2019-04-09 Jian Wang , Yunshan Zhong , Yachun Li , Chi Zhang , Yichen Wei

One-shot face re-enactment is a challenging task due to the identity mismatch between source and driving faces. Specifically, the suboptimally disentangled identity information of driving subjects would inevitably interfere with the…

Computer Vision and Pattern Recognition · Computer Science 2022-11-24 Yunfan Liu , Qi Li , Zhenan Sun , Tieniu Tan

There is a growing need for social robots and intelligent agents that can effectively interact with and support users. For the interactions to be seamless, the agents need to analyse social scenes and behavioural cues from their (robot's)…

Robotics · Computer Science 2025-10-28 Tongfei Bian , Mathieu Chollet , Tanaya Guha

Single-view 3D human reconstruction has garnered significant attention in recent years. Despite numerous advancements, prior research has concentrated on reconstructing 3D models from clear, close-up images of individual subjects, often…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Yizheng Song , Yiyu Zhuang , Qipeng Xu , Haixiang Wang , Jiahe Zhu , Jing Tian , Siyu Zhu , Hao Zhu

Recent years have seen significant progress in human image generation, particularly with the advancements in diffusion models. However, existing diffusion methods encounter challenges when producing consistent hand anatomy and the generated…

Computer Vision and Pattern Recognition · Computer Science 2024-05-01 Anton Pelykh , Ozge Mercanoglu Sincan , Richard Bowden

Given a user wearing a low frame rate wearable camera during a day, this work aims to automatically detect the moments when the user gets engaged into a social interaction solely by reviewing the automatically captured photos by the worn…

Computer Vision and Pattern Recognition · Computer Science 2017-05-15 Maedeh Aghaei , Mariella Dimiccoli , Petia Radeva

We address the problem of generating realistic 3D human-object interactions (HOIs) driven by textual prompts. To this end, we take a modular design and decompose the complex task into simpler sub-tasks. We first develop a dual-branch…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Xiaogang Peng , Yiming Xie , Zizhao Wu , Varun Jampani , Deqing Sun , Huaizu Jiang

Diffusion models have emerged as powerful generative models in the text-to-image domain. This paper studies their application as observation-to-action models for imitating human behaviour in sequential environments. Human behaviour is…

We propose a novel diffusion-based framework for reconstructing 3D geometry of hand-held objects from monocular RGB images by leveraging hand-object interaction as geometric guidance. Our method conditions a latent diffusion model on an…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Ayce Idil Aytekin , Helge Rhodin , Rishabh Dabral , Christian Theobalt

Human-human communication is like a delicate dance where listeners and speakers concurrently interact to maintain conversational dynamics. Hence, an effective model for generating listener nonverbal behaviors requires understanding the…

Computer Vision and Pattern Recognition · Computer Science 2024-07-19 Minh Tran , Di Chang , Maksim Siniukov , Mohammad Soleymani