English
Related papers

Related papers: Text2HOI: Text-guided 3D Motion Generation for Han…

200 papers

In this paper, a deep learning-based model for 3D human motion generation from the text is proposed via gesture action classification and an autoregressive model. The model focuses on generating special gestures that express human thinking,…

Computer Vision and Pattern Recognition · Computer Science 2022-11-21 Gwantae Kim , Youngsuk Ryu , Junyeop Lee , David K. Han , Jeongmin Bae , Hanseok Ko

Human-Object Interaction (HOI) detection is a fundamental visual task aiming at localizing and recognizing interactions between humans and objects. Existing works focus on the visual and linguistic features of humans and objects. However,…

Computer Vision and Pattern Recognition · Computer Science 2021-08-20 Tao He , Lianli Gao , Jingkuan Song , Yuan-Fang Li

Text-to-image models give rise to workflows which often begin with an exploration step, where users sift through a large collection of generated images. The global nature of the text-to-image generation process prevents users from narrowing…

Computer Vision and Pattern Recognition · Computer Science 2023-08-15 Or Patashnik , Daniel Garibi , Idan Azuri , Hadar Averbuch-Elor , Daniel Cohen-Or

We present Text2Gestures, a transformer-based learning method to interactively generate emotive full-body gestures for virtual agents aligned with natural language text inputs. Our method generates emotionally expressive gestures by…

Human-Computer Interaction · Computer Science 2024-11-26 Uttaran Bhattacharya , Nicholas Rewkowski , Abhishek Banerjee , Pooja Guhan , Aniket Bera , Dinesh Manocha

Human hands possess the dexterity to interact with diverse objects such as grasping specific parts of the objects and/or approaching them from desired directions. More importantly, humans can grasp objects of any shape without…

Robotics · Computer Science 2024-07-15 Hui Zhang , Sammy Christen , Zicong Fan , Otmar Hilliges , Jie Song

Generating realistic human motions that naturally respond to both spoken language and physical objects is crucial for interactive digital experiences. Current methods, however, address speech-driven gestures or object interactions…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Sreehari Rajan , Kunal Bhosikar , Charu Sharma

Generating interaction-centric videos, such as those depicting humans or robots interacting with objects, is crucial for embodied intelligence, as they provide rich and diverse visual priors for robot learning, manipulation policy training,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-24 Gen Li , Bo Zhao , Jianfei Yang , Laura Sevilla-Lara

Generating realistic 3D Human-Human Interaction (HHI) requires coherent modeling of the physical plausibility of the agents and their interaction semantics. Existing methods compress all motion information into a single latent…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Zichen Geng , Zeeshan Hayder , Bo Miao , Jian Liu , Wei Liu , Ajmal Mian

Generating digital humans that move realistically has many applications and is widely studied, but existing methods focus on the major limbs of the body, ignoring the hands and head. Hands have been separately studied, but the focus has…

Computer Vision and Pattern Recognition · Computer Science 2023-03-20 Omid Taheri , Vasileios Choutas , Michael J. Black , Dimitrios Tzionas

In 3D hand-object interaction (HOI) tasks, estimating precise joint poses of hands and objects from monocular RGB input remains highly challenging due to the inherent geometric ambiguity of RGB images and the severe mutual occlusions that…

Computer Vision and Pattern Recognition · Computer Science 2025-07-21 Yuechen Xie , Haobo Jiang , Jian Yang , Yigong Zhang , Jin Xie

We propose a new paradigm to automatically generate training data with accurate labels at scale using the text-to-image synthesis frameworks (e.g., DALL-E, Stable Diffusion, etc.). The proposed approach1 decouples training data generation…

Computer Vision and Pattern Recognition · Computer Science 2023-09-13 Yunhao Ge , Jiashu Xu , Brian Nlong Zhao , Neel Joshi , Laurent Itti , Vibhav Vineet

Recent progress in 3D hand--object interaction (HOI) generation has primarily focused on single--hand grasp synthesis, while bimanual manipulation remains significantly more challenging. Long--horizon planning instability, fine--grained…

Robotics · Computer Science 2026-03-11 Zhi Wang , Liu Liu , Ruonan Liu , Dan Guo , Meng Wang

Humans build 3D understandings of the world through active object exploration, using jointly their senses of vision and touch. However, in 3D shape reconstruction, most recent progress has relied on static datasets of limited sensory data…

Computer Vision and Pattern Recognition · Computer Science 2021-10-27 Edward J. Smith , David Meger , Luis Pineda , Roberto Calandra , Jitendra Malik , Adriana Romero , Michal Drozdzal

Robot-to-human object handover is an important step in many human robot collaboration tasks. A successful handover requires the robot to maintain a stable grasp on the object while making sure the human receives the object in a natural and…

Robotics · Computer Science 2024-10-01 Zixi Wang , Zeyi Liu , Nicolas Ouporov , Shuran Song

We present Text2Tex, a novel method for generating high-quality textures for 3D meshes from the given text prompts. Our method incorporates inpainting into a pre-trained depth-aware image diffusion model to progressively synthesize high…

Computer Vision and Pattern Recognition · Computer Science 2023-03-22 Dave Zhenyu Chen , Yawar Siddiqui , Hsin-Ying Lee , Sergey Tulyakov , Matthias Nießner

The Human-Object Interaction (HOI) task explores the dynamic interactions between humans and objects in physical environments, providing essential biomechanical and cognitive-behavioral foundations for fields such as robotics, virtual…

Computer Vision and Pattern Recognition · Computer Science 2025-08-11 Ruiyan Wang , Lin Zuo , Zonghao Lin , Qiang Wang , Zhengxue Cheng , Rong Xie , Jun Ling , Li Song

Text-to-image generation has made remarkable progress with the emergence of diffusion models. However, it is still a difficult task to generate images for street views based on text, mainly because the road topology of street scenes is…

Computer Vision and Pattern Recognition · Computer Science 2024-02-08 Jinming Su , Songen Gu , Yiting Duan , Xingyue Chen , Junfeng Luo

In recent years, 3D models have been utilized in many applications, such as auto-driver, 3D reconstruction, VR, and AR. However, the scarcity of 3D model data does not meet its practical demands. Thus, generating high-quality 3D models…

Computer Vision and Pattern Recognition · Computer Science 2023-05-26 Weizhi Nie , Ruidong Chen , Weijie Wang , Bruno Lepri , Nicu Sebe

Detecting human-object interactions (HOIs) is an intricate challenge in the field of computer vision. Existing methods for HOI detection heavily rely on appearance-based features, but these may not fully capture all the essential…

Computer Vision and Pattern Recognition · Computer Science 2023-07-04 Lijing Zhu , Qizhen Lan , Alvaro Velasquez , Houbing Song , Acharya Kamal , Qing Tian , Shuteng Niu

Reconstructing 3D Human-Object Interaction from an RGB image is essential for perceptive systems. Yet, this remains challenging as it requires capturing the subtle physical coupling between the body and objects. While current methods rely…

Computer Vision and Pattern Recognition · Computer Science 2026-04-23 Dimitrije Antić , Alvaro Budria , George Paschalidis , Sai Kumar Dwivedi , Dimitrios Tzionas