English
Related papers

Related papers: InteractMove: Text-Controlled Human-Object Interac…

200 papers

We present a method for teaching machines to understand and model the underlying spatial common sense of diverse human-object interactions in 3D in a self-supervised way. This is a challenging task, as there exist specific manifolds of the…

Computer Vision and Pattern Recognition · Computer Science 2023-09-06 Sookwan Han , Hanbyul Joo

Recent generative models can synthesize high-quality images, but they often fail to generate humans interacting with objects using their hands. This arises mostly from the model's misunderstanding of such interactions and the hardships of…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Patrick Kwon , Chen Chen , Hanbyul Joo

To be useful in everyday environments, robots must be able to observe and learn about objects. Recent datasets enable progress for classifying data into known object categories; however, it is unclear how to collect reliable object data…

Robotics · Computer Science 2019-01-18 Abhishek Venkataraman , Brent Griffin , Jason J. Corso

We present a fully automatic system that takes a 3D scene and generates plausible 3D human bodies that are posed naturally in that 3D scene. Given a 3D scene without people, humans can easily imagine how people could interact with the scene…

Computer Vision and Pattern Recognition · Computer Science 2020-04-21 Yan Zhang , Mohamed Hassan , Heiko Neumann , Michael J. Black , Siyu Tang

Generating realistic human videos remains a challenging task, with the most effective methods currently relying on a human motion sequence as a control signal. Existing approaches often use existing motion extracted from other videos, which…

Computer Vision and Pattern Recognition · Computer Science 2024-12-18 Hsin-Ping Huang , Yang Zhou , Jui-Hsien Wang , Difan Liu , Feng Liu , Ming-Hsuan Yang , Zhan Xu

Understanding how people interact with their surroundings and each other is essential for enabling robots to act in socially compliant and context-aware ways. While 3D Scene Graphs have emerged as a powerful semantic representation for…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Ermanno Bartoli , Dennis Rotondi , Buwei He , Patric Jensfelt , Kai O. Arras , Iolanda Leite

Grounding 3D object affordance seeks to locate objects' ''action possibilities'' regions in the 3D space, which serves as a link between perception and operation for embodied agents. Existing studies primarily focus on connecting visual…

Computer Vision and Pattern Recognition · Computer Science 2023-08-10 Yuhang Yang , Wei Zhai , Hongchen Luo , Yang Cao , Jiebo Luo , Zheng-Jun Zha

This paper investigates one of the most challenging tasks in dynamic manipulation -- catching large-momentum moving objects. Beyond the realm of quasi-static manipulation, dealing with highly dynamic objects can significantly improve the…

Robotics · Computer Science 2024-03-27 Lei Yan , Theodoros Stouraitis , João Moura , Wenfu Xu , Michael Gienger , Sethu Vijayakumar

Generating high-fidelity 3D indoor scenes remains a significant challenge due to data scarcity and the complexity of modeling intricate spatial relations. Current methods often struggle to scale beyond training distribution to dense scenes…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Xingjian Ran , Shujie Zhang , Weipeng Zhong , Li Luo , Bo Dai

The physical properties of an object, such as mass, significantly affect how we manipulate it with our hands. Surprisingly, this aspect has so far been neglected in prior work on 3D motion synthesis. To improve the naturalness of the…

Computer Vision and Pattern Recognition · Computer Science 2023-12-25 Soshi Shimada , Franziska Mueller , Jan Bednarik , Bardia Doosti , Bernd Bickel , Danhang Tang , Vladislav Golyanik , Jonathan Taylor , Christian Theobalt , Thabo Beeler

The Human-Object Interaction (HOI) task explores the dynamic interactions between humans and objects in physical environments, providing essential biomechanical and cognitive-behavioral foundations for fields such as robotics, virtual…

Computer Vision and Pattern Recognition · Computer Science 2025-08-11 Ruiyan Wang , Lin Zuo , Zonghao Lin , Qiang Wang , Zhengxue Cheng , Rong Xie , Jun Ling , Li Song

Learning 3D human-object interaction relation is pivotal to embodied AI and interaction modeling. Most existing methods approach the goal by learning to predict isolated interaction elements, e.g., human contact, object affordance, and…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Yuhang Yang , Wei Zhai , Hongchen Luo , Yang Cao , Zheng-Jun Zha

Articulated object generation has seen increasing advancements, yet existing models often lack the ability to be conditioned on text prompts. To address the significant gap between textual descriptions and 3D articulated object…

Computer Vision and Pattern Recognition · Computer Science 2025-12-04 Hao Sun , Lei Fan , Donglin Di , Shaohui Liu

Understanding 3D scenes requires flexible combinations of visual reasoning tasks, including depth estimation, novel view synthesis, and object manipulation, all of which are essential for perception and interaction. Existing approaches have…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Wanhee Lee , Klemen Kotar , Rahul Mysore Venkatesh , Jared Watrous , Honglin Chen , Khai Loong Aw , Daniel L. K. Yamins

We consider the problem of learning object arrangements in a 3D scene. The key idea here is to learn how objects relate to human poses based on their affordances, ease of use and reachability. In contrast to modeling object-object…

Machine Learning · Computer Science 2012-07-03 Yun Jiang , Marcus Lim , Ashutosh Saxena

Semantics has enabled 3D scene understanding and affordance-driven object interaction. However, robots operating in real-world environments face a critical limitation: they cannot anticipate how objects move. Long-horizon mobile…

As pretrained text-to-image diffusion models become increasingly powerful, recent efforts have been made to distill knowledge from these text-to-image pretrained models for optimizing a text-guided 3D model. Most of the existing methods…

Computer Vision and Pattern Recognition · Computer Science 2024-06-12 Gege Gao , Weiyang Liu , Anpei Chen , Andreas Geiger , Bernhard Schölkopf

In this paper, we introduce RoleMotion, a large-scale human motion dataset that encompasses a wealth of role-playing and functional motion data tailored to fit various specific scenes. Existing text datasets are mainly constructed…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Junran Peng , Yiheng Huang , Silei Shen , Zeji Wei , Jingwei Yang , Baojie Wang , Yonghao He , Chuanchen Luo , Man Zhang , Xucheng Yin , Wei Sui

Recent advances in 3D scene reconstruction and 4D human animation have broadened adoption, but integrating the two remains difficult. Key challenges include placing humans at plausible locations and scales without interpenetration, aligning…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Qingyang Liu , Bingjie Gao , Weiheng Huang , Jun Zhang , Zhongqian Sun , Yang Wei , Fengrui Liu , Zelin Peng , Qianli Ma , Shuai Yang , Zhaohe Liao , Haonan Zhao , Li Niu

Diffusion models when conditioned on text prompts, generate realistic-looking images with intricate details. But most of these pre-trained models fail to generate accurate images when it comes to human features like hands, teeth, etc. We…

Computer Vision and Pattern Recognition · Computer Science 2023-12-27 Gurusha Juneja , Sukrit Kumar