English
Related papers

Related papers: THOM: Generating Physically Plausible Hand-Object …

200 papers

Despite recent advancements in text-to-image generation, most existing methods struggle to create images with multiple objects and complex spatial relationships in the 3D world. To tackle this limitation, we introduce a generic AI system,…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Yanbo Ding , Shaobin Zhuang , Kunchang Li , Zhengrong Yue , Yu Qiao , Yali Wang

Humans can steadily and gently grasp unfamiliar objects based on tactile perception. Robots still face challenges in achieving similar performance due to the difficulty of learning accurate grasp-force predictions and force control…

Robotics · Computer Science 2025-02-05 Mingxuan Li , Lunwei Zhang , Tiemin Li , Yao Jiang

Human-robot interaction (HRI) encompasses a wide range of collaborative tasks, with handover being one of the most fundamental. As robots become more integrated into human environments, the potential for service robots to assist in handing…

Robotics · Computer Science 2024-09-27 Jiangshan Liu , Wenlong Dong , Jiankun Wang , Max Q. -H. Meng

Developing autonomous physical human-robot interaction (pHRI) systems is limited by the scarcity of large-scale training data to learn robust robot behaviors for real-world applications. In this paper, we introduce a zero-shot…

Robotics · Computer Science 2026-04-13 Junxiang Wang , Xinwen Xu , Tiancheng Wu , Julian Millan , Nir Pechuk , Zackory Erickson

Modeling 4D human-object interaction (HOI) is a compelling challenge in computer vision and an essential technology powering virtual and mixed-reality applications. While existing works have achieved promising results on specific HOI…

Computer Vision and Pattern Recognition · Computer Science 2026-05-01 Mengfei Zhang , Jinlu Zhang , Zhigang Tu

Previous efforts have managed to generate production-ready 3D assets from text or images. However, these methods primarily employ NeRF or 3D Gaussian representations, which are not adept at producing smooth, high-quality geometries required…

Graphics · Computer Science 2024-10-15 Rengan Xie , Wenting Zheng , Kai Huang , Yizheng Chen , Qi Wang , Qi Ye , Wei Chen , Yuchi Huo

Understanding how we grasp objects with our hands has important applications in areas like robotics and mixed reality. However, this challenging problem requires accurate modeling of the contact between hands and objects. To capture grasps,…

Computer Vision and Pattern Recognition · Computer Science 2024-03-29 Chandradeep Pokhariya , Ishaan N Shah , Angela Xing , Zekun Li , Kefan Chen , Avinash Sharma , Srinath Sridhar

Learning-based methods to understand and model hand-object interactions (HOI) require a large amount of high-quality HOI data. One way to create HOI data is to transfer hand poses from a source object to another based on the objects'…

Computer Vision and Pattern Recognition · Computer Science 2026-01-07 Qiaochu Wang , Chufeng Xiao , Manfred Lau , Hongbo Fu

Hand-Object Interaction (HOI) remains a core challenge in digital human video synthesis, where models must generate physically plausible contact and preserve object identity across frames. Although recent HOI reenactment approaches have…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Xuan Huang , Mochu Xiang , Zhelun Shen , Jinbo Wu , Chenming Wu , Chen Zhao , Kaisiyuan Wang , Hang Zhou , Shanshan Liu , Haocheng Feng , Wei He , Jingdong Wang

The recently emerging text-to-motion advances have spired numerous attempts for convenient and interactive human motion generation. Yet, existing methods are largely limited to generating body motions only without considering the rich…

Computer Vision and Pattern Recognition · Computer Science 2024-09-27 Wenqian Zhang , Molin Huang , Yuxuan Zhou , Juze Zhang , Jingyi Yu , Jingya Wang , Lan Xu

The Large Vision Language Model (VLM) has recently addressed remarkable progress in bridging two fundamental modalities. VLM, trained by a sufficiently large dataset, exhibits a comprehensive understanding of both visual and linguistic to…

Computer Vision and Pattern Recognition · Computer Science 2024-11-28 Donggoo Kang , Dasol Jeong , Hyunmin Lee , Sangwoo Park , Hasil Park , Sunkyu Kwon , Yeongjoon Kim , Joonki Paik

Within the context of human-robot interaction (HRI), Theory of Mind (ToM) is intended to serve as a user-friendly backend to the interface of robotic systems, enabling robots to infer and respond to human mental states. When integrated into…

Robotics · Computer Science 2026-02-20 Marie S. Bauer , Julia Gachot , Matthias Kerzel , Cornelius Weber , Stefan Wermter

We present a physics-based framework to simulate porous, deformable materials and interactive tools with haptic feedback that can reshape it. In order to allow the material to be moulded non-homogeneously, we propose an algorithm to change…

Graphics · Computer Science 2022-02-24 Avirup Mandal , Parag Chaudhuri , Subhasis Chaudhuri

Reconstructing 3D Human-Object Interaction from an RGB image is essential for perceptive systems. Yet, this remains challenging as it requires capturing the subtle physical coupling between the body and objects. While current methods rely…

Computer Vision and Pattern Recognition · Computer Science 2026-04-23 Dimitrije Antić , Alvaro Budria , George Paschalidis , Sai Kumar Dwivedi , Dimitrios Tzionas

Human-object interaction (HOI) detection has seen advancements with Vision Language Models (VLMs), but these methods often depend on extensive manual annotations. Vision Large Language Models (VLLMs) can inherently recognize and reason…

Computer Vision and Pattern Recognition · Computer Science 2024-10-22 Jianjun Gao , Chen Cai , Ruoyu Wang , Wenyang Liu , Kim-Hui Yap , Kratika Garg , Boon-Siew Han

Existing automatic approaches for 3D virtual character motion synthesis supporting scene interactions do not generalise well to new objects outside training distributions, even when trained on extensive motion capture datasets with diverse…

Computer Vision and Pattern Recognition · Computer Science 2024-02-16 Wanyue Zhang , Rishabh Dabral , Thomas Leimkühler , Vladislav Golyanik , Marc Habermann , Christian Theobalt

Spatial relationships between objects provide important information for text-based image retrieval. As users are more likely to describe a scene from a real world perspective, using 3D spatial relationships rather than 2D relationships that…

Computer Vision and Pattern Recognition · Computer Science 2017-04-13 Ang Li , Jin Sun , Joe Yue-Hei Ng , Ruichi Yu , Vlad I. Morariu , Larry S. Davis

We propose a novel system for robot-to-human object handover that emulates human coworker interactions. Unlike most existing studies that focus primarily on grasping strategies and motion planning, our system focus on 1. inferring human…

Robotics · Computer Science 2025-03-06 Hanxin Zhang , Abdulqader Dhafer , Zhou Daniel Hao , Hongbiao Dong

We introduce HoME: a Household Multimodal Environment for artificial agents to learn from vision, audio, semantics, physics, and interaction with objects and other agents, all within a realistic context. HoME integrates over 45,000 diverse…

Artificial Intelligence · Computer Science 2017-11-30 Simon Brodeur , Ethan Perez , Ankesh Anand , Florian Golemo , Luca Celotti , Florian Strub , Jean Rouat , Hugo Larochelle , Aaron Courville

Humans naturally interact with both others and the surrounding multiple objects, engaging in various social activities. However, recent advances in modeling human-object interactions mostly focus on perceiving isolated individuals and…

Computer Vision and Pattern Recognition · Computer Science 2024-04-03 Juze Zhang , Jingyan Zhang , Zining Song , Zhanhe Shi , Chengfeng Zhao , Ye Shi , Jingyi Yu , Lan Xu , Jingya Wang