English
Related papers

Related papers: MEgoHand: Multimodal Egocentric Hand-Object Intera…

200 papers

This paper introduces a framework, called EMOTION, for generating expressive motion sequences in humanoid robots, enhancing their ability to engage in humanlike non-verbal communication. Non-verbal cues such as facial expressions, gestures,…

Robotics · Computer Science 2024-10-31 Peide Huang , Yuhan Hu , Nataliya Nechyporenko , Daehwa Kim , Walter Talbott , Jian Zhang

We propose CG-HOI, the first method to address the task of generating dynamic 3D human-object interactions (HOIs) from text. We model the motion of both human and object in an interdependent fashion, as semantically rich human motion rarely…

Computer Vision and Pattern Recognition · Computer Science 2024-05-20 Christian Diller , Angela Dai

Our interaction with the world is an inherently multimodal experience. However, the understanding of human-to-object interactions has historically been addressed focusing on a single modality. In particular, a limited number of works have…

Computer Vision and Pattern Recognition · Computer Science 2019-10-16 Alejandro Cartas , Jordi Luque , Petia Radeva , Carlos Segura , Mariella Dimiccoli

We propose a multimodal, physically grounded approach for metric-scale amodal object reconstruction and pose estimation under severe hand occlusion. Unlike prior occlusion-aware 3D generation methods that rely only on vision, we leverage…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Gabriele Mario Caddeo , Pasquale Marra , Lorenzo Natale

The rapid development of Multimodal Large Language Models (MLLMs) has led to growing interest in egocentric video understanding, specifically the ability for MLLMs to recognize fine-grained hand-object interactions, track object state…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Yang Dai , Dian Jiao , Tianwei Lin , Wenqiao Zhang

Generating realistic and physically plausible 3D Human-Object Interactions (HOI) remains a key challenge in motion generation. One primary reason is that describing these physical constraints with words alone is difficult. To address this…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Songjin Cai , Linjie Zhong , Ling Guo , Changxing Ding

Egocentric motion capture with a head-mounted body-facing stereo camera is crucial for VR and AR applications but presents significant challenges such as heavy occlusions and limited annotated real-world data. Existing methods rely on…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Andrea Boscolo Camiletto , Jian Wang , Eduardo Alvarado , Rishabh Dabral , Thabo Beeler , Marc Habermann , Christian Theobalt

Previous works concerning single-view hand-held object reconstruction typically rely on supervision from 3D ground-truth models, which are hard to collect in real world. In contrast, readily accessible hand-object videos offer a promising…

Computer Vision and Pattern Recognition · Computer Science 2024-03-14 Chenyangguang Zhang , Guanlong Jiao , Yan Di , Gu Wang , Ziqin Huang , Ruida Zhang , Fabian Manhardt , Bowen Fu , Federico Tombari , Xiangyang Ji

Whole-body multi-modal human motion generation poses two primary challenges: creating an effective motion generation mechanism and integrating various modalities, such as text, speech, and music, into a cohesive framework. Unlike previous…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Zhe Li , Weihao Yuan , Weichao Shen , Siyu Zhu , Zilong Dong , Chang Xu

We present UnrealEgo, i.e., a new large-scale naturalistic dataset for egocentric 3D human pose estimation. UnrealEgo is based on an advanced concept of eyeglasses equipped with two fisheye cameras that can be used in unconstrained…

Computer Vision and Pattern Recognition · Computer Science 2022-08-03 Hiroyasu Akada , Jian Wang , Soshi Shimada , Masaki Takahashi , Christian Theobalt , Vladislav Golyanik

The rapid advancement of the automotive industry towards automated and semi-automated vehicles has rendered traditional methods of vehicle interaction, such as touch-based and voice command systems, inadequate for a widening range of…

Human-Computer Interaction · Computer Science 2024-02-08 Amr Gomaa , Guillermo Reyes , Michael Feld , Antonio Krüger

In this paper, we address the problem of estimating the hand pose from the egocentric view when the hand is interacting with objects. Specifically, we propose a method to label a dataset Ego-Siam which contains the egocentric images…

Computer Vision and Pattern Recognition · Computer Science 2021-10-01 Yao Lu , Walterio W. Mayol-Cuevas

Recent advances in video diffusion models have unlocked new potential for realistic audio-driven talking video generation. However, achieving seamless audio-lip synchronization, maintaining long-term identity consistency, and producing…

Computer Vision and Pattern Recognition · Computer Science 2024-12-06 Longtao Zheng , Yifan Zhang , Hanzhong Guo , Jiachun Pan , Zhenxiong Tan , Jiahao Lu , Chuanxin Tang , Bo An , Shuicheng Yan

Long context egocentric video understanding has recently attracted significant research attention, with augmented reality (AR) highlighted as one of its most important application domains. Nevertheless, the task remains highly challenging…

Machine Learning · Computer Science 2026-04-10 Qiance Tang , Ziqi Wang , Jieyu Lin , Ziyun Li , Barbara De Salvo , Sai Qian Zhang

Reconstructing a 3D hand mesh from a single RGB image is challenging due to complex articulations, self-occlusions, and depth ambiguities. Traditional discriminative methods, which learn a deterministic mapping from a 2D image to a single…

Computer Vision and Pattern Recognition · Computer Science 2025-03-20 Muhammad Usama Saleem , Ekkasit Pinyoanuntapong , Mayur Jagdishbhai Patel , Hongfei Xue , Ahmed Helmy , Srijan Das , Pu Wang

Objects manipulated by the hand (i.e., manipulanda) are particularly challenging to reconstruct from Internet videos. Not only does the hand occlude much of the object, but also the object is often only visible in a small number of image…

Computer Vision and Pattern Recognition · Computer Science 2026-01-01 Jane Wu , Georgios Pavlakos , Georgia Gkioxari , Jitendra Malik

Egocentric videos capture scenes from a wearer's viewpoint, resulting in dynamic backgrounds, frequent motion, and occlusions, posing challenges to accurate keystep recognition. We propose a flexible graph-learning framework for…

Computer Vision and Pattern Recognition · Computer Science 2026-02-11 Julia Lee Romero , Kyle Min , Subarna Tripathi , Morteza Karimzadeh

Synthesizing realistic human-object interaction (HOI) is essential for 3D computer vision and robotics, underpinning animation and embodied control. Existing approaches often require manually specified intermediate waypoints and place all…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Hwanhee Jung , Seunggwan Lee , Jeongyoon Yoon , SeungHyeon Kim , Giljoo Nam , Qixing Huang , Sangpil Kim

Current video models fail as world model as they lack fine-graiend control. General-purpose household robots require real-time fine motor control to handle delicate tasks and urgent situations. In this work, we introduce fine-grained…

Computer Vision and Pattern Recognition · Computer Science 2025-10-03 Yichen Li , Antonio Torralba

We present AssemblyHands, a large-scale benchmark dataset with accurate 3D hand pose annotations, to facilitate the study of egocentric activities with challenging hand-object interactions. The dataset includes synchronized egocentric and…

Computer Vision and Pattern Recognition · Computer Science 2023-04-25 Takehiko Ohkawa , Kun He , Fadime Sener , Tomas Hodan , Luan Tran , Cem Keskin
‹ Prev 1 8 9 10 Next ›