中文
相关论文

相关论文: Hand2World: Autoregressive Egocentric Interaction …

200 篇论文

Understanding and answering questions based on a user's pointing gesture is essential for next-generation egocentric AI assistants. However, current Multimodal Large Language Models (MLLMs) struggle with such tasks due to the lack of…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Yura Choi , Roy Miles , Rolandos Alexandros Potamias , Ismail Elezi , Jiankang Deng , Stefanos Zafeiriou

Understanding how humans would behave during hand-object interaction is vital for applications in service robot manipulation and extended reality. To achieve this, some recent works have been proposed to simultaneously forecast hand…

计算机视觉与模式识别 · 计算机科学 2025-11-17 Junyi Ma , Jingyi Xu , Xieyuanli Chen , Hesheng Wang

In this work, we are dedicated to a new task, i.e., hand-object interaction image generation, which aims to conditionally generate the hand-object image under the given hand, object and their interaction status. This task is challenging and…

计算机视觉与模式识别 · 计算机科学 2022-11-29 Hezhen Hu , Weilun Wang , Wengang Zhou , Houqiang Li

Accurate 3D tracking of hands and their interactions with the world in unconstrained settings remains a significant challenge for egocentric computer vision. With few exceptions, existing datasets are predominantly captured in controlled…

计算机视觉与模式识别 · 计算机科学 2025-10-06 Patrick Rim , Kun He , Kevin Harris , Braden Copple , Shangchen Han , Sizhe An , Ivan Shugurov , Tomas Hodan , He Wen , Xu Xie

Accurately identifying hands in images is a key sub-task for human activity understanding with wearable first-person point-of-view cameras. Traditional hand segmentation approaches rely on a large corpus of manually labeled data to generate…

计算机视觉与模式识别 · 计算机科学 2018-06-18 Yubo Zhang , Vishnu Naresh Boddeti , Kris M. Kitani

Despite remarkable progress in image generation models, generating realistic hands remains a persistent challenge due to their complex articulation, varying viewpoints, and frequent occlusions. We present FoundHand, a large-scale…

计算机视觉与模式识别 · 计算机科学 2024-12-06 Kefan Chen , Chaerin Min , Linguang Zhang , Shreyas Hampali , Cem Keskin , Srinath Sridhar

Learning human-object manipulation presents significant challenges due to its fine-grained and contact-rich nature of the motions involved. Traditional physics-based animation requires extensive modeling and manual setup, and more…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Quankai Gao , Jiawei Yang , Qiangeng Xu , Le Chen , Yue Wang

Talking head generation is to generate video based on a given source identity and target motion. However, current methods face several challenges that limit the quality and controllability of the generated videos. First, the generated face…

计算机视觉与模式识别 · 计算机科学 2023-11-03 Yue Gao , Yuan Zhou , Jinglu Wang , Xiao Li , Xiang Ming , Yan Lu

Recent approaches have demonstrated the promise of using diffusion models to generate interactive and explorable worlds. However, most of these methods face critical challenges such as excessively large parameter sizes, reliance on lengthy…

计算机视觉与模式识别 · 计算机科学 2025-12-29 Xiaofeng Mao , Zhen Li , Chuanhao Li , Xiaojie Xu , Kaining Ying , Tong He , Jiangmiao Pang , Yu Qiao , Kaipeng Zhang

World models, which predict future transitions from past observation and action sequences, have shown great promise for improving data efficiency in sequential decision-making. However, existing world models often require extensive…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Siqiao Huang , Jialong Wu , Qixing Zhou , Shangchen Miao , Mingsheng Long

Capturing interaction of hands with objects is important to autonomously detect human actions from egocentric videos. In this work, we present a pyramid video transformer with a dynamic class token generator for egocentric action…

计算机视觉与模式识别 · 计算机科学 2023-03-17 Chenbin Pan , Zhiqi Zhang , Senem Velipasalar , Yi Xu

3D world generation is essential for applications such as immersive content creation or autonomous driving simulation. Recent advances in 3D world generation have shown promising results; however, these methods are constrained by grid…

计算机视觉与模式识别 · 计算机科学 2026-05-04 Jaeyoung Chung , Suyoung Lee , Jianfeng Xiang , Jiaolong Yang , Kyoung Mu Lee

Multimodal AI agents are increasingly automating complex real-world workflows that involve online web execution. However, current web-agent benchmarks suffer from a critical limitation: they focus entirely on web-based interaction and…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Shoubin Yu , Lei Shu , Antoine Yang , Yao Fu , Srinivas Sunkara , Maria Wang , Jindong Chen , Mohit Bansal , Boqing Gong

Can conversational videos captured from multiple egocentric viewpoints reveal the map of a scene in a cost-efficient way? We seek to answer this question by proposing a new problem: efficiently building the map of a previously unseen 3D…

计算机视觉与模式识别 · 计算机科学 2023-04-24 Sagnik Majumder , Hao Jiang , Pierre Moulon , Ethan Henderson , Paul Calamia , Kristen Grauman , Vamsi Krishna Ithapu

Foundational world models must be both interactive and preserve spatiotemporal coherence for effective future planning with action choices. However, present models for long video generation have limited inherent world modeling capabilities…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Taiye Chen , Xun Hu , Zihan Ding , Chi Jin

Long context egocentric video understanding has recently attracted significant research attention, with augmented reality (AR) highlighted as one of its most important application domains. Nevertheless, the task remains highly challenging…

机器学习 · 计算机科学 2026-04-10 Qiance Tang , Ziqi Wang , Jieyu Lin , Ziyun Li , Barbara De Salvo , Sai Qian Zhang

We present InterHandGen, a novel framework that learns the generative prior of two-hand interaction. Sampling from our model yields plausible and diverse two-hand shapes in close interaction with or without an object. Our prior can be…

计算机视觉与模式识别 · 计算机科学 2024-03-27 Jihyun Lee , Shunsuke Saito , Giljoo Nam , Minhyuk Sung , Tae-Kyun Kim

In this paper, we propose a method to jointly determine the status of hand-object interaction. This is crucial for egocentric human activity understanding and interaction. From a computer vision perspective, we believe that determining…

计算机视觉与模式识别 · 计算机科学 2022-11-17 Yao Lu , Yanan Liu

With the rapid development of artificial intelligence technologies and wearable devices, egocentric vision understanding has emerged as a new and challenging research direction, gradually attracting widespread attention from both academia…

计算机视觉与模式识别 · 计算机科学 2025-09-25 Xiang Li , Heqian Qiu , Lanxiao Wang , Hanwen Zhang , Chenghao Qi , Linfeng Han , Huiyu Xiong , Hongliang Li

Text-to-image generation models have achieved remarkable advancements in recent years, aiming to produce realistic images from textual descriptions. However, these models often struggle with generating anatomically accurate representations…

计算机视觉与模式识别 · 计算机科学 2024-12-24 Haozhuo Zhang , Bin Zhu , Yu Cao , Yanbin Hao