中文
相关论文

相关论文: SFHand: Learning Embodied Manipulation by Streamin…

200 篇论文

Egocentric human videos provide a scalable source of manipulation demonstrations; however, deploying them on robots requires active viewpoint control to maintain task-critical visibility, which human viewpoint imitation often fails to…

机器人学 · 计算机科学 2026-02-27 Daesol Cho , Youngseok Jang , Danfei Xu , Sehoon Ha

High-fidelity hand gesture generation represents a significant challenge in human-centric generation tasks. Existing methods typically employ a single-view mesh-rendered image prior to enhancing gesture generation quality. However, the…

图形学 · 计算机科学 2025-08-07 Qifan Fu , Xu Chen , Muhammad Asad , Shanxin Yuan , Changjae Oh , Gregory Slabaugh

We present a self-trainable method, Mask2Hand, which learns to solve the challenging task of predicting 3D hand pose and shape from a 2D binary mask of hand silhouette/shadow without additional manually-annotated data. Given the intrinsic…

计算机视觉与模式识别 · 计算机科学 2022-07-04 Li-Jen Chang , Yu-Cheng Liao , Chia-Hui Lin , Hwann-Tzong Chen

We introduce Forecast-aware Gaussian Splatting (Forecast-GS), a predictive 3D representation framework for language-conditioned robotic manipulation. While recent manipulation systems have made progress by grounding language instructions…

机器人学 · 计算机科学 2026-05-13 Kaixin Jia , Jiacheng Xu

Most model-based 3D hand pose and shape estimation methods directly regress the parametric model parameters from an image to obtain 3D joints under weak supervision. However, these methods involve solving a complex optimization problem with…

计算机视觉与模式识别 · 计算机科学 2025-04-04 Shiyong Liu , Zhihao Li , Xiao Tang , Jianzhuang Liu

Understanding continuous video streams plays a fundamental role in real-time applications including embodied AI and autonomous driving. Unlike offline video understanding, streaming video understanding requires the ability to process video…

计算机视觉与模式识别 · 计算机科学 2025-07-23 Yibin Yan , Jilan Xu , Shangzhe Di , Yikun Liu , Yudi Shi , Qirui Chen , Zeqian Li , Yifei Huang , Weidi Xie

3D hand tracking methods based on monocular RGB videos are easily affected by motion blur, while event camera, a sensor with high temporal resolution and dynamic range, is naturally suitable for this task with sparse output and low power…

计算机视觉与模式识别 · 计算机科学 2023-03-01 Chuanlin Lan , Ziyuan Yin , Arindam Basu , Rosa H. M. Chan

Embodied control requires agents to leverage multi-modal pre-training to quickly learn how to act in new environments, where video demonstrations contain visual and motion details needed for low-level perception and control, and language…

机器学习 · 计算机科学 2023-04-20 Yao Mu , Shunyu Yao , Mingyu Ding , Ping Luo , Chuang Gan

For efficient human-agent interaction, an agent should proactively recognize their target user and prepare for upcoming interactions. We formulate this challenging problem as the novel task of jointly forecasting a person's intent to…

计算机视觉与模式识别 · 计算机科学 2025-05-09 Tongfei Bian , Yiming Ma , Mathieu Chollet , Victor Sanchez , Tanaya Guha

Despite remarkable progress in image generation models, generating realistic hands remains a persistent challenge due to their complex articulation, varying viewpoints, and frequent occlusions. We present FoundHand, a large-scale…

计算机视觉与模式识别 · 计算机科学 2024-12-06 Kefan Chen , Chaerin Min , Linguang Zhang , Shreyas Hampali , Cem Keskin , Srinath Sridhar

We study the problem of making 3D scene reconstructions interactive by asking the following question: can we predict the sounds of human hands physically interacting with a scene? First, we record a video of a human manipulating objects…

计算机视觉与模式识别 · 计算机科学 2025-06-12 Yiming Dou , Wonseok Oh , Yuqing Luo , Antonio Loquercio , Andrew Owens

Understanding how humans interact with the world necessitates accurate 3D hand pose estimation, a task complicated by the hand's high degree of articulation, frequent occlusions, self-occlusions, and rapid motions. While most existing…

计算机视觉与模式识别 · 计算机科学 2023-12-29 Enes Duran , Muhammed Kocabas , Vasileios Choutas , Zicong Fan , Michael J. Black

The development of embodied AI systems is increasingly constrained by the availability and structure of physical interaction data. Despite recent advances in vision-language-action (VLA) models, current pipelines suffer from high data…

机器人学 · 计算机科学 2026-03-24 Xinhai Sun , Xiang Shi , Menglin Zou , Wenlong Huang

While existing equivariant methods enhance data efficiency, they suffer from high computational intensity, reliance on single-modality inputs, and instability when combined with fast-sampling methods. In this work, we propose E3Flow, a…

机器人学 · 计算机科学 2026-03-25 Qinglun Zhang , Shen Cheng , Tian Dan , Haoqiang Fan , Guanghui Liu , Shuaicheng Liu

We study object interaction anticipation in egocentric videos. This task requires an understanding of the spatio-temporal context formed by past actions on objects, coined action context. We propose TransFusion, a multimodal…

计算机视觉与模式识别 · 计算机科学 2024-03-12 Razvan-George Pasca , Alexey Gavryushin , Muhammad Hamza , Yen-Ling Kuo , Kaichun Mo , Luc Van Gool , Otmar Hilliges , Xi Wang

We propose Hand-Eye Autonomous Delivery (HEAD), a framework that learns navigation, locomotion, and reaching skills for humanoids, directly from human motion and vision perception data. We take a modular approach where the high-level…

机器人学 · 计算机科学 2025-08-11 Sirui Chen , Yufei Ye , Zi-Ang Cao , Jennifer Lew , Pei Xu , C. Karen Liu

We propose EgoGrasp, the first method to reconstruct world-space hand-object interactions (W-HOI) from dynamic egoview videos, supporting open-vocabulary objects. Accurate W-HOI reconstruction is critical for embodied intelligence yet…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Hongming Fu , Wenjia Wang , Xiaozhen Qiao , Rolandos Alexandros Potamias , Taku Komura , Shuo Yang , Zheng Liu , Bo Zhao

Talking-head video editing aims to efficiently insert, delete, and substitute the word of a pre-recorded video through a text transcript editor. The key challenge for this task is obtaining an editing model that generates new talking-head…

多媒体 · 计算机科学 2023-09-21 Songlin Yang , Wei Wang , Jun Ling , Bo Peng , Xu Tan , Jing Dong

We propose the use of a proportional-derivative (PD) control based policy learned via reinforcement learning (RL) to estimate and forecast 3D human pose from egocentric videos. The method learns directly from unsegmented egocentric videos…

计算机视觉与模式识别 · 计算机科学 2019-08-06 Ye Yuan , Kris Kitani

Learning a generalizable bimanual manipulation policy is extremely challenging for embodied agents due to the large action space and the need for coordinated arm movements. Existing approaches rely on Vision-Language-Action (VLA) models to…

计算机视觉与模式识别 · 计算机科学 2025-06-02 Chenyou Fan , Fangzheng Yan , Chenjia Bai , Jiepeng Wang , Chi Zhang , Zhen Wang , Xuelong Li