中文
相关论文

相关论文: SFHand: Learning Embodied Manipulation by Streamin…

200 篇论文

We present a framework for pre-training of 3D hand pose estimation from in-the-wild hand images sharing with similar hand characteristics, dubbed SimHand. Pre-training with large-scale images achieves promising results in various tasks, but…

计算机视觉与模式识别 · 计算机科学 2025-05-06 Nie Lin , Takehiko Ohkawa , Yifei Huang , Mingfang Zhang , Minjie Cai , Ming Li , Ryosuke Furuta , Yoichi Sato

We hand the community HAND, a simple and time-efficient method for teaching robots new manipulation tasks through human hand demonstrations. Instead of relying on task-specific robot demonstrations collected via teleoperation, HAND uses…

机器人学 · 计算机科学 2025-10-28 Matthew Hong , Anthony Liang , Kevin Kim , Harshitha Rajaprakash , Jesse Thomason , Erdem Bıyık , Jesse Zhang

In surgical training for medical students, proficiency development relies on expert-led skill assessment, which is costly, time-limited, difficult to scale, and its expertise remains confined to institutions with available specialists.…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Le Ma , Thiago Freitas dos Santos , Nadia Magnenat-Thalmann , Katarzyna Wac

Embodied perception refers to the ability of an autonomous agent to perceive its environment so that it can (re)act. The responsiveness of the agent is largely governed by latency of its processing pipeline. While past work has studied the…

计算机视觉与模式识别 · 计算机科学 2020-08-26 Mengtian Li , Yu-Xiong Wang , Deva Ramanan

Sensing gloves have become important tools for teleoperation and robotic policy learning as they are able to provide rich signals like speed, acceleration and tactile feedback. A common approach to track gloved hands is to directly use the…

计算机视觉与模式识别 · 计算机科学 2026-02-06 Wenhui Cui , Ziyi Kou , Chuan Qin , Ergys Ristani , Li Guan

Estimating 3D interacting hand pose from a single RGB image is essential for understanding human actions. Unlike most previous works that directly predict the 3D poses of two interacting hands simultaneously, we propose to decompose the…

计算机视觉与模式识别 · 计算机科学 2022-07-25 Hao Meng , Sheng Jin , Wentao Liu , Chen Qian , Mengxiang Lin , Wanli Ouyang , Ping Luo

A large number of works in egocentric vision have concentrated on action and object recognition. Detection and segmentation of hands in first-person videos, however, has less been explored. For many applications in this domain, it is…

计算机视觉与模式识别 · 计算机科学 2018-03-30 Aisha Urooj Khan , Ali Borji

Despite recent progress in 3D hand reconstruction from monocular videos, most existing methods rely on data captured in well-controlled environments and therefore degrade in real-world settings with severe perturbations, such as hand-object…

计算机视觉与模式识别 · 计算机科学 2026-02-25 Hanhui Li , Xuan Huang , Wanquan Liu , Yuhao Cheng , Long Chen , Yiqiang Yan , Xiaodan Liang , Chenqiang Gao

Hand gesture serves as a crucial role during the expression of sign language. Current deep learning based methods for sign language understanding (SLU) are prone to over-fitting due to insufficient sign data resource and suffer limited…

计算机视觉与模式识别 · 计算机科学 2023-05-09 Hezhen Hu , Weichao Zhao , Wengang Zhou , Houqiang Li

Devising intelligent agents able to live in an environment and learn by observing the surroundings is a longstanding goal of Artificial Intelligence. From a bare Machine Learning perspective, challenges arise when the agent is prevented…

计算机视觉与模式识别 · 计算机科学 2022-04-27 Matteo Tiezzi , Simone Marullo , Lapo Faggi , Enrico Meloni , Alessandro Betti , Stefano Melacci

Predicting turn-taking in multiparty conversations has many practical applications in human-computer/robot interaction. However, the complexity of human communication makes it a challenging task. Recent advances have shown that synchronous…

计算机视觉与模式识别 · 计算机科学 2023-12-22 Mehdi Fatan , Emanuele Mincato , Dimitra Pintzou , Mariella Dimiccoli

In egocentric action recognition a single population model is typically trained and subsequently embodied on a head-mounted device, such as an augmented reality headset. While this model remains static for new users and environments, we…

计算机视觉与模式识别 · 计算机科学 2023-07-13 Matthias De Lange , Hamid Eghbalzadeh , Reuben Tan , Michael Iuzzolino , Franziska Meier , Karl Ridgeway

Egocentric interactive world models are essential for augmented reality and embodied AI, where visual generation must respond to user input with low latency, geometric consistency, and long-term stability. We study egocentric interaction…

计算机视觉与模式识别 · 计算机科学 2026-02-16 Yuxi Wang , Wenqi Ouyang , Tianyi Wei , Yi Dong , Zhiqi Shen , Xingang Pan

Motion-controllable video generation is crucial for egocentric applications in virtual reality and embodied AI. However, existing methods often struggle to achieve 3D-consistent fine-grained hand articulation. By adopting on 2D trajectories…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Chenyangguang Zhang , Botao Ye , Boqi Chen , Alexandros Delitzas , Fangjinhua Wang , Marc Pollefeys , Xi Wang

Real robot data collection for imitation learning has led to significant advancements in robotic manipulation. However, the requirement for robot hardware in the process fundamentally constrains the scale of the data. In this paper, we…

Hands are often severely occluded by objects, which makes 3D hand mesh estimation challenging. Previous works often have disregarded information at occluded regions. However, we argue that occluded regions have strong correlations with…

计算机视觉与模式识别 · 计算机科学 2022-03-29 JoonKyu Park , Yeonguk Oh , Gyeongsik Moon , Hongsuk Choi , Kyoung Mu Lee

Manipulation has long been a challenging task for robots, while humans can effortlessly perform complex interactions with objects, such as hanging a cup on the mug rack. A key reason is the lack of a large and uniform dataset for teaching…

机器人学 · 计算机科学 2025-06-09 Hongyan Zhi , Peihao Chen , Siyuan Zhou , Yubo Dong , Quanxi Wu , Lei Han , Mingkui Tan

Recent advances in conversational AI have been substantial, but developing real-time systems for perceptual task guidance remains challenging. These systems must provide interactive, proactive assistance based on streaming visual inputs,…

Generating realistic 3D hand motion from natural language is vital for VR, robotics, and human-computer interaction. Existing methods either focus on full-body motion, overlooking detailed hand gestures, or require explicit 3D object…

计算机视觉与模式识别 · 计算机科学 2026-03-26 Ching-Lam Cheng , Bin Zhu , Shengfeng He

In this work, we introduce the Virtual In-Hand Eye Transformer (VIHE), a novel method designed to enhance 3D manipulation capabilities through action-aware view rendering. VIHE autoregressively refines actions in multiple stages by…

机器人学 · 计算机科学 2024-03-20 Weiyao Wang , Yutian Lei , Shiyu Jin , Gregory D. Hager , Liangjun Zhang