English
Related papers

Related papers: VIHE: Virtual In-Hand Eye Transformer for 3D Robot…

200 papers

Deploying visual reinforcement learning (RL) policies in real-world manipulation is often hindered by camera viewpoint changes. A policy trained from a fixed front-facing camera may fail when the camera is shifted -- an unavoidable…

Robotics · Computer Science 2026-03-13 Zheng Li , Pei Qu , Yufei Jia , Shihui Zhou , Haizhou Ge , Jiahang Cao , Jinni Zhou , Guyue Zhou , Jun Ma

Recent advances in robot manipulation have leveraged pre-trained vision-language models (VLMs) and explored integrating 3D spatial signals into these models for effective action prediction, giving rise to the promising…

Robotics · Computer Science 2026-01-14 Zhenyang Liu , Yongchong Gu , Yikai Wang , Xiangyang Xue , Yanwei Fu

We present AVI-HT, an adaptive visual-IMU fusion approach for tracking 3D hand poses by jointly modeling the egocentric image with on-glove 6-DoF IMU signals. AVI-HT achieves significantly improved accuracy and availability, particularly in…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Ziyi Kou , Ankit Kumar , Mia Huang , Taylor Niehues , Vatsal Mehta , Ergys Ristani , Li Guan

Optimizing behaviors for dexterous manipulation has been a longstanding challenge in robotics, with a variety of methods from model-based control to model-free reinforcement learning having been previously explored in literature. Perhaps…

Robotics · Computer Science 2022-03-25 Sridhar Pandian Arunachalam , Sneha Silwal , Ben Evans , Lerrel Pinto

Learning to represent three dimensional (3D) human pose given a two dimensional (2D) image of a person, is a challenging problem. In order to make the problem less ambiguous it has become common practice to estimate 3D pose in the camera…

Computer Vision and Pattern Recognition · Computer Science 2024-07-10 Mara Levy , Abhinav Shrivastava

In this work, we aim to learn a unified vision-based policy for multi-fingered robot hands to manipulate a variety of objects in diverse poses. Though prior work has shown benefits of using human videos for policy learning, performance…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Zerui Chen , Shizhe Chen , Etienne Arlaud , Ivan Laptev , Cordelia Schmid

This article presents a novel telepresence system for advancing aerial manipulation in dynamic and unstructured environments. The proposed system not only features a haptic device, but also a virtual reality (VR) interface that provides…

Tactile and visual perception are both crucial for humans to perform fine-grained interactions with their environment. Developing similar multi-modal sensing capabilities for robots can significantly enhance and expand their manipulation…

Robotics · Computer Science 2025-01-08 Binghao Huang , Yixuan Wang , Xinyi Yang , Yiyue Luo , Yunzhu Li

Hand-eye calibration is an important task in vision-guided robotic systems and is crucial for determining the transformation matrix between the camera coordinate system and the robot end-effector. Existing methods, for multi-view robotic…

Robotics · Computer Science 2025-07-29 Ye Wang , Haodong Jing , Yang Liao , Yongqiang Ma , Nanning Zheng

Eye-tracking applications that utilize the human gaze in video understanding tasks have become increasingly important. To effectively automate the process of video analysis based on eye-tracking data, it is important to accurately replicate…

Computer Vision and Pattern Recognition · Computer Science 2024-04-15 Suleyman Ozdel , Yao Rong , Berat Mert Albaba , Yen-Ling Kuo , Xi Wang , Enkelejda Kasneci

Executing contact-rich manipulation tasks necessitates the fusion of tactile and visual feedback. However, the distinct nature of these modalities poses significant challenges. In this paper, we introduce a system that leverages visual and…

Robotics · Computer Science 2024-08-01 Ying Yuan , Haichuan Che , Yuzhe Qin , Binghao Huang , Zhao-Heng Yin , Kang-Won Lee , Yi Wu , Soo-Chul Lim , Xiaolong Wang

In robot learning, Vision Transformers (ViTs) are standard for visual perception, yet most methods discard valuable information by using only the final layer's features. We argue this provides an insufficient representation and propose the…

Computer Vision and Pattern Recognition · Computer Science 2026-02-02 Wenhao Li , Chengwei Ma , Weixin Mao

This paper addresses the scarcity of large-scale datasets for accurate object-in-hand pose estimation, which is crucial for robotic in-hand manipulation within the ``Perception-Planning-Control" paradigm. Specifically, we introduce VinT-6D,…

Robots can use Visual Imitation Learning (VIL) to learn manipulation tasks from video demonstrations. However, translating visual observations into actionable robot policies is challenging due to the high-dimensional nature of video data.…

Robotics · Computer Science 2025-01-22 Ananth Jonnavittula , Sagar Parekh , Dylan P. Losey

Visual observations from different viewpoints can significantly influence the performance of visuomotor policies in robotic manipulation. Among these, egocentric (in-hand) views often provide crucial information for precise control.…

Robotics · Computer Science 2025-09-22 Haoran Ding , Anqing Duan , Zezhou Sun , Dezhen Song , Yoshihiko Nakamura

Automatically assessing handwritten mathematical solutions is an important problem in educational technology with practical applications, but it remains a significant challenge due to the diverse formats, unstructured layouts, and symbolic…

Computation and Language · Computer Science 2025-10-28 Thu Phuong Nguyen , Duc M. Nguyen , Hyotaek Jeon , Hyunwook Lee , Hyunmin Song , Sungahn Ko , Taehwan Kim

Robotic manipulation in complex scenes demands precise perception of task-relevant details, yet fixed or suboptimal viewpoints often impair fine-grained perception and induce occlusions, constraining imitation-learned policies. We present…

Eye-in-hand cameras have shown promise in enabling greater sample efficiency and generalization in vision-based robotic manipulation. However, for robotic imitation, it is still expensive to have a human teleoperator collect large amounts…

Robotics · Computer Science 2023-07-13 Moo Jin Kim , Jiajun Wu , Chelsea Finn

Learning representations in the joint domain of vision and touch can improve manipulation dexterity, robustness, and sample-complexity by exploiting mutual information and complementary cues. Here, we present Visuo-Tactile Transformers…

Robotics · Computer Science 2022-10-04 Yizhou Chen , Andrea Sipos , Mark Van der Merwe , Nima Fazeli

The current interacting hand (IH) datasets are relatively simplistic in terms of background and texture, with hand joints being annotated by a machine annotator, which may result in inaccuracies, and the diversity of pose distribution is…

Computer Vision and Pattern Recognition · Computer Science 2023-09-29 Lijun Li , Linrui Tian , Xindi Zhang , Qi Wang , Bang Zhang , Mengyuan Liu , Chen Chen