English
Related papers

Related papers: Robot See Robot Do: Imitating Articulated Object M…

200 papers

Recent advancements in 3D robotic manipulation have improved grasping of everyday objects, but transparent and specular materials remain challenging due to depth sensing limitations. While several 3D reconstruction and depth completion…

Robotics · Computer Science 2025-06-23 Mingxu Zhang , Xiaoqi Li , Jiahui Xu , Kaichen Zhou , Hojin Bae , Yan Shen , Chuyan Xiong , Hao Dong

In this paper, we propose a novel framework that allows therapists to teach robot-assisted rehabilitation exercises remotely via RGB-D video. Our system encodes demonstrations as 6-DoF body-centric trajectories using Cartesian Dynamic…

Robotics · Computer Science 2026-03-17 Ali Alabbas , Camillo Murgia , Joanne Regan , Philip Long

This work focuses on the 3D reconstruction of non-rigid objects based on monocular RGB video sequences. Concretely, we aim at building high-fidelity models for generic object categories and casually captured scenes. To this end, we do not…

Computer Vision and Pattern Recognition · Computer Science 2023-08-22 Yikai Wang , Yinpeng Dong , Fuchun Sun , Xiao Yang

We aim to teach robots to perform simple object manipulation tasks by watching a single video demonstration. Towards this goal, we propose an optimization approach that outputs a coarse and temporally evolving 3D scene to mimic the action…

Robotics · Computer Science 2022-08-04 Vladimir Petrik , Mohammad Nomaan Qureshi , Josef Sivic , Makarand Tapaswi

Imitation Learning from monocular video demonstrations provides a scalable approach for teaching complex skills to humanoid robots. However, translating human motion to humanoids requires overcoming significant morphological mismatches.…

Learning from human demonstrations has exhibited remarkable achievements in robot manipulation. However, the challenge remains to develop a robot system that matches human capabilities and data efficiency in learning and generalizability,…

Robotics · Computer Science 2024-01-05 Dingkun Guo

Robots operating in the real world must plan through environments that deform, yield, and reconfigure under contact, requiring interaction-aware 3D representations that extend beyond static geometric occupancy. To address this, we introduce…

Robotics · Computer Science 2026-02-16 Pavan Mantripragada , Siddhanth Deshmukh , Eadom Dessalene , Manas Desai , Yiannis Aloimonos

In this paper, we introduce a method to automatically reconstruct the 3D motion of a person interacting with an object from a single RGB video. Our method estimates the 3D poses of the person and the object, contact positions, and forces…

Computer Vision and Pattern Recognition · Computer Science 2019-06-18 Zongmian Li , Jiri Sedlar , Justin Carpentier , Ivan Laptev , Nicolas Mansard , Josef Sivic

The ability to successfully grasp objects is crucial in robotics, as it enables several interactive downstream applications. To this end, most approaches either compute the full 6D pose for the object of interest or learn to predict a set…

One of the most efficient ways for a learning-based robotic arm to learn to process complex tasks as human, is to directly learn from observing how human complete those tasks, and then imitate. Our idea is based on success of Deep…

Robotics · Computer Science 2018-10-05 Cheng Xuan , Zhiqiang Tang , Jinxin Xu

Our work aims to obtain 3D reconstruction of hands and manipulated objects from monocular videos. Reconstructing hand-object manipulations holds a great potential for robotics and learning from human demonstrations. The supervised learning…

Computer Vision and Pattern Recognition · Computer Science 2022-03-15 Yana Hasson , Gül Varol , Ivan Laptev , Cordelia Schmid

Reconstructing people, objects, and their interactions in 3D is a long-standing goal for intelligent systems. Often the input is RGB video from a moving camera, making the task ill-posed; depth is ambiguous, humans and objects occlude each…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Lixin Xue , Chengwei Zheng , Georgios Paschalidis , Chen Guo , Manuel Kaufmann , Juan Zarate , Dimitrios Tzionas

We study the problem of teaching humanoid robots manipulation skills by imitating from single video demonstrations. We introduce OKAMI, a method that generates a manipulation plan from a single RGB-D video and derives a policy for…

Robotics · Computer Science 2024-10-16 Jinhan Li , Yifeng Zhu , Yuqi Xie , Zhenyu Jiang , Mingyo Seo , Georgios Pavlakos , Yuke Zhu

Understanding articulated objects from monocular video is a crucial yet challenging task in robotics and digital twin creation. Existing methods often rely on complex multi-view setups, high-fidelity object scans, or fragile long-term point…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Arslan Artykov , Tom Ravaud , Corentin Sautier , Vincent Lepetit

The motivation of this paper is to develop a smart system using multi-modal vision for next-generation mechanical assembly. It includes two phases where in the first phase human beings teach the assembly structure to a robot and in the…

Robotics · Computer Science 2016-01-27 Weiwei Wan , Feng Lu , Zepei Wu , Kensuke Harada

We tackle the problem of monocular 3D reconstruction of articulated objects like humans and animals. We contribute DensePose 3D, a method that can learn such reconstructions in a weakly supervised fashion from 2D image annotations only.…

Computer Vision and Pattern Recognition · Computer Science 2021-09-02 Roman Shapovalov , David Novotny , Benjamin Graham , Patrick Labatut , Andrea Vedaldi

As interest grows in world models that predict future states from current observations and actions, accurately modeling part-level dynamics has become increasingly relevant for various applications. Existing approaches, such as…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Mingju Gao , Yike Pan , Huan-ang Gao , Zongzheng Zhang , Wenyi Li , Hao Dong , Hao Tang , Li Yi , Hao Zhao

Manipulating unseen articulated objects through visual feedback is a critical but challenging task for real robots. Existing learning-based solutions mainly focus on visual affordance learning or other pre-trained visual models to guide…

Robotics · Computer Science 2024-04-29 Pengwei Xie , Rui Chen , Siang Chen , Yuzhe Qin , Fanbo Xiang , Tianyu Sun , Jing Xu , Guijin Wang , Hao Su

We introduce a novel system for human-to-robot trajectory transfer that enables robots to manipulate objects by learning from human demonstration videos. The system consists of four modules. The first module is a data collection module that…

Robotics · Computer Science 2025-10-27 Sai Haneesh Allu , Jishnu Jaykumar P , Ninad Khargonkar , Tyler Summers , Jian Yao , Yu Xiang

Perceiving accurate 3D object shape is important for robots to interact with the physical world. Current research along this direction has been primarily relying on visual observations. Vision, however useful, has inherent limitations due…

Computer Vision and Pattern Recognition · Computer Science 2018-08-10 Shaoxiong Wang , Jiajun Wu , Xingyuan Sun , Wenzhen Yuan , William T. Freeman , Joshua B. Tenenbaum , Edward H. Adelson
‹ Prev 1 2 3 10 Next ›