English
Related papers

Related papers: Masquerade: Learning from In-the-wild Human Videos…

200 papers

A key challenge in manipulation is learning a policy that can robustly generalize to diverse visual environments. A promising mechanism for learning robust policies is to leverage video generative models, which are pretrained on large-scale…

Robots can use Visual Imitation Learning (VIL) to learn manipulation tasks from video demonstrations. However, translating visual observations into actionable robot policies is challenging due to the high-dimensional nature of video data.…

Robotics · Computer Science 2025-01-22 Ananth Jonnavittula , Sagar Parekh , Dylan P. Losey

Human behaviors in the real world naturally encode rich, long-term contextual information that can be leveraged to train embodied agents for perception, understanding, and acting. However, existing capture systems typically rely on costly…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Wenjia Wang , Liang Pan , Huaijin Pi , Yuke Lou , Xuqian Ren , Yifan Wu , Zhouyingcheng Liao , Lei Yang , Rishabh Dabral , Christian Theobalt , Taku Komura

We present DexMan, an automated framework that converts human visual demonstrations into bimanual dexterous manipulation skills for humanoid robots in simulation. Operating directly on third-person videos of humans manipulating rigid…

Robotics · Computer Science 2025-10-10 Jhen Hsieh , Kuan-Hsun Tu , Kuo-Han Hung , Tsung-Wei Ke

We propose a method to reconstruct global human trajectories from videos in the wild. Our optimization method decouples the camera and human motion, which allows us to place people in the same world coordinate frame. Most existing methods…

Computer Vision and Pattern Recognition · Computer Science 2023-03-22 Vickie Ye , Georgios Pavlakos , Jitendra Malik , Angjoo Kanazawa

Humans inherently possess generalizable visual representations that empower them to efficiently explore and interact with the environments in manipulation tasks. We advocate that such a representation automatically arises from…

Human-robot teaming (HRT) systems often rely on large-scale datasets of human and robot interactions, especially for close-proximity collaboration tasks such as human-robot handovers. Learning robot manipulation policies from raw,…

Robotics · Computer Science 2025-08-14 Yuekun Wu , Yik Lung Pang , Andrea Cavallaro , Changjae Oh

Visual feedback is critical for motor skill acquisition in sports and rehabilitation, and psychological studies show that observing near-perfect versions of one's own performance accelerates learning more effectively than watching expert…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Arjun Somayazulu , Kristen Grauman

Human videos are a scalable source of training data for robot learning. However, humans and robots significantly differ in embodiment, making many human actions infeasible for direct execution on a robot. Still, these demonstrations convey…

Many recent advances in robotic manipulation have come through imitation learning, yet these rely largely on mimicking a particularly hard-to-acquire form of demonstrations: those collected on the same robot in the same room with the same…

Robotics · Computer Science 2025-04-01 Junyao Shi , Zhuolun Zhao , Tianyou Wang , Ian Pedroza , Amy Luo , Jie Wang , Jason Ma , Dinesh Jayaraman

Learning generic skills for humanoid robots interacting with 3D scenes by mimicking human data is a key research challenge with significant implications for robotics and real-world applications. However, existing methodologies and…

Robotics · Computer Science 2024-12-24 Yun Liu , Bowen Yang , Licheng Zhong , He Wang , Li Yi

Existing research on avatar creation is typically limited to laboratory datasets, which require high costs against scalability and exhibit insufficient representation of the real world. On the other hand, the web abounds with off-the-shelf…

Computer Vision and Pattern Recognition · Computer Science 2025-03-13 Zihao Huang , Shoukang Hu , Guangcong Wang , Tianqi Liu , Yuhang Zang , Zhiguo Cao , Wei Li , Ziwei Liu

Rendering the visual appearance of moving humans from occluded monocular videos is a challenging task. Most existing research renders 3D humans under ideal conditions, requiring a clear and unobstructed scene. Those methods cannot be used…

Computer Vision and Pattern Recognition · Computer Science 2025-08-18 Tiange Xiang , Adam Sun , Scott Delp , Kazuki Kozuka , Li Fei-Fei , Ehsan Adeli

Humanoid robots are envisioned as embodied intelligent agents capable of performing a wide range of human-level loco-manipulation tasks, particularly in scenarios requiring strenuous and repetitive labor. However, learning these skills is…

Robotics · Computer Science 2024-12-20 Junjia Liu , Zhuo Li , Minghao Yu , Zhipeng Dong , Sylvain Calinon , Darwin Caldwell , Fei Chen

Recent work in visual representation learning for robotics demonstrates the viability of learning from large video datasets of humans performing everyday tasks. Leveraging methods such as masked autoencoding and contrastive learning, these…

Human demonstrations as prompts are a powerful way to program robots to do long-horizon manipulation tasks. However, translating these demonstrations into robot-executable actions presents significant challenges due to execution mismatches…

Robotics · Computer Science 2025-04-01 Kushal Kedia , Prithwish Dan , Angela Chao , Maximus Adrian Pace , Sanjiban Choudhury

We present a method for teaching dexterous manipulation tasks to robots from human hand motion demonstrations. Unlike existing approaches that solely rely on kinematics information without taking into account the plausibility of robot and…

Robotics · Computer Science 2025-01-09 Sungjae Park , Seungho Lee , Mingi Choi , Jiye Lee , Jeonghwan Kim , Jisoo Kim , Hanbyul Joo

We present a universal motion representation that encompasses a comprehensive range of motor skills for physics-based humanoid control. Due to the high dimensionality of humanoids and the inherent difficulties in reinforcement learning,…

Computer Vision and Pattern Recognition · Computer Science 2024-04-15 Zhengyi Luo , Jinkun Cao , Josh Merel , Alexander Winkler , Jing Huang , Kris Kitani , Weipeng Xu

Learning new robot tasks on new platforms and in new scenes from only a handful of demonstrations remains challenging. While videos of other embodiments - humans and different robots - are abundant, differences in embodiment, camera, and…

Cross-embodiment video generation aims to transfer motions across different humanoid embodiments, such as human-to-robot and robot-to-robot, enabling scalable data generation for embodied intelligence. A major challenge in this setting is…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Yiren Song , Xiyao Deng , Pei Yang , Yihan Wang , Mike Zheng Shou
‹ Prev 1 3 4 5 6 7 10 Next ›