English
Related papers

Related papers: Masquerade: Learning from In-the-wild Human Videos…

200 papers

Data scarcity fundamentally limits the generalization of bimanual dexterous manipulation, as real-world data collection for dexterous hands is expensive and labor-intensive. Human manipulation videos, as a direct carrier of manipulation…

Robotics · Computer Science 2026-02-11 Juncheng Mu , Sizhe Yang , Yiming Bao , Hojin Bae , Tianming Wei , Linning Xu , Boyi Li , Huazhe Xu , Jiangmiao Pang

A significant bottleneck in humanoid policy learning is the acquisition of large-scale, diverse datasets, as collecting reliable real-world data remains both difficult and cost-prohibitive. To address this limitation, we introduce…

Robotics · Computer Science 2025-10-06 Rui Zhong , Yizhe Sun , Junjie Wen , Jinming Li , Chuang Cheng , Wei Dai , Zhiwen Zeng , Huimin Lu , Yichen Zhu , Yi Xu

Despite progress in human motion capture, existing multi-view methods often face challenges in estimating the 3D pose and shape of multiple closely interacting people. This difficulty arises from reliance on accurate 2D joint estimations,…

Computer Vision and Pattern Recognition · Computer Science 2024-08-21 Feichi Lu , Zijian Dong , Jie Song , Otmar Hilliges

Robots operating in complex and uncertain environments face considerable challenges. Advanced robotic systems often rely on extensive datasets to learn manipulation tasks. In contrast, when humans are faced with unfamiliar tasks, such as…

Robotics · Computer Science 2025-11-10 Yichen Zhu , Feifei Feng

Recent advances in video diffusion models have enabled realistic and controllable human image animation with temporal coherence. Although generating reasonable results, existing methods often overlook the need for regional supervision in…

Computer Vision and Pattern Recognition · Computer Science 2024-10-01 Zhongcong Xu , Chaoyue Song , Guoxian Song , Jianfeng Zhang , Jun Hao Liew , Hongyi Xu , You Xie , Linjie Luo , Guosheng Lin , Jiashi Feng , Mike Zheng Shou

This paper presents a novel Expressive Facial Control (ExFace) method based on Diffusion Transformers, which achieves precise mapping from human facial blendshapes to bionic robot motor control. By incorporating an innovative model…

Robotics · Computer Science 2026-02-19 Dong Zhang , Jingwei Peng , Yuyang Jiao , Jiayuan Gu , Jingyi Yu , Jiahao Chen

Manipulation has long been a challenging task for robots, while humans can effortlessly perform complex interactions with objects, such as hanging a cup on the mug rack. A key reason is the lack of a large and uniform dataset for teaching…

Robotics · Computer Science 2025-06-09 Hongyan Zhi , Peihao Chen , Siyuan Zhou , Yubo Dong , Quanxi Wu , Lei Han , Mingkui Tan

Multi-finger robotic hand manipulation and grasping are challenging due to the high-dimensional action space and the difficulty of acquiring large-scale training data. Existing approaches largely rely on human teleoperation with wearable…

Large behavior models have shown strong dexterous manipulation capabilities by extending imitation learning to large-scale training on multi-task robot data, yet their generalization remains limited by the insufficient robot data coverage.…

We propose a novel framework for learning high-level cognitive capabilities in robot manipulation tasks, such as making a smiley face using building blocks. These tasks often involve complex multi-step reasoning, presenting significant…

Robotics · Computer Science 2023-05-31 Chuhao Jin , Wenhui Tan , Jiange Yang , Bei Liu , Ruihua Song , Limin Wang , Jianlong Fu

The ability to learn manipulation skills by watching videos of humans has the potential to unlock a new source of highly scalable data for robot learning. Here, we tackle prehensile manipulation, in which tasks involve grasping an object…

Robotics · Computer Science 2026-02-16 Albert J. Zhai , Kuo-Hao Zeng , Jiasen Lu , Ali Farhadi , Shenlong Wang , Wei-Chiu Ma

Can we learn robot manipulation for everyday tasks, only by watching videos of humans doing arbitrary tasks in different unstructured settings? Unlike widely adopted strategies of learning task-specific behaviors or direct imitation of a…

Robotics · Computer Science 2023-02-07 Homanga Bharadhwaj , Abhinav Gupta , Shubham Tulsiani , Vikash Kumar

Learning robot manipulation from abundant human videos offers a scalable alternative to costly robot-specific data collection. However, domain gaps across visual, morphological, and physical aspects hinder direct imitation. To effectively…

Robotics · Computer Science 2025-09-16 Yangcen Liu , Woo Chul Shin , Yunhai Han , Zhenyang Chen , Harish Ravichandar , Danfei Xu

Humans are adept at learning new tasks by watching a few instructional videos. On the other hand, robots that learn new actions either require a lot of effort through trial and error, or use expert demonstrations that are challenging to…

Robotics · Computer Science 2020-11-16 Vladimír Petrík , Makarand Tapaswi , Ivan Laptev , Josef Sivic

Video representation learning has seen tremendous progress in recent years. This has been driven by many factors, including the scale of training and the success of visual models trained contrastively with language. While these factors have…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Mantas Skackauskas , Xinyue Hao , Laura Sevilla-Lara

Masked autoencoding has shown excellent performance on self-supervised video representation learning. Temporal redundancy has led to a high masking ratio and customized masking strategy in VideoMAE. In this paper, we aim to further improve…

Computer Vision and Pattern Recognition · Computer Science 2023-08-22 Bingkun Huang , Zhiyu Zhao , Guozhen Zhang , Yu Qiao , Limin Wang

Visual pre-training with large-scale real-world data has made great progress in recent years, showing great potential in robot learning with pixel observations. However, the recipes of visual pre-training for robot manipulation tasks are…

Robotics · Computer Science 2023-08-08 Ya Jing , Xuelin Zhu , Xingbin Liu , Qie Sima , Taozheng Yang , Yunhai Feng , Tao Kong

The ability to learn from human demonstration endows robots with the ability to automate various tasks. However, directly learning from human demonstration is challenging since the structure of the human hand can be very different from the…

Robotics · Computer Science 2022-12-09 Xingyu Liu , Deepak Pathak , Kris M. Kitani

Humanoid robots could be versatile and intuitive human avatars that operate remotely in inaccessible places: the robot could reproduce in the remote location the movements of an operator equipped with a wearable motion capture device while…

Robotics · Computer Science 2022-03-29 Luigi Penco , Jean-Baptiste Mouret , Serena Ivaldi

Vision is well-known for its use in manipulation, especially using visual servoing. Due to the 3D nature of the world, using multiple camera views and merging them creates better representations for Q-learning and in turn, trains more…

Machine Learning · Computer Science 2025-09-01 Abdulaziz Almuzairee , Rohan Patil , Dwait Bhatt , Henrik I. Christensen