English
Related papers

Related papers: Object-centric 3D Motion Field for Robot Learning …

200 papers

Learning from videos offers a promising path toward generalist robots by providing rich visual and temporal priors beyond what real robot datasets contain. While existing video generative models produce impressive visual predictions, they…

Artificial Intelligence · Computer Science 2025-12-24 Hung-Chieh Fang , Kuo-Han Hung , Chu-Rong Chen , Po-Jung Chou , Chun-Kai Yang , Po-Chen Ko , Yu-Chiang Wang , Yueh-Hua Wu , Min-Hung Chen , Shao-Hua Sun

We address the challenge of acquiring real-world manipulation skills with a scalable framework. We hold the belief that identifying an appropriate prediction target capable of leveraging large-scale datasets is crucial for achieving…

Robotics · Computer Science 2024-09-24 Chengbo Yuan , Chuan Wen , Tong Zhang , Yang Gao

A robot's ability to act is fundamentally constrained by what it can perceive. Many existing approaches to visual representation learning utilize general-purpose training criteria, e.g. image reconstruction, smoothness in latent space, or…

We present a novel approach to weakly supervised object detection. Instead of annotated images, our method only requires two short videos to learn to detect a new object: 1) a video of a moving object and 2) one or more "negative" videos of…

Computer Vision and Pattern Recognition · Computer Science 2019-10-01 Rico Jonschkowski , Austin Stone

In recent years there has been considerable interest in human action recognition. Several approaches have been developed in order to enhance the automatic video analysis. Although some developments have been achieved by the computer vision…

Computer Vision and Pattern Recognition · Computer Science 2012-07-09 Núbia Rosa da Silva , Odemir Martinez Bruno

Estimating human pose from video is a task that receives considerable attention due to its applicability in numerous 3D fields. The complexity of prior knowledge of human body movements poses a challenge to neural network models in the task…

Computer Vision and Pattern Recognition · Computer Science 2023-09-26 Wenshuo Chen , Xiang Zhou , Zhengdi Yu , Weixi Gu , Kai Zhang

We present DexMan, an automated framework that converts human visual demonstrations into bimanual dexterous manipulation skills for humanoid robots in simulation. Operating directly on third-person videos of humans manipulating rigid…

Robotics · Computer Science 2025-10-10 Jhen Hsieh , Kuan-Hsun Tu , Kuo-Han Hung , Tsung-Wei Ke

The accuracy of monocular 3D human pose estimation depends on the viewpoint from which the image is captured. While freely moving cameras, such as on drones, provide control over this viewpoint, automatically positioning them at the…

Computer Vision and Pattern Recognition · Computer Science 2020-06-19 Sena Kiciroglu , Helge Rhodin , Sudipta N. Sinha , Mathieu Salzmann , Pascal Fua

Derived from rapid advances in computer vision and machine learning, video analysis tasks have been moving from inferring the present state to predicting the future state. Vision-based action recognition and prediction from videos are such…

Computer Vision and Pattern Recognition · Computer Science 2022-02-15 Yu Kong , Yun Fu

Training general-purpose robots requires learning from large and diverse data sources. Current approaches rely heavily on teleoperated demonstrations which are difficult to scale. We present a scalable framework for training manipulation…

Robotics · Computer Science 2026-05-29 Marion Lepert , Jiaying Fang , Jeannette Bohg

The ability to specify robot commands by a non-expert user is critical for building generalist agents capable of solving a large variety of tasks. One convenient way to specify the intended robot goal is by a video of a person demonstrating…

Robotics · Computer Science 2023-05-11 Elliot Chane-Sane , Cordelia Schmid , Ivan Laptev

Learning to use tools or objects in common scenes, particularly handling them in various ways as instructed, is a key challenge for developing interactive robots. Training models to generate such manipulation trajectories requires a large…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Tomoya Yoshida , Shuhei Kurita , Taichi Nishimura , Shinsuke Mori

Understanding 3D motion from videos presents inherent challenges due to the diverse types of movement, ranging from rigid and deformable objects to articulated structures. To overcome this, we propose Liv3Stroke, a novel approach for…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Jaeah Lee , Changwoon Choi , Young Min Kim , Jaesik Park

3D perceptual representations are well suited for robot manipulation as they easily encode occlusions and simplify spatial reasoning. Many manipulation tasks require high spatial precision in end-effector pose prediction, which typically…

Robotics · Computer Science 2023-10-23 Theophile Gervet , Zhou Xian , Nikolaos Gkanatsios , Katerina Fragkiadaki

Embodied learning for object-centric robotic manipulation is a rapidly developing and challenging area in embodied AI. It is crucial for advancing next-generation intelligent robots and has garnered significant interest recently. Unlike…

Robotics · Computer Science 2025-01-15 Ying Zheng , Lei Yao , Yuejiao Su , Yi Zhang , Yi Wang , Sicheng Zhao , Yiyi Zhang , Lap-Pui Chau

The growing interest in embodied intelligence has brought ego-centric perspectives to contemporary research. One significant challenge within this realm is the accurate localization and tracking of objects in ego-centric videos, primarily…

Computer Vision and Pattern Recognition · Computer Science 2024-10-14 Shengyu Hao , Wenhao Chai , Zhonghan Zhao , Meiqi Sun , Wendi Hu , Jieyang Zhou , Yixian Zhao , Qi Li , Yizhou Wang , Xi Li , Gaoang Wang

Training a deep network policy for robot manipulation is notoriously costly and time consuming as it depends on collecting a significant amount of real world data. To work well in the real world, the policy needs to see many instances of…

Robotics · Computer Science 2019-06-24 Xinchen Yan , Mohi Khansari , Jasmine Hsu , Yuanzheng Gong , Yunfei Bai , Sören Pirk , Honglak Lee

Bridging the gap between motion models and reality is crucial by using limited data to deploy robots in the real world. Deep learning is expected to be generalized to diverse situations while reducing feature design costs through end-to-end…

Robotics · Computer Science 2024-03-15 Kanata Suzuki , Hiroshi Ito , Tatsuro Yamada , Kei Kase , Tetsuya Ogata

The existing Motion Imitation models typically require expert data obtained through MoCap devices, but the vast amount of training data needed is difficult to acquire, necessitating substantial investments of financial resources, manpower,…

Robotics · Computer Science 2024-05-03 Liu Qiyuan

Current state-of-the-art video models process a video clip as a long sequence of spatio-temporal tokens. However, they do not explicitly model objects, their interactions across the video, and instead process all the tokens in the video. In…

Computer Vision and Pattern Recognition · Computer Science 2023-06-21 Xingyi Zhou , Anurag Arnab , Chen Sun , Cordelia Schmid