English
Related papers

Related papers: Vidar: Embodied Video Diffusion Model for Generali…

200 papers

Diffusion models (DMs) have recently achieved impressive photorealism in image and video generation. However, their application to image animation remains limited, even when trained on large-scale datasets. Two primary challenges contribute…

Computer Vision and Pattern Recognition · Computer Science 2026-01-01 Zhenhao Li , Shaohan Yi , Zheng Liu , Leonartinus Gao , Minh Ngoc Le , Ambrose Ling , Zhuoran Wang , Md Amirul Islam , Zhixiang Chi , Yuanhao Yu

Reward and representation learning are two long-standing challenges for learning an expanding set of robot manipulation skills from sensory observations. Given the inherent cost and scarcity of in-domain, task-specific robot data, learning…

Robotics · Computer Science 2023-03-08 Yecheng Jason Ma , Shagun Sodhani , Dinesh Jayaraman , Osbert Bastani , Vikash Kumar , Amy Zhang

Learning directly from human demonstration videos is a key milestone toward scalable and generalizable robot learning. Yet existing methods rely on intermediate representations such as keypoints or trajectories, introducing information loss…

Computer Vision and Pattern Recognition · Computer Science 2025-12-22 Yiren Song , Cheng Liu , Weijia Mao , Mike Zheng Shou

Large language models (LLMs) have demonstrated that large-scale pretraining enables systems to adapt rapidly to new problems with little supervision in the language domain. This success, however, has not translated as effectively to the…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Pablo Acuaviva , Aram Davtyan , Mariam Hassan , Sebastian Stapf , Ahmad Rahimi , Alexandre Alahi , Paolo Favaro

The ability to specify robot commands by a non-expert user is critical for building generalist agents capable of solving a large variety of tasks. One convenient way to specify the intended robot goal is by a video of a person demonstrating…

Robotics · Computer Science 2023-05-11 Elliot Chane-Sane , Cordelia Schmid , Ivan Laptev

Recently, end-to-end robotic manipulation models have gained significant attention for their generalizability and scalability. However, they often suffer from limited robustness to camera viewpoint changes when training with a fixed camera.…

Robotics · Computer Science 2026-04-24 Songen Gu , Yuhang Zheng , Weize Li , Yupeng Zheng , Yating Feng , Xiang Li , Yilun Chen , Pengfei Li , Wenchao Ding

Learning generalizable visual representations across different embodied environments is essential for effective robotic manipulation in real-world scenarios. However, the limited scale and diversity of robot demonstration data pose a…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Jiaming Zhou , Teli Ma , Kun-Yu Lin , Zifan Wang , Ronghe Qiu , Junwei Liang

Generalization in embodied AI is hindered by the "seeing-to-doing gap," which stems from data scarcity and embodiment heterogeneity. To address this, we pioneer "pointing" as a unified, embodiment-agnostic intermediate representation,…

Robotics · Computer Science 2026-04-07 Yifu Yuan , Haiqin Cui , Yaoting Huang , Yibin Chen , Fei Ni , Zibin Dong , Pengyi Li , Yan Zheng , Hongyao Tang , Jianye Hao

Despite the central role of action in embodied intelligence, learning transferable action representations from visual transitions remains a fundamental challenge, particularly when world models must generalize across embodiments under…

Robotics · Computer Science 2026-05-19 Hongjia Liu , Fan Feng , Minghao Fu , Xinyue Wang , Haofei Lu , Biwei Huang

Pre-training for Reinforcement Learning (RL) with purely video data is a valuable yet challenging problem. Although in-the-wild videos are readily available and inhere a vast amount of prior world knowledge, the absence of action…

Computer Vision and Pattern Recognition · Computer Science 2024-11-06 Hao Luo , Bohan Zhou , Zongqing Lu

The generalization ability of imitation learning policies for robotic manipulation is fundamentally constrained by the diversity of expert demonstrations, while collecting demonstrations across varied environments is costly and difficult in…

Robotics · Computer Science 2026-04-02 Yichen Xie , Yixiao Wang , Shuqi Zhao , Cheng-En Wu , Masayoshi Tomizuka , Jianwen Xie , Hao-Shu Fang

Imitation Learning can train robots to perform complex and diverse manipulation tasks, but learned policies are brittle with observations outside of the training distribution. 3D scene representations that incorporate observations from…

Computer Vision and Pattern Recognition · Computer Science 2025-05-19 Albert Wilcox , Mohamed Ghanem , Masoud Moghani , Pierre Barroso , Benjamin Joffe , Animesh Garg

Learning robust and scalable visual representations from massive multi-view video data remains a challenge in computer vision and autonomous driving. Existing pre-training methods either rely on expensive supervised learning with 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-03-14 Jialv Zou , Bencheng Liao , Qian Zhang , Wenyu Liu , Xinggang Wang

Cross-embodiment learning seeks to build generalist robots that operate across diverse morphologies, but differences in action spaces and kinematics hinder data sharing and policy transfer. This raises a central question: Is there any…

Robotics · Computer Science 2025-11-11 Zihao He , Bo Ai , Tongzhou Mu , Yulin Liu , Weikang Wan , Jiawei Fu , Yilun Du , Henrik I. Christensen , Hao Su

World models, which predict future transitions from past observation and action sequences, have shown great promise for improving data efficiency in sequential decision-making. However, existing world models often require extensive…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Siqiao Huang , Jialong Wu , Qixing Zhou , Shangchen Miao , Mingsheng Long

Diffusion models have emerged as a powerful generative method for synthesizing high-quality and diverse set of images. In this paper, we propose a video generation method based on diffusion models, where the effects of motion are modeled in…

Computer Vision and Pattern Recognition · Computer Science 2022-12-02 Kangfu Mei , Vishal M. Patel

Predictive manipulation has recently gained considerable attention in the Embodied AI community due to its potential to improve robot policy performance by leveraging predicted states. However, generating accurate future visual states of…

Robotics · Computer Science 2025-09-15 Yuhang Huang , Jiazhao Zhang , Shilong Zou , Xinwang Liu , Ruizhen Hu , Kai Xu

Humans can recognize the same actions despite large context and viewpoint variations, such as differences between species (walking in spiders vs. horses), viewpoints (egocentric vs. third-person), and contexts (real life vs movies). Current…

Computer Vision and Pattern Recognition · Computer Science 2025-09-24 Rogerio Guimaraes , Frank Xiao , Pietro Perona , Markus Marks

We present a deep imitation learning framework for robotic bimanual manipulation in a continuous state-action space. A core challenge is to generalize the manipulation skills to objects in different locations. We hypothesize that modeling…

Recent progress has shown that video diffusion models (VDMs) can be repurposed for diverse multimodal graphics tasks. However, existing methods often train separate models for each problem setting, which fixes the input-output mapping and…

Computer Vision and Pattern Recognition · Computer Science 2026-05-04 Houyuan Chen , Hong Li , Xianghao Kong , Tianrui Zhu , Shaocong Xu , Weiqing Xiao , Yuwei Guo , Chongjie Ye , Lvmin Zhang , Hao Zhao , Anyi Rao