English
Related papers

Related papers: Robotic Manipulation by Imitating Generated Videos…

200 papers

Video diffusion models (VDMs) have advanced significantly in recent years, enabling the generation of highly realistic videos and drawing the attention of the community in their potential as world simulators. However, despite their…

Computer Vision and Pattern Recognition · Computer Science 2025-04-07 Xindi Yang , Baolu Li , Yiming Zhang , Zhenfei Yin , Lei Bai , Liqian Ma , Zhiyong Wang , Jianfei Cai , Tien-Tsin Wong , Huchuan Lu , Xu Jia

Generating human videos with realistic and controllable motions is a challenging task. While existing methods can generate visually compelling videos, they lack separate control over four key video elements: foreground subject, background…

Computer Vision and Pattern Recognition · Computer Science 2025-08-13 Jingyun Liang , Jingkai Zhou , Shikai Li , Chenjie Cao , Lei Sun , Yichen Qian , Weihua Chen , Fan Wang

Video generation models offer a promising imagination mechanism for robot manipulation by predicting long-horizon future observations, but effectively exploiting these imagined futures for action execution remains challenging. Existing…

Robotics · Computer Science 2026-05-13 Yajie Li , Bozhou Zhang , Chun Gu , Zipei Ma , Jiahui Zhang , Jiankang Deng , Xiatian Zhu , Li Zhang

Robot manipulation research still suffers from significant data scarcity: even the largest robot datasets are orders of magnitude smaller and less diverse than those that fueled recent breakthroughs in language and vision. We introduce…

Robotics · Computer Science 2026-05-29 Marion Lepert , Jiaying Fang , Jeannette Bohg

The rapid advances in generative AI models have empowered the creation of highly realistic images with arbitrary content, raising concerns about potential misuse and harm, such as Deepfakes. Current research focuses on training detectors…

Computer Vision and Pattern Recognition · Computer Science 2024-05-31 Zhiyuan He , Pin-Yu Chen , Tsung-Yi Ho

In this paper, we introduce RAVID, the first framework for AI-generated image detection that leverages visual retrieval-augmented generation (RAG). While RAG methods have shown promise in mitigating factual inaccuracies in foundation…

Computer Vision and Pattern Recognition · Computer Science 2025-08-07 Mamadou Keita , Wassim Hamidouche , Hessen Bougueffa Eutamene , Abdelmalik Taleb-Ahmed , Abdenour Hadid

Humanoid loco-manipulation in unstructured environments demands tight integration of egocentric perception and whole-body control. However, existing approaches either depend on external motion capture systems or fail to generalize across…

Robotics · Computer Science 2025-11-14 Shaofeng Yin , Yanjie Ze , Hong-Xing Yu , C. Karen Liu , Jiajun Wu

Recent unsupervised pre-training methods have shown to be effective on language and vision domains by learning useful representations for multiple downstream tasks. In this paper, we investigate if such unsupervised pre-training methods can…

Computer Vision and Pattern Recognition · Computer Science 2022-06-20 Younggyo Seo , Kimin Lee , Stephen James , Pieter Abbeel

We introduce layered controllable video generation, where we, without any supervision, decompose the initial frame of a video into foreground and background layers, with which the user can control the video generation process by simply…

Computer Vision and Pattern Recognition · Computer Science 2022-10-05 Jiahui Huang , Yuhe Jin , Kwang Moo Yi , Leonid Sigal

Recently, natural language has been the primary medium for human-robot interaction. However, its inherent lack of spatial precision introduces challenges for robotic task definition such as ambiguity and verbosity. Moreover, in some public…

Robotics · Computer Science 2025-07-29 Yanbang Li , Ziyang Gong , Haoyang Li , Xiaoqi Huang , Haolan Kang , Guangping Bai , Xianzheng Ma

Video diffusion models provide powerful real-world simulators for embodied AI but remain limited in controllability for robotic manipulation. Recent works on trajectory-conditioned video generation address this gap but often rely on 2D…

Computer Vision and Pattern Recognition · Computer Science 2025-12-17 Yang Bai , Liudi Yang , George Eskandar , Fengyi Shen , Mohammad Altillawi , Ziyuan Liu , Gitta Kutyniok

We explore the potential of large-scale generative video models for autonomous driving, introducing an open-source auto-regressive video model (VaViM) and its companion video-action model (VaVAM) to investigate how video pre-training…

We present FloVD, a novel video diffusion model for camera-controllable video generation. FloVD leverages optical flow to represent the motions of the camera and moving objects. This approach offers two key benefits. Since optical flow can…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Wonjoon Jin , Qi Dai , Chong Luo , Seung-Hwan Baek , Sunghyun Cho

We propose a new concept, Evolution 6.0, which represents the evolution of robotics driven by Generative AI. When a robot lacks the necessary tools to accomplish a task requested by a human, it autonomously designs the required instruments…

Learning from human video demonstrations offers a scalable alternative to teleoperation or kinesthetic teaching, but poses challenges for robot manipulators due to embodiment differences and joint feasibility constraints. We address this…

Robotics · Computer Science 2025-09-26 Xiaoxiang Dong , Matthew Johnson-Roberson , Weiming Zhi

When performing 3D manipulation tasks, robots have to execute action planning based on perceptions from multiple fixed cameras. The multi-camera setup introduces substantial redundancy and irrelevant information, which increases…

Robotics · Computer Science 2025-12-19 Yixiang Chen , Yan Huang , Keji He , Peiyan Li , Liang Wang

Teaching robots dexterous manipulation skills often requires collecting hundreds of demonstrations using wearables or teleoperation, a process that is challenging to scale. Videos of human-object interactions are easier to collect and…

Robotics · Computer Science 2025-08-19 Tyler Ga Wei Lum , Olivia Y. Lee , C. Karen Liu , Jeannette Bohg

Vision-Language Models (VLMs) acquire real-world knowledge and general reasoning ability through Internet-scale image-text corpora. They can augment robotic systems with scene understanding and task planning, and assist visuomotor policies…

Robotics · Computer Science 2025-06-23 Kaiyuan Chen , Shuangyu Xie , Zehan Ma , Pannag R Sanketi , Ken Goldberg

We introduce LaViLa, a new approach to learning video-language representations by leveraging Large Language Models (LLMs). We repurpose pre-trained LLMs to be conditioned on visual input, and finetune them to create automatic video…

Computer Vision and Pattern Recognition · Computer Science 2022-12-09 Yue Zhao , Ishan Misra , Philipp Krähenbühl , Rohit Girdhar

Human motion synthesis is an important problem with applications in graphics, gaming and simulation environments for robotics. Existing methods require accurate motion capture data for training, which is costly to obtain. Instead, we…

Computer Vision and Pattern Recognition · Computer Science 2022-08-15 Kevin Xie , Tingwu Wang , Umar Iqbal , Yunrong Guo , Sanja Fidler , Florian Shkurti