中文
相关论文

相关论文: Learning Video Generation for Robotic Manipulation…

200 篇论文

We propose ReCamDriving, a purely vision-based, camera-controlled novel-trajectory video generation framework. While repair-based methods fail to restore complex artifacts and LiDAR-based approaches rely on sparse and incomplete cues,…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Yaokun Li , Shuaixian Wang , Mantang Guo , Jiehui Huang , Taojun Ding , Mu Hu , Kaixuan Wang , Shaojie Shen , Guang Tan

We present Perceive-Represent-Generate (PRG), a novel three-stage framework that maps perceptual information of different modalities (e.g., visual or sound), corresponding to a sequence of instructions, to an adequate sequence of movements…

机器人学 · 计算机科学 2022-10-25 Fábio Vital , Miguel Vasco , Alberto Sardinha , Francisco Melo

Data scarcity continues to be a major challenge in the field of robotic manipulation. Although diffusion models provide a promising solution for generating robotic manipulation videos, existing methods largely depend on 2D trajectories,…

机器人学 · 计算机科学 2025-11-14 Ying Li , Xiaobao Wei , Xiaowei Chi , Yuming Li , Zhongyu Zhao , Hao Wang , Ningning Ma , Ming Lu , Sirui Han , Shanghang Zhang

In recent years, diffusion models have achieved tremendous success in the field of video generation, with controllable video generation receiving significant attention. However, existing control methods still face two limitations: Firstly,…

计算机视觉与模式识别 · 计算机科学 2024-10-15 Zhang Wan , Sheng Tang , Jiawei Wei , Ruize Zhang , Juan Cao

Recent advances in diffusion models bring new vitality to visual content creation. However, current text-to-video generation models still face significant challenges such as high training costs, substantial data requirements, and…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Sicong Feng , Jielong Yang , Li Peng

Manipulating articulated objects requires multiple robot arms in general. It is challenging to enable multiple robot arms to collaboratively complete manipulation tasks on articulated objects. In this paper, we present $\textbf{V-MAO}$, a…

机器人学 · 计算机科学 2021-11-09 Xingyu Liu , Kris M. Kitani

Research on video generation has recently made tremendous progress, enabling high-quality videos to be generated from text prompts or images. Adding control to the video generation process is an important goal moving forward and recent…

计算机视觉与模式识别 · 计算机科学 2024-05-28 Zhengfei Kuang , Shengqu Cai , Hao He , Yinghao Xu , Hongsheng Li , Leonidas Guibas , Gordon Wetzstein

The emergence of diffusion models has greatly propelled the progress in image and video generation. Recently, some efforts have been made in controllable video generation, including text-to-video generation and video motion control, among…

计算机视觉与模式识别 · 计算机科学 2024-05-02 Teng Hu , Jiangning Zhang , Ran Yi , Yating Wang , Hongrui Huang , Jieyu Weng , Yabiao Wang , Lizhuang Ma

Camera control, which achieves diverse visual effects by changing camera position and pose, has attracted widespread attention. However, existing methods face challenges such as complex interaction and limited control capabilities. To…

计算机视觉与模式识别 · 计算机科学 2025-04-04 Xiaoda Yang , Jiayang Xu , Kaixuan Luan , Xinyu Zhan , Hongshun Qiu , Shijun Shi , Hao Li , Shuai Yang , Li Zhang , Checheng Yu , Cewu Lu , Lixin Yang

We introduce a novel formulation for incorporating visual feedback in controlling robots. We define a generative model from actions to image observations of features on the end-effector. Inference in the model allows us to infer the robot…

机器人学 · 计算机科学 2020-03-11 Nishad Gothoskar , Miguel Lázaro-Gredilla , Abhishek Agarwal , Yasemin Bekiroglu , Dileep George

Despite tremendous recent progress, generative video models still struggle to capture real-world motion, dynamics, and physics. We show that this limitation arises from the conventional pixel reconstruction objective, which biases models…

计算机视觉与模式识别 · 计算机科学 2025-05-27 Hila Chefer , Uriel Singer , Amit Zohar , Yuval Kirstain , Adam Polyak , Yaniv Taigman , Lior Wolf , Shelly Sheynin

Video generation has advanced rapidly, producing photorealistic videos from text or image prompts. Meanwhile, film production and social robotics increasingly demand multi-person videos with rich social interactions, including…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Liangyang Ouyang , Ruicong Liu , Caixin Kang , Yifei Huang , Yoichi Sato

Simple as it seems, moving an object to another location within an image is, in fact, a challenging image-editing task that requires re-harmonizing the lighting, adjusting the pose based on perspective, accurately filling occluded regions,…

图形学 · 计算机科学 2025-03-12 Xin Yu , Tianyu Wang , Soo Ye Kim , Paul Guerrero , Xi Chen , Qing Liu , Zhe Lin , Xiaojuan Qi

This article suggests a reasoning-guided vision-language-motion diffusion framework (RG-VLMD) for generating instruction-aware co-speech gestures for humanoid robots in educational scenarios. The system integrates multi-modal affective…

机器人学 · 计算机科学 2026-03-20 Fuze Sun , Lingyu Li , Lekan Dai , Xinyu Fan

Despite remarkable achievements in video synthesis, achieving granular control over complex dynamics, such as nuanced movement among multiple interacting objects, still presents a significant hurdle for dynamic world modeling, compounded by…

计算机视觉与模式识别 · 计算机科学 2024-03-21 Pengxiang Li , Kai Chen , Zhili Liu , Ruiyuan Gao , Lanqing Hong , Guo Zhou , Hua Yao , Dit-Yan Yeung , Huchuan Lu , Xu Jia

Trajectory-based motion control has emerged as an intuitive and efficient approach for controllable video generation. However, the existing trajectory-based approaches are usually limited to only generating the motion trajectory of the…

计算机视觉与模式识别 · 计算机科学 2025-02-28 Yuhao Li , Mirana Claire Angel , Salman Khan , Yu Zhu , Jinqiu Sun , Yanning Zhang , Fahad Shahbaz Khan

We propose a unified framework for motion control in video generation that seamlessly integrates camera movement, object-level translation, and fine-grained local motion using trajectory-based inputs. In contrast to prior methods that…

计算机视觉与模式识别 · 计算机科学 2025-06-11 Angtian Wang , Haibin Huang , Jacob Zhiyuan Fang , Yiding Yang , Chongyang Ma

In recent years, there has been a significant effort dedicated to developing efficient, robust, and general human-to-robot handover systems. However, the area of flexible handover in the context of complex and continuous objects' motion…

机器人学 · 计算机科学 2023-08-31 Gu Zhang , Hao-Shu Fang , Hongjie Fang , Cewu Lu

Existing robot policies predominantly adopt the task-centric approach, requiring end-to-end task data collection. This results in limited generalization to new tasks and difficulties in pinpointing errors within long-horizon, multi-stage…

Adaptive control for real-time manipulation requires quick estimation and prediction of object properties. While robot learning in this area primarily focuses on using vision, many tasks cannot rely on vision due to object occlusion. Here,…

机器人学 · 计算机科学 2021-10-12 Ahalya Prabhakar , Stanislas Furrer , Lorenzo Panchetti , Maxence Perret , Aude Billard