English
Related papers

Related papers: VAG: Dual-Stream Video-Action Generation for Embod…

200 papers

We revisit human motion synthesis, a task useful in various real world applications, in this paper. Whereas a number of methods have been developed previously for this task, they are often limited in two aspects: focusing on the poses while…

Computer Vision and Pattern Recognition · Computer Science 2021-06-01 Jingbo Wang , Sijie Yan , Bo Dai , Dahua LIn

Building an efficient and physically consistent world model from limited observations is a long standing challenge in vision and robotics. Many existing world modeling pipelines are based on implicit generative models, which are hard to…

Computer Vision and Pattern Recognition · Computer Science 2025-06-06 Wenhao Hu , Xuexiang Wen , Xi Li , Gaoang Wang

Video-to-video synthesis (vid2vid) aims for converting high-level semantic inputs to photorealistic videos. While existing vid2vid methods can achieve short-term temporal consistency, they fail to ensure the long-term one. This is because…

Computer Vision and Pattern Recognition · Computer Science 2020-07-17 Arun Mallya , Ting-Chun Wang , Karan Sapra , Ming-Yu Liu

Generative video modeling has emerged as a compelling tool to zero-shot reason about plausible physical interactions for open-world manipulation. Yet, it remains a challenge to translate such human-led motions into the low-level actions…

Robotics · Computer Science 2026-01-01 Karthik Dharmarajan , Wenlong Huang , Jiajun Wu , Li Fei-Fei , Ruohan Zhang

Recent video generation models have achieved remarkable progress and are now deployed in film, social media production, and advertising. Beyond their creative potential, such models also hold promise as world simulators for robotics and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 David Romero , Ariana Bermudez , Viacheslav Iablochnikov , Hao Li , Fabio Pizzati , Ivan Laptev

Diffusion Policy has dominated action generation due to its strong capabilities for modeling multi-modal action distributions, but its multi-step denoising processes make it impractical for real-time visuomotor control. Existing…

Robotics · Computer Science 2026-05-18 Kangye Ji , Jianbo Zhou , Yuan Meng , Ye Li , Hanyun Cui , Zhi Wang

Video generative models have emerged as a promising robotics backbone, capable of generating videos that depict the completion of complex tasks across embodiments and environments. Recent work proposes robot foundation models that jointly…

Robotics · Computer Science 2026-05-28 Sizhe Lester Li , Evan Kim , Xingjian Bai , Tong Zhao , Tao Pang , Max Simchowitz , Vincent Sitzmann

Scaling general-purpose manipulation to new robot embodiments remains challenging: each platform typically needs large, homogeneous demonstrations, and end-to-end pixel-to-action pipelines may degenerate under background and viewpoint…

Machine Learning · Computer Science 2025-12-23 Yao Feng , Hengkai Tan , Xinyi Mao , Chendong Xiang , Guodong Liu , Shuhe Huang , Hang Su , Jun Zhu

Diffusion-based video generation has achieved significant progress, yet generating multiple actions that occur sequentially remains a formidable task. Directly generating a video with sequential actions can be extremely challenging due to…

Computer Vision and Pattern Recognition · Computer Science 2024-05-29 Bowen Zhang , Xiaofei Xie , Haotian Lu , Na Ma , Tianlin Li , Qing Guo

World Models (WMs) have emerged as a promising approach for post-training Vision-Language-Action (VLA) policies to improve robustness and generalization under environmental changes. However, most WM-based post-training methods rely on…

Computer Vision and Pattern Recognition · Computer Science 2026-03-12 An Dinh Vuong , Tuan Van Vo , Abdullah Sohail , Haoran Ding , Liang Ma , Xiaodan Liang , Anqing Duan , Ivan Laptev , Ian Reid

Being able to generate realistic trajectory options is at the core of increasing the degree of automation of road vehicles. While model-driven, rule-based, and classical learning-based methods are widely used to tackle these tasks at…

Computer Vision and Pattern Recognition · Computer Science 2025-06-23 Annajoyce Mariani , Kira Maag , Hanno Gottschalk

Recent methods have made notable progress in the visual quality of hand-object interaction video synthesis. However, most approaches rely on 2D control signals that lack spatial expressiveness and limit the utilization of synthetic 3D…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Mingjin Chen , Junhao Chen , Zhaoxin Fan , Yujian Lee , Zichen Dang , Lili Wang , Yawen Cui , Lap-Pui Chau , Yi Wang

The field of robotics has made significant strides toward developing generalist robot manipulation policies. However, evaluating these policies in real-world scenarios remains time-consuming and challenging, particularly as the number of…

Robotics · Computer Science 2025-05-27 Yaxuan Li , Yichen Zhu , Junjie Wen , Chaomin Shen , Yi Xu

4D driving simulation is essential for developing realistic autonomous driving simulators. Despite advancements in existing methods for generating driving scenes, significant challenges remain in view transformation and spatial-temporal…

Computer Vision and Pattern Recognition · Computer Science 2024-12-12 Lening Wang , Wenzhao Zheng , Dalong Du , Yunpeng Zhang , Yilong Ren , Han Jiang , Zhiyong Cui , Haiyang Yu , Jie Zhou , Jiwen Lu , Shanghang Zhang

Anticipating actions before they occur is a core challenge in action understanding research. While conventional methods rely on extracting and aggregating temporal information from videos, as humans we can often predict upcoming actions by…

Computer Vision and Pattern Recognition · Computer Science 2025-12-03 Manuel Benavent-Lledo , Konstantinos Bacharidis , Victoria Manousaki , Konstantinos Papoutsakis , Antonis Argyros , Jose Garcia-Rodriguez

Endoscopic videos from multicentres often have different imaging conditions, e.g., color and illumination, which make the models trained on one domain usually fail to generalize well to another. Domain adaptation is one of the potential…

Computer Vision and Pattern Recognition · Computer Science 2020-04-20 Jiawei Chen , Yuexiang Li , Kai Ma , Yefeng Zheng

Cross-embodiment video generation aims to transfer motions across different humanoid embodiments, such as human-to-robot and robot-to-robot, enabling scalable data generation for embodied intelligence. A major challenge in this setting is…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Yiren Song , Xiyao Deng , Pei Yang , Yihan Wang , Mike Zheng Shou

We present Galaxea Open-World Dataset, a large-scale, diverse collection of robot behaviors recorded in authentic human living and working environments. All demonstrations are gathered using a consistent robotic embodiment, paired with…

Robotics · Computer Science 2025-09-03 Tao Jiang , Tianyuan Yuan , Yicheng Liu , Chenhao Lu , Jianning Cui , Xiao Liu , Shuiqi Cheng , Jiyang Gao , Huazhe Xu , Hang Zhao

Video colorization aims to transform grayscale videos into vivid color representations while maintaining temporal consistency and structural integrity. Existing video colorization methods often suffer from color bleeding and lack…

Computer Vision and Pattern Recognition · Computer Science 2025-01-17 Zixun Fang , Zhiheng Liu , Kai Zhu , Yu Liu , Ka Leong Cheng , Wei Zhai , Yang Cao , Zheng-Jun Zha

Reliable anticipation of traffic accidents is essential for advancing autonomous driving systems. However, this objective is limited by two fundamental challenges: the scarcity of diverse, high-quality training data and the frequent absence…

Computer Vision and Pattern Recognition · Computer Science 2025-07-18 Yanchen Guan , Haicheng Liao , Chengyue Wang , Xingcheng Liu , Jiaxun Zhang , Zhenning Li