English
Related papers

Related papers: Video2Act: A Dual-System Video Diffusion Policy wi…

200 papers

Vision-Language Action (VLA) models significantly advance robotic manipulation by leveraging the strong perception capabilities of pretrained vision-language models (VLMs). By integrating action modules into these pretrained models, VLA…

Computer Vision and Pattern Recognition · Computer Science 2025-10-20 Shaoqi Dong , Chaoyou Fu , Haihan Gao , Yi-Fan Zhang , Chi Yan , Chu Wu , Xiaoyu Liu , Yunhang Shen , Jing Huo , Deqiang Jiang , Haoyu Cao , Yang Gao , Xing Sun , Ran He , Caifeng Shan

This paper presents DNAct, a language-conditioned multi-task policy framework that integrates neural rendering pre-training and diffusion training to enforce multi-modality learning in action sequence spaces. To learn a generalizable…

Robotics · Computer Science 2024-03-11 Ge Yan , Yueh-Hua Wu , Xiaolong Wang

Video action models (VAMs) have emerged as a promising paradigm for robot learning, owing to their powerful visual foresight for complex manipulation tasks. However, current VAMs, typically relying on either slow multi-step video generation…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Haodong Yan , Zhide Zhong , Jiaguan Zhu , Junjie He , Weilin Yuan , Wenxuan Song , Xin Gong , Yingjie Cai , Guanyi Zhao , Xu Yan , Bingbing Liu , Ying-Cong Chen , Haoang Li

Recently, end-to-end robotic manipulation models have gained significant attention for their generalizability and scalability. However, they often suffer from limited robustness to camera viewpoint changes when training with a fixed camera.…

Robotics · Computer Science 2026-04-24 Songen Gu , Yuhang Zheng , Weize Li , Yupeng Zheng , Yating Feng , Xiang Li , Yilun Chen , Pengfei Li , Wenchao Ding

Text-to-video (T2V) synthesis has gained increasing attention in the community, in which the recently emerged diffusion models (DMs) have promisingly shown stronger performance than the past approaches. While existing state-of-the-art DMs…

Artificial Intelligence · Computer Science 2024-03-20 Hao Fei , Shengqiong Wu , Wei Ji , Hanwang Zhang , Tat-Seng Chua

Vision-Language Models (VLMs) typically rely on static initial frames for video reasoning, restricting their ability to incorporate essential dynamic information as the reasoning process evolves. Existing methods that augment…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Martin Q. Ma , Yuxiao Qu , Aditya Agrawal , Willis Guo , Paul Pu Liang , Ruslan Salakhutdinov , Louis-Philippe Morency

Multi-agent applications have recently gained significant popularity. In many computer vision tasks, a network of agents, such as a team of robots with cameras, could work collaboratively to perceive the environment for efficient and…

Computer Vision and Pattern Recognition · Computer Science 2023-05-30 Shuyue Lan , Zhilu Wang , Ermin Wei , Amit K. Roy-Chowdhury , Qi Zhu

Video motion transfer aims to synthesize videos by generating visual content according to a text prompt while transferring the motion pattern observed in a reference video. Recent methods predominantly use the Diffusion Transformer (DiT)…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Yue Ma , Zhikai Wang , Tianhao Ren , Mingzhe Zheng , Hongyu Liu , Jiayi Guo , Kunyu Feng , Yuxuan Xue , Zixiang Zhao , Konrad Schindler , Qifeng Chen , Linfeng Zhang

We present Action Agent, a two-stage framework that unifies agentic navigation video generation with flow-constrained diffusion control for multi-embodiment robot navigation. In Stage I, a large language model (LLM) acts as an orchestration…

Robotics · Computer Science 2026-05-05 Jeffrin Sam , Nguyen Khang , Yara Mahmoud , Miguel Altamirano Cabrera , Dzmitry Tsetserukou

Portrait animation aims to generate photo-realistic videos from a single source image by reenacting the expression and pose from a driving video. While early methods relied on 3D morphable models or feature warping techniques, they often…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Mallikarjun B. R. , Fei Yin , Vikram Voleti , Nikita Drobyshev , Maksim Lapin , Aaryaman Vasishta , Varun Jampani

Diffusion models have obtained substantial progress in image-to-video generation. However, in this paper, we find that these models tend to generate videos with less motion than expected. We attribute this to the issue called conditional…

Computer Vision and Pattern Recognition · Computer Science 2024-11-07 Min Zhao , Hongzhou Zhu , Chendong Xiang , Kaiwen Zheng , Chongxuan Li , Jun Zhu

Predicting future motion trajectories is a critical capability across domains such as robotics, autonomous systems, and human activity forecasting, enabling safer and more intelligent decision-making. This paper proposes a novel, efficient,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-05 Zesen Zhong , Duomin Zhang , Yijia Li

A key challenge with procedure planning in instructional videos lies in how to handle a large decision space consisting of a multitude of action types that belong to various tasks. To understand real-world video content, an AI agent must…

Computer Vision and Pattern Recognition · Computer Science 2023-09-15 Fen Fang , Yun Liu , Ali Koksal , Qianli Xu , Joo-Hwee Lim

We present Im2Flow2Act, a scalable learning framework that enables robots to acquire real-world manipulation skills without the need of real-world robot training data. The key idea behind Im2Flow2Act is to use object flow as the…

Robotics · Computer Science 2024-10-07 Mengda Xu , Zhenjia Xu , Yinghao Xu , Cheng Chi , Gordon Wetzstein , Manuela Veloso , Shuran Song

Diffusion Models (DMs) have demonstrated state-of-the-art performance in content generation without requiring adversarial training. These models are trained using a two-step process. First, a forward - diffusion - process gradually adds…

Computer Vision and Pattern Recognition · Computer Science 2024-03-13 Anwaar Ulhaq , Naveed Akhtar

A generalist robotic policy needs both semantic understanding for task planning and the ability to interact with the environment through predictive capabilities. To tackle this, we present MM-ACT, a unified Vision-Language-Action (VLA)…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Haotian Liang , Xinyi Chen , Bin Wang , Mingkang Chen , Yitian Liu , Yuhao Zhang , Zanxin Chen , Tianshuo Yang , Yilun Chen , Jiangmiao Pang , Dong Liu , Xiaokang Yang , Yao Mu , Wenqi Shao , Ping Luo

Humans can recognize the same actions despite large context and viewpoint variations, such as differences between species (walking in spiders vs. horses), viewpoints (egocentric vs. third-person), and contexts (real life vs movies). Current…

Computer Vision and Pattern Recognition · Computer Science 2025-09-24 Rogerio Guimaraes , Frank Xiao , Pietro Perona , Markus Marks

Despite the remarkable progress in deep generative models, synthesizing high-resolution and temporally coherent videos still remains a challenge due to their high-dimensionality and complex temporal dynamics along with large spatial…

Computer Vision and Pattern Recognition · Computer Science 2023-03-31 Sihyun Yu , Kihyuk Sohn , Subin Kim , Jinwoo Shin

This study introduces an efficient and effective method, MeDM, that utilizes pre-trained image Diffusion Models for video-to-video translation with consistent temporal flow. The proposed framework can render videos from scene position…

Computer Vision and Pattern Recognition · Computer Science 2023-12-21 Ernie Chu , Tzuhsuan Huang , Shuo-Yen Lin , Jun-Cheng Chen

Learning transferable knowledge from unlabeled video data and applying it in new environments is a fundamental capability of intelligent agents. This work presents VideoWorld 2, which extends VideoWorld and offers the first investigation…

Computer Vision and Pattern Recognition · Computer Science 2026-02-11 Zhongwei Ren , Yunchao Wei , Xiao Yu , Guixun Luo , Yao Zhao , Bingyi Kang , Jiashi Feng , Xiaojie Jin
‹ Prev 1 3 4 5 6 7 10 Next ›