English
Related papers

Related papers: FlowDreamer: A RGB-D World Model with Flow-based M…

200 papers

Performing facial expression transfer under one-shot setting has been increasing in popularity among research community with a focus on precise control of expressions. Existing techniques showcase compelling results in perceiving…

Computer Vision and Pattern Recognition · Computer Science 2024-04-24 Siddharth Nijhawan , Takuya Yashima , Tamaki Kojima

We present a novel framework for dynamic radiance field prediction given monocular video streams. Unlike previous methods that primarily focus on predicting future frames, our method goes a step further by generating explicit 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-01-29 Di Qi , Tong Yang , Beining Wang , Xiangyu Zhang , Wenqiang Zhang

Robot manipulation critically depends on perception that preserves the action-relevant aspects of a scene. Yet most robot learning pipelines are built upon visual encoders pre-trained for static recognition or vision-language alignment,…

Learning from Demonstration (LfD) offers a promising paradigm for robot skill acquisition. Recent approaches attempt to extract manipulation commands directly from video demonstrations, yet face two critical challenges: (1) general video…

Robotics · Computer Science 2026-02-24 Thanh Nguyen Canh , Thanh-Tuan Tran , Haolan Zhang , Ziyan Gao , Nak Young Chong , Xiem HoangVan

Recent advances in vision-language-action (VLA) models have shown promise in integrating image generation with action prediction to improve generalization and reasoning in robot manipulation. However, existing methods are limited to…

Computer Vision and Pattern Recognition · Computer Science 2025-08-27 Wenyao Zhang , Hongsi Liu , Zekun Qi , Yunnan Wang , Xinqiang Yu , Jiazhao Zhang , Runpei Dong , Jiawei He , Fan Lu , He Wang , Zhizheng Zhang , Li Yi , Wenjun Zeng , Xin Jin

Is a deep learning model capable of understanding systems governed by certain first principle laws by only observing the system's output? Can deep learning learn the underlying physics and honor the physics when making predictions? The…

Computational Physics · Physics 2020-06-11 Rohan Thavarajah , Xiang Zhai , Zheren Ma , David Castineira

Recent success in legged robot locomotion is attributed to the integration of reinforcement learning and physical simulators. However, these policies often encounter challenges when deployed in real-world environments due to sim-to-real…

Robotics · Computer Science 2025-06-04 Shaoting Zhu , Linzhan Mou , Derun Li , Baijun Ye , Runhan Huang , Hang Zhao

Predicting future frames of a video sequence has been a problem of high interest in the field of Computer Vision as it caters to a multitude of applications. The ability to predict, anticipate and reason about future events is the essence…

Computer Vision and Pattern Recognition · Computer Science 2020-09-04 Jasmeen Kaur , Sukhendu Das

Scene flow methods based on deep learning have achieved impressive performance. However, current top-performing methods still struggle with ill-posed regions, such as extensive flat regions or occlusions, due to insufficient local evidence.…

Computer Vision and Pattern Recognition · Computer Science 2024-12-24 Min Lin , Gangwei Xu , Yun Wang , Xianqi Wang , Xin Yang

Planning with world models offers a powerful paradigm for robotic control. Conventional approaches train a model to predict future frames conditioned on current frames and actions, which can then be used for planning. However, the objective…

Machine Learning · Computer Science 2025-10-23 Jacob Berg , Chuning Zhu , Yanda Bao , Ishan Durugkar , Abhishek Gupta

Pre-training for Reinforcement Learning (RL) with purely video data is a valuable yet challenging problem. Although in-the-wild videos are readily available and inhere a vast amount of prior world knowledge, the absence of action…

Computer Vision and Pattern Recognition · Computer Science 2024-11-06 Hao Luo , Bohan Zhou , Zongqing Lu

Generative manipulation policies can fail catastrophically under deployment-time distribution shift, yet many failures are near-misses: the robot reaches almost-correct poses and would succeed with a small corrective motion. We propose…

Robotics · Computer Science 2026-03-05 Edgar Welte , Yitian Shi , Rosa Wolf , Maximillian Gilles , Rania Rayyes

Most end-to-end Multi-Object Tracking (MOT) methods face the problems of low accuracy and poor generalization ability. Although traditional filter-based methods can achieve better results, they are difficult to be endowed with optimal…

Computer Vision and Pattern Recognition · Computer Science 2021-03-08 Guangyao Zhai , Xin Kong , Jinhao Cui , Yong Liu , Zhen Yang

While deep learning has shown tremendous success in a wide range of domains, it remains a grand challenge to incorporate physical principles in a systematic manner to the design, training, and inference of such models. In this paper, we aim…

Computational Physics · Physics 2020-06-16 Rui Wang , Karthik Kashinath , Mustafa Mustafa , Adrian Albert , Rose Yu

Video diffusion models provide powerful real-world simulators for embodied AI but remain limited in controllability for robotic manipulation. Recent works on trajectory-conditioned video generation address this gap but often rely on 2D…

Computer Vision and Pattern Recognition · Computer Science 2025-12-17 Yang Bai , Liudi Yang , George Eskandar , Fengyi Shen , Mohammad Altillawi , Ziyuan Liu , Gitta Kutyniok

We propose an action-conditioned dynamics model that predicts scene changes caused by object and agent interactions in a viewpoint-invariant 3D neural scene representation space, inferred from RGB-D videos. In this 3D feature space, objects…

Robotics · Computer Science 2020-12-29 Hsiao-Yu Fish Tung , Zhou Xian , Mihir Prabhudesai , Shamit Lal , Katerina Fragkiadaki

Data scarcity continues to be a major challenge in the field of robotic manipulation. Although diffusion models provide a promising solution for generating robotic manipulation videos, existing methods largely depend on 2D trajectories,…

Robotics · Computer Science 2025-11-14 Ying Li , Xiaobao Wei , Xiaowei Chi , Yuming Li , Zhongyu Zhao , Hao Wang , Ningning Ma , Ming Lu , Sirui Han , Shanghang Zhang

Text-to-video models have demonstrated substantial potential in robotic decision-making, enabling the imagination of realistic plans of future actions as well as accurate environment simulation. However, one major issue in such models is…

Robotics · Computer Science 2024-04-19 Siyuan Zhou , Yilun Du , Jiaben Chen , Yandong Li , Dit-Yan Yeung , Chuang Gan

We present a novel deep learning architecture for probabilistic future prediction from video. We predict the future semantics, geometry and motion of complex real-world urban scenes and use this representation to control an autonomous…

Computer Vision and Pattern Recognition · Computer Science 2020-07-20 Anthony Hu , Fergal Cotter , Nikhil Mohan , Corina Gurau , Alex Kendall

Existing research studies on vision and language grounding for robot navigation focus on improving model-free deep reinforcement learning (DRL) models in synthetic environments. However, model-free DRL models do not consider the dynamics in…

Computer Vision and Pattern Recognition · Computer Science 2018-07-27 Xin Wang , Wenhan Xiong , Hongmin Wang , William Yang Wang