English
Related papers

Related papers: Future Optical Flow Prediction Improves Robot Cont…

200 papers

Recent advances in FlowMatching-based Vision-Language-Action (VLA) frameworks have demonstrated remarkable advantages in generating high-frequency action chunks, particularly for highly dexterous robotic manipulation tasks. Despite these…

Robotics · Computer Science 2026-03-03 Yang Chen , Xiaoguang Ma , Bin Zhao

Robot learning holds tremendous promise to unlock the full potential of flexible, general, and dexterous robot systems, as well as to address some of the deepest questions in artificial intelligence. However, bringing robot learning to the…

We study the problem of synthesizing a number of likely future frames from a single input image. In contrast to traditional methods that have tackled this problem in a deterministic or non-parametric way, we propose to model future frames…

Computer Vision and Pattern Recognition · Computer Science 2019-08-13 Tianfan Xue , Jiajun Wu , Katherine L. Bouman , William T. Freeman

Vision-language models (VLMs) have demonstrated strong performance in 2D scene understanding and generation, but extending this unification to the physical world remains an open challenge. Existing 3D and 4D approaches typically embed scene…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Hanyu Zhou , Gim Hee Lee

Long video generation has gained increasing attention due to its widespread applications in fields such as entertainment and simulation. Despite advances, synthesizing temporally coherent and visually compelling long sequences remains a…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Jiahao Chen , Hangjie Yuan , Yichen Qian , Jingyun Liang , Jiazheng Xing , Pengwei Liu , Weihua Chen , Fan Wang , Bing Su

Expanding on neural operators, we propose a novel framework for stochastic process learning across arbitrary domains. In particular, we develop operator flow matching (OFM) for learning stochastic process priors on function spaces. OFM…

Machine Learning · Computer Science 2025-10-14 Yaozhong Shi , Zachary E. Ross , Domniki Asimaki , Kamyar Azizzadenesheli

Current 4D generation methods have achieved noteworthy efficacy with the aid of advanced diffusion generative models. However, these methods lack multi-view spatial-temporal modeling and encounter challenges in integrating diverse prior…

Computer Vision and Pattern Recognition · Computer Science 2024-10-23 Haiyu Zhang , Xinyuan Chen , Yaohui Wang , Xihui Liu , Yunhong Wang , Yu Qiao

Learning a generalist embodied agent capable of completing multiple tasks poses challenges, primarily stemming from the scarcity of action-labeled robotic datasets. In contrast, a vast amount of human videos exist, capturing intricate tasks…

Machine Learning · Computer Science 2024-10-10 Haoran He , Chenjia Bai , Ling Pan , Weinan Zhang , Bin Zhao , Xuelong Li

Imitation Learning (IL) enables robots to learn complex skills from demonstrations without explicit task modeling, but it typically requires large amounts of demonstrations, creating significant collection costs. Prior work has investigated…

Robotics · Computer Science 2026-03-02 Runze Tang , Penny Sweetser

Recent Vision-Language-Action (VLA) models equipped with Flow Matching (FM) action heads achieve state-of-the-art performance in complex robot manipulation. However, the multi-step iterative ODE solving required by FM introduces inference…

Robotics · Computer Science 2026-03-19 Zhou Fang , Jiaqi Wang , Yi Zhou , Qiongfeng Shi

Vision-language-action (VLA) models have shown great potential in building generalist robots, but still face a dilemma-misalignment of 2D image forecasting and 3D action prediction. Besides, such a vision-action entangled training manner…

Robotics · Computer Science 2026-04-21 Wenyao Zhang , Bozhou Zhang , Zekun Qi , Wenjun Zeng , Xin Jin , Li Zhang

As virtual reality gains popularity, the demand for controllable creation of immersive and dynamic omnidirectional videos (ODVs) is increasing. While previous text-to-ODV generation methods achieve impressive results, they struggle with…

Computer Vision and Pattern Recognition · Computer Science 2024-12-13 Weiqi Li , Shijie Zhao , Chong Mou , Xuhan Sheng , Zhenyu Zhang , Qian Wang , Junlin Li , Li Zhang , Jian Zhang

Although Model Predictive Control (MPC) can effectively predict the future states of a system and thus is widely used in robotic manipulation tasks, it does not have the capability of environmental perception, leading to the failure in some…

Robotics · Computer Science 2024-07-16 Wentao Zhao , Jiaming Chen , Ziyu Meng , Donghui Mao , Ran Song , Wei Zhang

Learning with neural networks from a continuous stream of visual information presents several challenges due to the non-i.i.d. nature of the data. However, it also offers novel opportunities to develop representations that are consistent…

Computer Vision and Pattern Recognition · Computer Science 2024-09-19 Simone Marullo , Matteo Tiezzi , Marco Gori , Stefano Melacci

Generative modeling aims to transform random noise into structured outputs. In this work, we enhance video diffusion models by allowing motion control via structured latent noise sampling. This is achieved by just a change in data: we…

Computer Vision and Pattern Recognition · Computer Science 2025-08-07 Ryan Burgert , Yuancheng Xu , Wenqi Xian , Oliver Pilarski , Pascal Clausen , Mingming He , Li Ma , Yitong Deng , Lingxiao Li , Mohsen Mousavi , Michael Ryoo , Paul Debevec , Ning Yu

Talking head generation is a significant research topic that still faces numerous challenges. Previous works often adopt generative adversarial networks or regression models, which are plagued by generation quality and average facial shape…

Computer Vision and Pattern Recognition · Computer Science 2024-08-20 Ziyu Yao , Xuxin Cheng , Zhiqi Huang

Text-conditioned diffusion models have emerged as a promising tool for neural video generation. However, current models still struggle with intricate spatiotemporal prompts and often generate restricted or incorrect motion. To address these…

Computer Vision and Pattern Recognition · Computer Science 2024-05-07 Long Lian , Baifeng Shi , Adam Yala , Trevor Darrell , Boyi Li

Although Multimodal Large Language Models (MLLMs) demonstrate strong omni-modal perception, their ability to forecast future events from audio-visual cues remains largely unexplored, as existing benchmarks focus mainly on retrospective…

Computation and Language · Computer Science 2026-01-21 Qian Chen , Jinlan Fu , Changsong Li , See-Kiong Ng , Xipeng Qiu

Most prior motion prediction endeavors in autonomous driving have inadequately encoded future scenarios, leading to predictions that may fail to accurately capture the diverse movements of agents (e.g., vehicles or pedestrians). To address…

Computer Vision and Pattern Recognition · Computer Science 2024-06-21 Mingkun Wang , Xiaoguang Ren , Ruochun Jin , Minglong Li , Xiaochuan Zhang , Changqian Yu , Mingxu Wang , Wenjing Yang

We present a neural network structure, FramePack, to train next-frame (or next-frame-section) prediction models for video generation. FramePack compresses input frame contexts with frame-wise importance so that more frames can be encoded…

Computer Vision and Pattern Recognition · Computer Science 2025-10-16 Lvmin Zhang , Shengqu Cai , Muyang Li , Gordon Wetzstein , Maneesh Agrawala
‹ Prev 1 8 9 10 Next ›