English
Related papers

Related papers: DiLA: Disentangled Latent Action World Models

200 papers

Recent advances in Vision-Language-Action (VLA) models demonstrate that visual signals can effectively complement sparse action supervisions. However, letting VLA directly predict high-dimensional visual states can distribute model capacity…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Yi Yang , Xueqi Li , Yiyang Chen , Jin Song , Yihan Wang , Zipeng Xiao , Jiadi Su , You Qiaoben , Pengfei Liu , Zhijie Deng

End-to-end (E2E) autonomous driving has recently attracted increasing interest in unifying Vision-Language-Action (VLA) with World Models to enhance decision-making and forward-looking imagination. However, existing methods fail to…

Computer Vision and Pattern Recognition · Computer Science 2026-02-09 Feiyang jia , Lin Liu , Ziying Song , Caiyan Jia , Hangjun Ye , Xiaoshuai Hao , Long Chen

A world model enables an intelligent agent to imagine, predict, and reason about how the world evolves in response to its actions, and accordingly to plan and strategize. While recent video generation models produce realistic visual…

Generative video models, a leading approach to world modeling, face fundamental limitations. They often violate physical and logical rules, lack interactivity, and operate as opaque black boxes ill-suited for building structured, queryable…

Computer Vision and Pattern Recognition · Computer Science 2025-12-15 Felix O'Mahony , Roberto Cipolla , Ayush Tewari

In this work, we present a novel approach to multi-view action recognition where we guide learned action representations to be separated from view-relevant information in a video. When trying to classify action instances captured from…

Computer Vision and Pattern Recognition · Computer Science 2023-12-12 Nyle Siddiqui , Praveen Tirupattur , Mubarak Shah

Recent advances in training-free video editing have enabled lightweight and precise cross-frame generation by leveraging pre-trained text-to-image diffusion models. However, existing methods often rely on heuristic frame selection to…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Zhangkai Wu , Xuhui Fan , Zhongyuan Xie , Kaize Shi , Longbing Cao

Recent advances in diffusion transformers have empowered video generation models to generate high-quality video clips from texts or images. However, world models with the ability to predict long-horizon futures from past observations and…

Computer Vision and Pattern Recognition · Computer Science 2026-01-28 Yixuan Zhu , Jiaqi Feng , Wenzhao Zheng , Yuan Gao , Xin Tao , Pengfei Wan , Jie Zhou , Jiwen Lu

We introduce \textbf{LaMP}, a dual-expert Vision-Language-Action framework that embeds dense 3D scene flow as a latent motion prior for robotic manipulation. Existing VLA models regress actions directly from 2D semantic visual features,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Xinkai Wang , Chenyi Wang , Yifu Xu , Mingzhe Ye , Fu-Cheng Zhang , Jialin Tian , Xinyu Zhan , Lifeng Zhu , Cewu Lu , Lixin Yang

Leveraging vast amounts of unlabeled internet video data for embodied AI is currently bottlenecked by the lack of action labels and the presence of action-correlated visual distractors. Although recent latent action policy optimization…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Albina Klepach , Alexander Nikulin , Ilya Zisman , Denis Tarasov , Alexander Derevyagin , Andrei Polubarov , Nikita Lyubaykin , Igor Kiselev , Vladislav Kurenkov

Leveraging future observation modeling to facilitate action generation presents a promising avenue for enhancing the capabilities of Vision-Language-Action (VLA) models. However, existing approaches struggle to strike a balance between…

Vision-Language-Action (VLA) models aim to control robots for manipulation from visual observations and natural-language instructions. However, existing hierarchical and autoregressive paradigms often introduce architectural overhead,…

Disentangled representation learning (DRL) aims to break down observed data into core intrinsic factors for a profound understanding of the data. In real-world scenarios, manually defining and labeling these factors are non-trivial, making…

Machine Learning · Computer Science 2024-11-01 Youngjun Jun , Jiwoo Park , Kyobin Choo , Tae Eun Choi , Seong Jae Hwang

Vision-Language-Action (VLA) models have achieved remarkable progress in robotic manipulation by mapping multimodal observations and instructions directly to actions. However, they typically mimic expert trajectories without predictive…

Deploying safety-critical agents requires anticipating the consequences of actions before they are executed. While world models offer a paradigm for this proactive foresight, current approaches relying on visual simulation incur prohibitive…

Learning to disentangle the hidden factors of variations within a set of observations is a key task for artificial intelligence. We present a unified formulation for class and content disentanglement and use it to illustrate the limitations…

Machine Learning · Computer Science 2020-02-19 Aviv Gabbay , Yedid Hoshen

Pre-training large models on vast amounts of web data has proven to be an effective approach for obtaining powerful, general models in domains such as language and vision. However, this paradigm has not yet taken hold in reinforcement…

Machine Learning · Computer Science 2024-03-28 Dominik Schmidt , Minqi Jiang

Contrastive language-image pretraining (CLIP) has demonstrated remarkable success in various image tasks. However, how to extend CLIP with effective temporal modeling is still an open and crucial problem. Existing factorized or joint…

Computer Vision and Pattern Recognition · Computer Science 2023-08-16 Shuyuan Tu , Qi Dai , Zuxuan Wu , Zhi-Qi Cheng , Han Hu , Yu-Gang Jiang

Vision-language models (VLMs) have shown strong performance on static visual understanding, yet they still struggle with dynamic spatial reasoning that requires imagining how scenes evolve under egocentric motion. Recent efforts address…

Computer Vision and Pattern Recognition · Computer Science 2026-04-30 Wanyue Zhang , Wenxiang Wu , Wang Xu , Jiaxin Luo , Helu Zhi , Yibin Huang , Shuo Ren , Zitao Liu , Jiajun Zhang

Vision-and-language (V-L) tasks require the system to understand both vision content and natural language, thus learning fine-grained joint representations of vision and language (a.k.a. V-L representations) is of paramount importance.…

Computer Vision and Pattern Recognition · Computer Science 2022-11-01 Fenglin Liu , Xian Wu , Shen Ge , Xuancheng Ren , Wei Fan , Xu Sun , Yuexian Zou

Effective interaction modeling and behavior prediction of dynamic agents play a significant role in interactive motion planning for autonomous robots. Although existing methods have improved prediction accuracy, few research efforts have…

Robotics · Computer Science 2024-01-09 Victoria M. Dax , Jiachen Li , Enna Sachdeva , Nakul Agarwal , Mykel J. Kochenderfer
‹ Prev 1 3 4 5 6 7 10 Next ›