English
Related papers

Related papers: World2Act: Latent Action Post-Training via Skill-C…

200 papers

In this work we study how explicit world-modeling objectives affect the internal representations and downstream capability of Transformers across different training stages. We use a controlled 2x2x2 Rubik's Cube and ask: (1) how does…

Machine Learning · Computer Science 2025-12-04 Prakhar Gupta , Henry Conklin , Sarah-Jane Leslie , Andrew Lee

Generative world models (WMs) can now simulate worlds with striking visual realism, which naturally raises the question of whether they can endow embodied agents with predictive perception for decision making. Progress on this question has…

Vision-Language-Action (VLA) models aim to control robots for manipulation from visual observations and natural-language instructions. However, existing hierarchical and autoregressive paradigms often introduce architectural overhead,…

Large scale pre-training on text and image data along with diverse robot demonstrations has helped Vision Language Action models (VLAs) to generalize to novel tasks, objects and scenes. However, these models are still susceptible to failure…

Robotics · Computer Science 2026-03-24 Kartikay Milind Pangaonkar , Prabin Rath , Omkar Patil , Nakul Gopalan

Dual-system Vision-Language-Action (VLA) models achieve state-of-the-art robotic manipulation but are bottlenecked by the VLM backbone, which must execute at every control step while producing temporally redundant features. We propose…

Vision-language-action (VLA) models increasingly rely on auxiliary world modules to plan over long horizons, yet how such modules should be parameterized on top of a pretrained VLA remains an open design question. Existing…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Zuojin Tang , Shengchao Yuan , Xiaoxin Bai , Zhiyuan Jing , De Ma , Gang Pan , Bin Liu

Vision-Language-Action (VLA) policies excel in aligning language, perception, and robot control. However, most VLAs are trained purely by imitation, which overfits to demonstrations, and is brittle under distribution shift. Reinforcement…

Robotics · Computer Science 2025-11-26 Jiahui Zhang , Ze Huang , Chun Gu , Zipei Ma , Li Zhang

Learning transferable knowledge from unlabeled video data and applying it in new environments is a fundamental capability of intelligent agents. This work presents VideoWorld 2, which extends VideoWorld and offers the first investigation…

Computer Vision and Pattern Recognition · Computer Science 2026-02-11 Zhongwei Ren , Yunchao Wei , Xiao Yu , Guixun Luo , Yao Zhao , Bingyi Kang , Jiashi Feng , Xiaojie Jin

Learned language-conditioned robot policies often struggle to effectively adapt to new real-world tasks even when pre-trained across a diverse set of instructions. We propose a novel approach for few-shot adaptation to unseen tasks that…

Robotics · Computer Science 2025-01-09 Vivek Myers , Bill Chunyuan Zheng , Oier Mees , Sergey Levine , Kuan Fang

Vision-language-action (VLA) models achieve strong in-distribution performance but degrade sharply under novel camera viewpoints and visual perturbations. We show that this brittleness primarily arises from misalignment in Spatial Modeling,…

Robotics · Computer Science 2026-04-01 Weiqi Li , Quande Zhang , Ruifeng Zhai , Liang Lin , Guangrun Wang

Reliable benchmarking is critical for advancing Vision-Language-Action (VLA) models, as it reveals their generalization, robustness, and alignment of perception with language-driven manipulation tasks. However, existing benchmarks often…

Computer Vision and Pattern Recognition · Computer Science 2026-02-09 Guodong Wang , Chenkai Zhang , Qingjie Liu , Jinjin Zhang , Jiancheng Cai , Junjie Liu , Xinmin Liu

Existing LLMs exhibit remarkable performance on various NLP tasks, but still struggle with complex real-world tasks, even equipped with advanced strategies like CoT and ReAct. In this work, we propose the CoAct framework, which transfers…

Computation and Language · Computer Science 2024-06-21 Xinming Hou , Mingming Yang , Wenxiang Jiao , Xing Wang , Zhaopeng Tu , Wayne Xin Zhao

World models have become crucial for autonomous driving, as they learn how scenarios evolve over time to address the long-tail challenges of the real world. However, current approaches relegate world models to limited roles: they operate…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Tianze Xia , Yongkang Li , Lijun Zhou , Jingfeng Yao , Kaixin Xiong , Haiyang Sun , Bing Wang , Kun Ma , Guang Chen , Hangjun Ye , Wenyu Liu , Xinggang Wang

Current research on Vision-Language-Action (VLA) models predominantly focuses on enhancing generalization through established reasoning techniques. While effective, these improvements invariably increase computational complexity and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-06 Riccardo Andrea Izzo , Gianluca Bardaro , Matteo Matteucci

Visual Retrieval-Augmented Generation (VRAG) enhances Vision-Language Models (VLMs) by incorporating external visual documents to address a given query. Existing VRAG frameworks usually depend on rigid, pre-defined external tools to extend…

Artificial Intelligence · Computer Science 2026-04-10 Yuqi Xiong , Chunyi Peng , Zhipeng Xu , Zhenghao Liu , Zulong Chen , Yukun Yan , Shuo Wang , Yu Gu , Ge Yu

Recently, post-training methods based on reinforcement learning, with a particular focus on Group Relative Policy Optimization (GRPO), have emerged as the robust paradigm for further advancement of text-to-image (T2I) models. However, these…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Haoyuan Sun , Jing Wang , Yuxin Song , Yu Lu , Bo Fang , Yifu Luo , Jun Yin , Pengyu Zeng , Miao Zhang , Tiantian Zhang , Xueqian Wang , Shijian Lu

Large language models (LLMs) have demonstrated that large-scale pretraining enables systems to adapt rapidly to new problems with little supervision in the language domain. This success, however, has not translated as effectively to the…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Pablo Acuaviva , Aram Davtyan , Mariam Hassan , Sebastian Stapf , Ahmad Rahimi , Alexandre Alahi , Paolo Favaro

Large Language Models (LLMs) have recently shown strong reasoning and generalization capabilities, motivating their use as decision-making policies in complex environments. StarCraft II (SC2), with its massive state-action space and partial…

Artificial Intelligence · Computer Science 2026-02-17 Yixin Zhang , Ziyi Wang , Yiming Rong , Haoxi Wang , Jinling Jiang , Shuang Xu , Haoran Wu , Shiyu Zhou , Bo Xu

World Action Models (WAMs) are an emerging family of policies that tie robot action generation to future-observation modeling. In this work, we focus on the joint video--action modeling paradigm, where actions and imagined future…

Predictive manipulation has recently gained considerable attention in the Embodied AI community due to its potential to improve robot policy performance by leveraging predicted states. However, generating accurate future visual states of…

Robotics · Computer Science 2025-09-15 Yuhang Huang , Jiazhao Zhang , Shilong Zou , Xinwang Liu , Ruizhen Hu , Kai Xu
‹ Prev 1 8 9 10 Next ›