English
Related papers

Related papers: RoboWM-Bench: A Benchmark for Evaluating World Mod…

200 papers

Generalizing language-conditioned robotic policies to new tasks remains a significant challenge, hampered by the lack of suitable simulation benchmarks. In this paper, we address this gap by introducing GemBench, a novel benchmark to assess…

Robotics · Computer Science 2025-03-04 Ricardo Garcia , Shizhe Chen , Cordelia Schmid

Video-Action Models (VAMs) have emerged as a promising framework for embodied intelligence, learning implicit world dynamics from raw video streams to produce temporally consistent action predictions. Although such models demonstrate strong…

Automatically generating training supervision for embodied tasks is crucial, as manual designing is tedious and not scalable. While prior works use large language models (LLMs) or vision-language models (VLMs) to generate rewards, these…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Xiaowen Qiu , Yian Wang , Jiting Cai , Zhehuan Chen , Chunru Lin , Tsun-Hsuan Wang , Chuang Gan

Generative world models are increasingly used for video generation, where learned simulators are expected to capture the physical rules that govern real-world dynamics. However, evaluating whether generated videos actually follow these…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Juyi Lin , Arash Akbari , Yumei He , Lin Zhao , Haichao Zhang , Arman Akbari , Xingchen Xu , Zoe Y. Lu , Enfu Nan , Hokin Deng , Edmund Yeh , Sarah Ostadabbas , Yun Fu , Jennifer Dy , Pu Zhao , Yanzhi Wang

Model-based reinforcement learning (MBRL) has achieved remarkable success in robotics due to its high sample efficiency and planning capability. However, extending MBRL to physical multi-robot cooperation remains challenging due to the…

Robotics · Computer Science 2026-04-07 Zijie Zhao , Honglei Guo , Shengqian Chen , Kaixuan Xu , Bo Jiang , Yuanheng Zhu , Dongbin Zhao

Video-based large language models (Video-LLMs) have been recently introduced, targeting both fundamental improvements in perception and comprehension, and a diverse range of user inquiries. In pursuit of the ultimate goal of achieving…

Computer Vision and Pattern Recognition · Computer Science 2023-11-29 Munan Ning , Bin Zhu , Yujia Xie , Bin Lin , Jiaxi Cui , Lu Yuan , Dongdong Chen , Li Yuan

Embodied AI requires agents that perceive, act, and anticipate how actions reshape future world states. World models serve as internal simulators that capture environment dynamics, enabling forward and counterfactual rollouts to support…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Xinqing Li , Xin He , Le Zhang , Min Wu , Xiaoli Li , Yun Liu

We introduce Robowheel, a data engine that converts human hand object interaction (HOI) videos into training-ready supervision for cross morphology robotic learning. From monocular RGB or RGB-D inputs, we perform high precision HOI…

Understanding the physical world is a fundamental challenge in embodied AI, critical for enabling agents to perform complex tasks and operate safely in real-world environments. While Vision-Language Models (VLMs) have shown great promise in…

Computer Vision and Pattern Recognition · Computer Science 2025-01-30 Wei Chow , Jiageng Mao , Boyi Li , Daniel Seita , Vitor Guizilini , Yue Wang

Post-training is essential for turning pretrained generalist robot policies into reliable task-specific controllers, but existing human-in-the-loop pipelines remain tied to physical execution: each correction requires robot time, scene…

Robotics · Computer Science 2026-05-06 Yaxuan Li , Zhongyi Zhou , Yefei Chen , Yanjiang Guo , Jiaming Liu , Shanghang Zhang , Jianyu Chen , Yichen Zhu

Robotic manipulation policies have made rapid progress in recent years, yet most existing approaches give limited consideration to memory capabilities. Consequently, they struggle to solve tasks that require reasoning over historical…

Utilizing Vision-Language Models (VLMs) for robotic manipulation represents a novel paradigm, aiming to enhance the model's ability to generalize to new objects and instructions. However, due to variations in camera specifications and…

Robotics · Computer Science 2024-09-13 Fanfan Liu , Feng Yan , Liming Zheng , Chengjian Feng , Yiyang Huang , Lin Ma

Learning from visual data opens the potential to accrue a large range of manipulation behaviors by leveraging human demonstrations without specifying each of them mathematically, but rather through natural task specification. In this paper,…

Robotics · Computer Science 2021-11-16 Haoyu Xiong , Quanzhou Li , Yun-Chun Chen , Homanga Bharadhwaj , Samarth Sinha , Animesh Garg

Enabling robots to execute long-horizon manipulation tasks from free-form language instructions remains a fundamental challenge in embodied AI. While vision-language models (VLMs) have shown promise as high-level planners, their deployment…

Robotics · Computer Science 2025-10-01 Zitong Bo , Yue Hu , Jinming Ma , Mingliang Zhou , Junhui Yin , Yachen Kang , Yuqi Liu , Tong Wu , Diyun Xiang , Hao Chen

Learning to execute long-horizon mobile manipulation tasks is crucial for advancing robotics in household and workplace settings. However, current approaches are typically data-inefficient, underscoring the need for improved models that…

Action-conditioned robot world models generate future video frames of the manipulated scene given a robot action sequence, offering a promising alternative for simulating tasks that are difficult to model with traditional physics engines.…

Robotics · Computer Science 2026-03-27 Jai Bardhan , Patrik Drozdik , Josef Sivic , Vladimir Petrik

Humans and animals excel in combining information from multiple sensory modalities, controlling their complex bodies, adapting to growth, failures, or using tools. These capabilities are also highly desirable in robots. They are displayed…

Robotics · Computer Science 2022-11-08 Matej Hoffmann

Recently, there has been a growing interest in rescue robots due to their vital role in addressing emergency scenarios and providing crucial support in challenging or hazardous situations where human intervention is difficult. However, very…

Robotics · Computer Science 2025-05-19 Qianwen Zhao , Rajarshi Roy , Chad Spurlock , Kevin Lister , Long Wang

Simulating robot-world interactions is a cornerstone of Embodied AI. Recently, a few works have shown promise in leveraging video generations to transcend the rigid visual/physical constraints of traditional simulators. However, they…

Robotics · Computer Science 2026-03-18 Mutian Xu , Tianbao Zhang , Tianqi Liu , Zhaoxi Chen , Xiaoguang Han , Ziwei Liu

Automatic identification of events and recurrent behavior analysis are critical for video surveillance. However, most existing content-based video retrieval benchmarks focus on scene-level similarity and do not evaluate the action…

Computer Vision and Pattern Recognition · Computer Science 2026-01-12 Oriol Rabasseda , Zenjie Li , Kamal Nasrollahi , Sergio Escalera