English
Related papers

Related papers: OA-WAM: Object-Addressable World Action Model for …

200 papers

Recently, by introducing large-scale dataset and strong transformer network, video-language pre-training has shown great success especially for retrieval. Yet, existing video-language transformer models do not explicitly fine-grained…

Computer Vision and Pattern Recognition · Computer Science 2022-05-19 Alex Jinpeng Wang , Yixiao Ge , Guanyu Cai , Rui Yan , Xudong Lin , Ying Shan , Xiaohu Qie , Mike Zheng Shou

Interactive perception enables robots to manipulate the environment and objects to bring them into states that benefit the perception process. Deformable objects pose challenges to this due to significant manipulation difficulty and…

Vision-language-action (VLA) models show potential for general robotic tasks, but remain challenging in spatiotemporally coherent manipulation, which requires fine-grained representations. Typically, existing methods embed 3D positions into…

Computer Vision and Pattern Recognition · Computer Science 2025-11-24 Hanyu Zhou , Chuanhao Ma , Gim Hee Lee

Magnetically actuated microrobots have been used as wireless, non-contact manipulation tools at microscales, making them promising for minimally invasive applications. However, their control remains challenging due to indirect actuation,…

Robotics · Computer Science 2026-05-28 Yongchen Wang , Kangyi Lu , Lan Wei , Dandan Zhang

Action-conditioned robot world models generate future video frames of the manipulated scene given a robot action sequence, offering a promising alternative for simulating tasks that are difficult to model with traditional physics engines.…

Robotics · Computer Science 2026-03-27 Jai Bardhan , Patrik Drozdik , Josef Sivic , Vladimir Petrik

The advancement of large Vision-Language-Action (VLA) models has significantly improved robotic manipulation in terms of language-guided task execution and generalization to unseen scenarios. While existing VLAs adapted from pretrained…

Object-centric world models (OCWM) aim to decompose visual scenes into object-level representations, providing structured abstractions that could improve compositional generalization and data efficiency in reinforcement learning. We…

Artificial Intelligence · Computer Science 2025-11-12 Stefano Ferraro , Akihiro Nakano , Masahiro Suzuki , Yutaka Matsuo

Recent vision-language-action (VLA) models have significantly advanced robotic manipulation by unifying perception, reasoning, and control. To achieve such integration, recent studies adopt a predictive paradigm that models future visual…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Yijie Zhu , Jie He , Rui Shao , Kaishen Yuan , Tao Tan , Xiaochen Yuan , Zitong Yu

Joint optimization of poses and features has been extensively studied and demonstrated to yield more accurate results in feature-based SLAM problems. However, research on jointly optimizing poses and non-feature-based maps remains limited.…

Robotics · Computer Science 2025-07-14 Yingyu Wang , Liang Zhao , Shoudong Huang

Reinforcement learning (RL) can refine Vision-Language-Action (VLA) policies beyond behavior cloning, but real-world RL remains expensive due to extensive rollouts, resets, supervision, and safety risks. Action-conditioned video world…

Robotics · Computer Science 2026-05-26 Xiaokang Liu , Zechen Bai , Hai Ci , Kevin Yuchen Ma , Mike Zheng Shou

Prevalent Vision-Language-Action (VLA) models are typically built upon Multimodal Large Language Models (MLLMs) and demonstrate exceptional proficiency in semantic understanding, but they inherently lack the capability to deduce physical…

This paper focuses on the problem of predicting the future position of a target road user given its current state, consisting of position and velocity. A weighted average approach is adopted, where the weights are determined from data…

Computational Engineering, Finance, and Science · Computer Science 2022-04-22 Angelos Toytziaridis , Paolo Falcone , Jonas Sjöberg

Generalizable articulated object manipulation is essential for home-assistant robots. Recent efforts focus on imitation learning from demonstrations or reinforcement learning in simulation, however, due to the prohibitive costs of…

Robotics · Computer Science 2024-02-22 Wenke Xia , Dong Wang , Xincheng Pang , Zhigang Wang , Bin Zhao , Di Hu , Xuelong Li

The binding problem in artificial neural networks is actively explored with the goal of achieving human-level recognition skills through the comprehension of the world in terms of symbol-like entities. Especially in the field of computer…

Computer Vision and Pattern Recognition · Computer Science 2024-03-05 Jinwoo Kim , Janghyuk Choi , Jaehyun Kang , Changyeon Lee , Ho-Jin Choi , Seon Joo Kim

Imitation learning for mobile manipulation is a key challenge in the field of robotic manipulation. However, current mobile manipulation frameworks typically decouple navigation and manipulation, executing manipulation only after reaching a…

Robotics · Computer Science 2025-07-16 Wang Zhicheng , Satoshi Yagi , Satoshi Yamamori , Jun Morimoto

Typical end-to-end formulations for learning robotic navigation involve predicting a small set of steering command actions (e.g., step forward, turn left, turn right, etc.) from images of the current state (e.g., a bird's-eye view of a SLAM…

Robotics · Computer Science 2020-10-13 Jimmy Wu , Xingyuan Sun , Andy Zeng , Shuran Song , Johnny Lee , Szymon Rusinkiewicz , Thomas Funkhouser

Integrating AI into the physical layer is a cornerstone of 6G networks. However, current data-driven approaches struggle to generalize across dynamic environments because they lack an intrinsic understanding of electromagnetic wave…

Networking and Internet Architecture · Computer Science 2026-03-27 Ziqi Chen , Yi Ren , Yixuan Huang , Qi Sun , Nan Li , Yuhong Huang , Chih-Lin I , Yifan Li , Liang Xia

Prevailing Vision-Language-Action Models (VLAs) for robotic manipulation are built upon vision-language backbones pretrained on large-scale, but disconnected static web data. As a result, despite improved semantic generalization, the policy…

Robotics · Computer Science 2025-12-22 Jonas Pai , Liam Achenbach , Victoriano Montesinos , Benedek Forrai , Oier Mees , Elvis Nava

Autonomous Vehicle (AV) perception systems require more than simply seeing, via e.g., object detection or scene segmentation. They need a holistic understanding of what is happening within the scene for safe interaction with other road…

Computer Vision and Pattern Recognition · Computer Science 2024-11-11 Salman Khan , Izzeddin Teeti , Reza Javanmard Alitappeh , Mihaela C. Stoian , Eleonora Giunchiglia , Gurkirt Singh , Andrew Bradley , Fabio Cuzzolin

Modeling dexterous hand-object interactions is challenging as it requires understanding how subtle finger motions influence the environment through contact with objects. While recent world models address interaction modeling, they typically…

‹ Prev 1 8 9 10 Next ›