English
Related papers

Related papers: Vision in Action: Learning Active Perception from …

200 papers

Unmanned Aerial Vehicles (UAV) have been standing out due to the wide range of applications in which they can be used autonomously. However, they need intelligent systems capable of providing a greater understanding of what they perceive to…

Robotics · Computer Science 2022-09-15 Matheus G. Mateus , Ricardo B. Grando , Paulo L. J. Drews-Jr

Vision-Language-Action (VLA) models have emerged as a popular paradigm for learning robot manipulation policies that can follow language instructions and generalize to novel scenarios. Recent works have begun to explore the incorporation of…

Robot teleoperation gains great success in various situations, including chemical pollution rescue, disaster relief, and long-distance manipulation. In this article, we propose a virtual reality (VR) based robot teleoperation system to…

Robotics · Computer Science 2023-08-03 Lingxiao Meng , Jiangshan Liu , Wei Chai , Jiankun Wang , Max Q. -H. Meng

This research introduces the Bi-VLA (Vision-Language-Action) model, a novel system designed for bimanual robotic dexterous manipulation that seamlessly integrates vision for scene understanding, language comprehension for translating human…

Designing robotic tasks for co-manipulation necessitates to exploit not only proprioceptive but also exteroceptive information for improved safety and autonomy. Following such instinct, this research proposes to formulate intuitive robotic…

Systems and Control · Electrical Eng. & Systems 2021-04-02 Sunny Katyara , Fanny Ficuciello , Tao Teng , Fei Chen , Bruno Siciliano , Darwin G. Caldwell

A common vision from science fiction is that robots will one day inhabit our physical spaces, sense the world as we do, assist our physical labours, and communicate with us through natural language. Here we study how to design artificial…

Vision-Language-Action (VLA) models extend vision-language models to embodied control by mapping natural-language instructions and visual observations to robot actions. Despite their capabilities, VLA systems face significant challenges due…

Robotics · Computer Science 2025-10-24 Weifan Guan , Qinghao Hu , Aosheng Li , Jian Cheng

We present a pixel-based deep active inference algorithm (PixelAI) inspired by human body perception and action. Our algorithm combines the free-energy principle from neuroscience, rooted in variational inference, with deep convolutional…

Computer Vision and Pattern Recognition · Computer Science 2021-02-08 Cansu Sancaktar , Marcel van Gerven , Pablo Lanillos

This paper presents a vision-based learning-by-demonstration approach to enable robots to learn and complete a manipulation task cooperatively. With this method, a vision system is involved in both the task demonstration and reproduction…

Robotics · Computer Science 2017-06-05 Bidan Huang , Menglong Ye , Su-Lin Lee , Guang-Zhong Yang

We present O2A, a novel method for learning to perform robotic manipulation tasks from a single (one-shot) third-person demonstration video. To our knowledge, it is the first time this has been done for a single demonstration. The key…

Robotics · Computer Science 2021-08-05 Leo Pauly , Wisdom C. Agboh , David C. Hogg , Raul Fuentes

Human perception of surroundings is often guided by the various poses present within the environment. Many computer vision tasks, such as human action recognition and robot imitation learning, rely on pose-based entities like human…

Computer Vision and Pattern Recognition · Computer Science 2023-06-16 Dominick Reilly , Aman Chadha , Srijan Das

Providing artificial agents with the same computational models of biological systems is a way to understand how intelligent behaviours may emerge. We present an active inference body perception and action model working for the first time in…

Robotics · Computer Science 2021-02-08 Guillermo Oliver , Pablo Lanillos , Gordon Cheng

In real-world scenarios, many robotic manipulation tasks are hindered by occlusions and limited fields of view, posing significant challenges for passive observation-based models that rely on fixed or wrist-mounted cameras. In this paper,…

Robotics · Computer Science 2025-02-13 Guokang Wang , Hang Li , Shuyuan Zhang , Di Guo , Yanhong Liu , Huaping Liu

The paper focuses on an immersive teleoperation system that enhances operator's ability to actively perceive the robot's surroundings. A consumer-grade HTC Vive VR system was used to synchronize the operator's hand and head movements with a…

A fundamental challenge of shared autonomy is to use high-DoF robots to assist, rather than hinder, humans by first inferring user intent and then empowering the user to achieve their intent. Although successful, prior methods either rely…

Robotics · Computer Science 2025-01-16 Atharv Belsare , Zohre Karimi , Connor Mattson , Daniel S. Brown

In robot learning, Vision Transformers (ViTs) are standard for visual perception, yet most methods discard valuable information by using only the final layer's features. We argue this provides an insufficient representation and propose the…

Computer Vision and Pattern Recognition · Computer Science 2026-02-02 Wenhao Li , Chengwei Ma , Weixin Mao

Learning from human demonstration is an effective approach for learning complex manipulation skills. However, existing approaches heavily focus on learning from passive human demonstration data for its simplicity in data collection.…

Robotics · Computer Science 2025-03-12 Philipp Wu , Yide Shentu , Qiayuan Liao , Ding Jin , Menglong Guo , Koushil Sreenath , Xingyu Lin , Pieter Abbeel

Can we turn a video prediction model into a robot policy? Videos, including those of humans or teleoperated robots, capture rich physical interactions. However, most of them lack labeled actions, which limits their use in robot learning. We…

Robotics · Computer Science 2026-03-31 Sandeep Routray , Hengkai Pan , Unnat Jain , Shikhar Bahl , Deepak Pathak

While imitation learning for vision based autonomous mobile robot navigation has recently received a great deal of attention in the research community, existing approaches typically require state action demonstrations that were gathered…

Robotics · Computer Science 2022-03-30 Haresh Karnan , Garrett Warnell , Xuesu Xiao , Peter Stone

Egocentric human videos provide a scalable source of manipulation demonstrations; however, deploying them on robots requires active viewpoint control to maintain task-critical visibility, which human viewpoint imitation often fails to…

Robotics · Computer Science 2026-02-27 Daesol Cho , Youngseok Jang , Danfei Xu , Sehoon Ha