中文
相关论文

相关论文: Vision in Action: Learning Active Perception from …

200 篇论文

Several recent studies have demonstrated the promise of deep visuomotor policies for robot manipulator control. Despite impressive progress, these systems are known to be vulnerable to physical disturbances, such as accidental or…

机器人学 · 计算机科学 2018-11-30 Pooya Abolghasemi , Amir Mazaheri , Mubarak Shah , Ladislau Bölöni

Existing models of human visual attention are generally unable to incorporate direct task guidance and therefore cannot model an intent or goal when exploring a scene. To integrate guidance of any downstream visual task into attention…

计算机视觉与模式识别 · 计算机科学 2022-11-23 Leo Schwinn , Doina Precup , Bjoern Eskofier , Dario Zanca

Accurate and robust localization is a fundamental need for mobile agents. Visual-inertial odometry (VIO) algorithms exploit the information from camera and inertial sensors to estimate position and translation. Recent deep learning based…

计算机视觉与模式识别 · 计算机科学 2022-09-20 Zheming Tu , Changhao Chen , Xianfei Pan , Ruochen Liu , Jiarui Cui , Jun Mao

Current research on Vision-Language-Action (VLA) models predominantly focuses on enhancing generalization through established reasoning techniques. While effective, these improvements invariably increase computational complexity and…

计算机视觉与模式识别 · 计算机科学 2026-03-06 Riccardo Andrea Izzo , Gianluca Bardaro , Matteo Matteucci

Large-scale, high-quality multimodal demonstrations are essential for robot learning of contact-rich dexterous manipulation. While human-centric data collection systems lower the barrier to scaling, they struggle to capture the tactile…

机器人学 · 计算机科学 2026-03-19 Xitong Chen , Yifeng Pan , Min Li , Xiaotian Ding

This paper addresses the problem of human operator intent recognition during teleoperated robot navigation. In this context, recognition of the operator's intended navigational goal, could enable an artificial intelligence (AI) agent to…

机器人学 · 计算机科学 2021-09-27 Dimitris Panagopoulos , Giannis Petousakis , Rustam Stolkin , Grigoris Nikolaou , Manolis Chiou

This paper presents a novel layered framework that integrates visual foundation models to improve robot manipulation tasks and motion planning. The framework consists of five layers: Perception, Cognition, Planning, Execution, and Learning.…

机器人学 · 计算机科学 2023-09-21 Chen Yang , Peng Zhou , Jiaming Qi

Human actions manipulating articulated objects, such as opening and closing a drawer, can be categorized into multiple modalities we define as interaction modes. Traditional robot learning approaches lack discrete representations of these…

机器人学 · 计算机科学 2024-10-29 Liquan Wang , Ankit Goyal , Haoping Xu , Animesh Garg

Recent advances in embodied intelligence have leveraged massive scaling of data and model parameters to master natural-language command following and multi-task control. In contrast, biological systems demonstrate an innate ability to…

机器人学 · 计算机科学 2026-01-22 Weiyu Guo , He Zhang , Pengteng Li , Tiefu Cai , Ziyang Chen , Yandong Guo , Xiao He , Yongkui Yang , Ying Sun , Hui Xiong

Visual search of relevant targets in the environment is a crucial robot skill. We propose a preliminary framework for the execution monitor of a robot task, taking care of the robot attitude to visually searching the environment for targets…

Navigation in an unknown environment consists of multiple separable subtasks, such as collecting information about the surroundings and navigating to the current goal. In the case of pure visual navigation, all these subtasks need to…

机器人学 · 计算机科学 2016-02-17 Tuomas Välimäki , Risto Ritala

Robotic vision for human-robot interaction and collaboration is a critical process for robots to collect and interpret detailed information related to human actions, goals, and preferences, enabling robots to provide more useful services to…

机器人学 · 计算机科学 2023-07-31 Nicole Robinson , Brendan Tidd , Dylan Campbell , Dana Kulić , Peter Corke

Learning visuomotor control policies in robotic systems is a fundamental problem when aiming for long-term behavioral autonomy. Recent supervised-learning-based vision and motion perception systems, however, are often separately built with…

机器人学 · 计算机科学 2020-06-17 Marvin Chancán , Michael Milford

Visual observations from different viewpoints can significantly influence the performance of visuomotor policies in robotic manipulation. Among these, egocentric (in-hand) views often provide crucial information for precise control.…

机器人学 · 计算机科学 2025-09-22 Haoran Ding , Anqing Duan , Zezhou Sun , Dezhen Song , Yoshihiko Nakamura

Pre-trained perception models excel in generic image domains but degrade significantly in novel environments like indoor scenes. The conventional remedy is fine-tuning on downstream data which incurs catastrophic forgetting of prior…

计算机视觉与模式识别 · 计算机科学 2026-03-02 Tianci Tang , Tielong Cai , Hongwei Wang , Gaoang Wang

Employing a teleoperation system for gathering demonstrations offers the potential for more efficient learning of robot manipulation. However, teleoperating a robot arm equipped with a dexterous hand or gripper, via a teleoperation system…

机器人学 · 计算机科学 2024-10-22 Shengcheng Luo , Quanquan Peng , Jun Lv , Kaiwen Hong , Katherine Rose Driggs-Campbell , Cewu Lu , Yong-Lu Li

In this paper, we address the problem of vision-based obstacle avoidance for robotic manipulators. This topic poses challenges for both perception and motion generation. While most work in the field aims at improving one of those aspects,…

机器人学 · 计算机科学 2020-11-02 Elie Aljalbout , Ji Chen , Konstantin Ritt , Maximilian Ulmer , Sami Haddadin

Human action Recognition for unknown views is a challenging task. We propose a view-invariant deep human action recognition framework, which is a novel integration of two important action cues: motion and shape temporal dynamics (STD). The…

计算机视觉与模式识别 · 计算机科学 2020-01-22 Chhavi Dhiman , Dinesh Kumar Vishwakarma

To have a robot actively supporting a human during a collaborative task, it is crucial that robots are able to identify the current action in order to predict the next one. Common approaches make use of high-level knowledge, such as object…

机器人学 · 计算机科学 2017-03-08 Markus Eich , Sareh Shirazi , Gordon Wyeth

We study the task of embodied visual active learning, where an agent is set to explore a 3d environment with the goal to acquire visual scene understanding by actively selecting views for which to request annotation. While accurate on some…

计算机视觉与模式识别 · 计算机科学 2020-12-18 David Nilsson , Aleksis Pirinen , Erik Gärtner , Cristian Sminchisescu