中文
相关论文

相关论文: Vision in Action: Learning Active Perception from …

200 篇论文

While Vision-Language-Action (VLA) models have demonstrated remarkable success in robotic manipulation, their application has largely been confined to low-degree-of-freedom end-effectors performing simple, vision-guided pick-and-place…

机器人学 · 计算机科学 2026-03-10 Tutian Tang , Xingyu Ji , Wanli Xing , Ce Hao , Wenqiang Xu , Lin Shao , Cewu Lu , Qiaojun Yu , Jiangmiao Pang , Kaifeng Zhang

\textbf{BEAVR} is an open-source, bimanual, multi-embodiment Virtual Reality (VR) teleoperation system for robots, designed to unify real-time control, data recording, and policy learning across heterogeneous robotic platforms. BEAVR…

机器人学 · 计算机科学 2025-08-14 Alejandro Posadas-Nava , Alejandro Carrasco , Richard Linares

The operation of telerobotic systems can be a challenging task, requiring intuitive and efficient interfaces to enable inexperienced users to attain a high level of proficiency. Body-Machine Interfaces (BoMI) represent a promising…

机器人学 · 计算机科学 2021-02-02 Matteo Macchini , Manana Lortkipanidze , Fabrizio Schiano , Dario Floreano

Visual imitation learning (VIL) provides an efficient and intuitive strategy for robotic systems to acquire novel skills. Recent advancements in Vision Language Models (VLMs) have demonstrated remarkable performance in vision and language…

机器人学 · 计算机科学 2024-11-01 Guanyan Chen , Meiling Wang , Te Cui , Yao Mu , Haoyang Lu , Tianxing Zhou , Zicai Peng , Mengxiao Hu , Haizhou Li , Yuan Li , Yi Yang , Yufeng Yue

Virtual Reality (VR) interfaces are increasingly used as remote visualization media in telerobotics. Remote environments captured through RGB-D cameras and visualized using VR interfaces can enhance operators' situational awareness and…

人机交互 · 计算机科学 2022-10-12 Y. T. Tefera , D. Mazzanti , S. Anastasi , D. G. Caldwell , P. Fiorini , N. Deshpande

In this paper we present a neurosymbolic architecture for coupling language-guided visual reasoning with robot manipulation. A non-expert human user can prompt the robot using unconstrained natural language, providing a referring expression…

机器人学 · 计算机科学 2025-12-16 Georgios Tziafas , Hamidreza Kasaei

Effective human-robot interaction (HRI) in multi-object teleoperation tasks faces significant challenges due to perceptual ambiguities in virtual reality (VR) environments and the limitations of single-modality intention recognition. This…

机器人学 · 计算机科学 2025-09-03 Chi Sun , Xian Wang , Abhishek Kumar , Chengbin Cui , Lik-Hang Lee

Autonomous driving has long relied on modular "Perception-Decision-Action" pipelines, where hand-crafted interfaces and rule-based components often break down in complex or long-tailed scenarios. Their cascaded design further propagates…

Human videos contain rich manipulation priors, but using them for robot learning remains difficult because raw observations entangle scene understanding, human motion, and embodiment-specific action. We introduce MoT-HRA, a hierarchical…

机器人学 · 计算机科学 2026-05-22 Yifan Xie , YuAn Wang , Guangyu Chen , Jinkun Liu , Yu Sun , Wenbo Ding

For robots to operate effectively and safely alongside humans, they must be able to understand the progress of ongoing actions. This ability, known as action progress prediction, is critical for tasks ranging from timely assistance to…

机器人学 · 计算机科学 2026-03-03 Elena Zoppellari , Federico Becattini , Marco Fiorucci , Lamberto Ballan

Recent vision-language-action (VLA) models and world action models (WAMs) advance robotic manipulation by enriching intermediate representations with auxiliary spatial features or future visual-state prediction. However, these…

机器人学 · 计算机科学 2026-05-26 Xinzhe Chen , Sihua Ren , Liqi Huang , Haowen Sun , Mingyang Li , Xingyu Chen , Zeyang Liu , Xuguang Lan

Active recognition enables robots to intelligently explore novel observations, thereby acquiring more information while circumventing undesired viewing conditions. Recent approaches favor learning policies from simulated or collected data,…

计算机视觉与模式识别 · 计算机科学 2023-11-27 Lei Fan , Mingfu Liang , Yunxuan Li , Gang Hua , Ying Wu

In robotic bimanual teleoperation, multimodal sensory feedback plays a crucial role, providing operators with a more immersive operating experience, reducing cognitive burden, and improving operating efficiency. In this study, we develop an…

人机交互 · 计算机科学 2025-01-03 Han Xu , Mingqi Chen , Gaofeng Li , Lei Wei , Shichi Peng , Haoliang Xu , Qiang Li

One of the fundamental goals of visual perception is to allow agents to meaningfully interact with their environment. In this paper, we take a step towards that long-term goal -- we extract highly localized actionable information related to…

计算机视觉与模式识别 · 计算机科学 2021-08-12 Kaichun Mo , Leonidas Guibas , Mustafa Mukadam , Abhinav Gupta , Shubham Tulsiani

Autonomous manipulation of articulated objects remains a fundamental challenge for robots in human environments. Vision-based methods can infer hidden kinematics but can yield imprecise estimates on unfamiliar objects. Tactile approaches…

机器人学 · 计算机科学 2026-04-03 Leiyao Cui , Zihang Zhao , Sirui Xie , Wenhuan Zhang , Zhi Han , Yixin Zhu

This paper describes MAIA, a Multimodal Automated Interpretability Agent. MAIA is a system that uses neural models to automate neural model understanding tasks like feature interpretation and failure mode discovery. It equips a pre-trained…

Despite advances in Vision-Language-Action (VLA) models, robotic manipulation struggles with fine-grained tasks because current models lack mechanisms for active visual attention allocation. Human gaze naturally encodes intent, planning,…

计算机视觉与模式识别 · 计算机科学 2026-04-08 Anupam Pani , Yanchao Yang

Aerial manipulation (AM) promises to move Unmanned Aerial Vehicles (UAVs) beyond passive inspection to contact-rich tasks such as grasping, assembly, and in-situ maintenance. Most prior AM demonstrations rely on external motion capture…

机器人学 · 计算机科学 2026-03-05 Yuanzhu Zhan , Yufei Jiang , Muqing Cao , Junyi Geng

Visually-guided underwater robots are deployed alongside human divers for cooperative exploration, inspection, and monitoring tasks in numerous shallow-water and coastal-water applications. The most essential capability of such companion…

机器人学 · 计算机科学 2021-07-30 Md Jahidul Islam

Vision and voice are two vital keys for agents' interaction and learning. In this paper, we present a novel indoor navigation model called Memory Vision-Voice Indoor Navigation (MVV-IN), which receives voice commands and analyzes multimodal…

计算机视觉与模式识别 · 计算机科学 2020-09-02 Liqi Yan , Dongfang Liu , Yaoxian Song , Changbin Yu