中文
相关论文

相关论文: Pay attention! - Robustifying a Deep Visuomotor Po…

200 篇论文

By and large, existing computational models of visual attention tacitly assume perfect vision and full access to the stimulus and thereby deviate from foveated biological vision. Moreover, modeling top-down attention is generally reduced to…

计算机视觉与模式识别 · 计算机科学 2022-11-22 Leo Schwinn , Doina Precup , Björn Eskofier , Dario Zanca

Vision-Language-Action (VLA) models have demonstrated significant advantages in robotic manipulation. However, their reliance on vision and language often leads to suboptimal performance in tasks involving visual occlusion, fine-grained…

Reinforcement learning (RL) fine-tuning has shown promise for Vision-Language-Action (VLA) models in robotic manipulation, but deployment-time visual shifts pose practical challenges. A key difficulty is that standard task rewards supervise…

机器人学 · 计算机科学 2026-05-14 Yuanfang Peng , Jingjing Fu , Chuheng Zhang , Li Zhao , Jiang Bian , Mingyu Liu , Ling Zhang , Jun Zhang , Rui Wang

Learned visuomotor policies are capable of performing increasingly complex manipulation tasks. However, most of these policies are trained on data collected from limited robot positions and camera viewpoints. This leads to poor…

机器人学 · 计算机科学 2025-09-29 Jingyun Yang , Isabella Huang , Brandon Vu , Max Bajracharya , Rika Antonova , Jeannette Bohg

We address the problem of safely solving complex bimanual robot manipulation tasks with sparse rewards. Such challenging tasks can be decomposed into sub-tasks that are accomplishable by different robots concurrently or sequentially for…

机器学习 · 计算机科学 2021-10-07 Minghao Zhang , Pingcheng Jian , Yi Wu , Huazhe Xu , Xiaolong Wang

Vision-Language-Action (VLA) models are vulnerable to adversarial attacks, yet universal and transferable attacks remain underexplored, as most existing patches overfit to a single model and fail in black-box settings. To address this gap,…

计算机视觉与模式识别 · 计算机科学 2026-03-11 Hui Lu , Yi Yu , Yiming Yang , Chenyu Yi , Qixin Zhang , Bingquan Shen , Alex C. Kot , Xudong Jiang

Vision-Language-Action models have recently emerged as a powerful paradigm for general-purpose robot learning, enabling agents to map visual observations and natural-language instructions into executable robotic actions. Though popular,…

Real world traffic sign recognition is an important step towards building autonomous vehicles, most of which highly dependent on Deep Neural Networks (DNNs). Recent studies demonstrated that DNNs are surprisingly susceptible to adversarial…

计算机视觉与模式识别 · 计算机科学 2021-08-16 Xinghao Yang , Weifeng Liu , Shengli Zhang , Wei Liu , Dacheng Tao

Humanoid robots have the potential capability to perform a diverse range of manipulation tasks, but this is based on a robust and precise standing controller. Existing methods are either ill-suited to precisely control high-dimensional…

机器人学 · 计算机科学 2025-08-04 Zhenghan Chen , Haocheng Xu , Haodong Zhang , Liang Zhang , He Li , Dongqi Wang , Jiyu Yu , Yifei Yang , Zhongxiang Zhou , Rong Xiong

While Vision-Language-Action (VLA) models generalize well to generic instructions, they struggle with personalized commands such as "bring my cup," where the robot must act on one specific instance among visually similar objects. We study…

机器人学 · 计算机科学 2026-01-30 Sangoh Lee , Sangwoo Mo , Wook-Shin Han

Transformers are state-of-the-art models for a variety of sequence modeling tasks. At their core is an attention function which models pairwise interactions between the inputs at every timestep. While attention is powerful, it does not…

计算与语言 · 计算机科学 2021-03-23 Hao Peng , Nikolaos Pappas , Dani Yogatama , Roy Schwartz , Noah A. Smith , Lingpeng Kong

Contact-rich manipulation requires not only vision-dominant task semantics but also closed-loop reactions to force/torque (F/T) transients. Yet, generative visuomotor policies are typically constrained to low-frequency updates due to…

机器人学 · 计算机科学 2026-03-10 Mingxin Wang , Zhirun Yue , Renhao Lu , Yizhe Li , Zihan Wang , Guoping Pan , Kangkang Dong , Jun Cheng , Yi Cheng , Houde Liu

Imitation learning has demonstrated significant potential in performing high-precision manipulation tasks using visual feedback. However, it is common practice in imitation learning for cameras to be fixed in place, resulting in issues like…

机器人学 · 计算机科学 2025-03-11 Ian Chuang , Andrew Lee , Dechen Gao , M-Mahdi Naddaf-Sh , Iman Soltani

Vision-Language-Action (VLA) models have recently emerged as powerful general-purpose policies for robotic manipulation, benefiting from large-scale multi-modal pre-training. However, they often fail to generalize reliably in…

机器人学 · 计算机科学 2025-12-02 Hongyin Zhang , Shuo Zhang , Junxi Jin , Qixin Zeng , Runze Li , Donglin Wang

Learning generalizable skills in robotic manipulation has long been challenging due to real-world sized observation and action spaces. One method for addressing this problem is attention focus -- the robot learns where to attend its sensors…

机器人学 · 计算机科学 2020-03-05 Marcus Gualtieri , Robert Platt

Vision-language-action (VLA) models trained on large-scale internet data and robot demonstrations have the potential to serve as generalist robot policies. However, despite their large-scale training, VLAs are often brittle to…

机器人学 · 计算机科学 2024-10-04 Asher J. Hancock , Allen Z. Ren , Anirudha Majumdar

Vision--Language--Action (VLA) policies have shown strong progress in mapping language instructions and visual observations to robotic actions, yet their reliability degrades in cluttered scenes with distractors. By analyzing failure cases,…

计算机视觉与模式识别 · 计算机科学 2026-03-26 Jiaying Zhou , Zhihao Zhan , Ruifeng Zhai , Qinhan Lyu , Hao Liu , Keze Wang , Liang Lin , Guangrun Wang

Visuomotor policies learned from demonstrations often overfit to nuisance visual factors in raw RGB observations, resulting in brittle behavior under appearance shifts such as background changes and object recoloring. We propose a…

Attention models have had a significant positive impact on deep learning across a range of tasks. However previous attempts at integrating attention with reinforcement learning have failed to produce significant improvements. We propose the…

机器学习 · 计算机科学 2019-04-09 Anthony Manchin , Ehsan Abbasnejad , Anton van den Hengel

Generalization in robotic manipulation remains a critical challenge, particularly when scaling to new environments with limited demonstrations. This paper introduces CAGE, a novel robotic manipulation policy designed to overcome these…

机器人学 · 计算机科学 2024-12-09 Shangning Xia , Hongjie Fang , Cewu Lu , Hao-Shu Fang