中文
相关论文

相关论文: Large Pre-Trained Models for Bimanual Manipulation…

200 篇论文

Generating 3D models from multi-view 2D RGB images has gained significant attention, extending the capabilities of technologies like Virtual Reality, Robotic Vision, and human-machine interaction. In this paper, we introduce a hybrid…

计算机视觉与模式识别 · 计算机科学 2024-12-03 Ajith Balakrishnan , Sreeja S , Linu Shine

As Vision Transformers (ViTs) are increasingly adopted in sensitive vision applications, there is a growing demand for improved interpretability. This has led to efforts to forward-align these models with carefully annotated abstract,…

计算机视觉与模式识别 · 计算机科学 2025-02-05 Sanchit Sinha , Guangzhi Xiong , Aidong Zhang

Deep imitation learning is promising for solving dexterous manipulation tasks because it does not require an environment model and pre-programmed robot behavior. However, its application to dual-arm manipulation tasks remains challenging.…

机器人学 · 计算机科学 2025-05-23 Heecheol Kim , Yoshiyuki Ohmura , Yasuo Kuniyoshi

The understanding of where humans look in a scene is a problem of great interest in visual perception and computer vision. When eye-tracking devices are not a viable option, models of human attention can be used to predict fixations. In…

计算机视觉与模式识别 · 计算机科学 2018-07-30 Dario Zanca , Marco Gori

A detailed environment representation is a crucial component of automated vehicles. Using single range sensor scans, data is often too sparse and subject to occlusions. Therefore, we present a method to augment occupancy grid maps from…

机器人学 · 计算机科学 2018-12-06 Sascha Wirges , Felix Hartenbach , Christoph Stiller

The recently proposed data augmentation TransMix employs attention labels to help visual transformers (ViT) achieve better robustness and performance. However, TransMix is deficient in two aspects: 1) The image cropping method of TransMix…

计算机视觉与模式识别 · 计算机科学 2023-08-08 Qihao Zhao , Yangyu Huang , Wei Hu , Fan Zhang , Jun Liu

We introduce DinoLizer, a DINOv2-based model for localizing manipulated regions in generative inpainting. Our method builds on a DINOv2 model pretrained to detect synthetic images on the B-Free dataset. We add a linear classification head…

计算机视觉与模式识别 · 计算机科学 2025-11-27 Minh Thong Doi , Jan Butora , Vincent Itier , Jérémie Boulanger , Patrick Bas

Pretrained Vision Transformers (ViTs) such as DINOv2 and MAE provide generic image features that can be applied to a variety of downstream tasks such as retrieval, classification, and segmentation. However, such representations tend to…

计算机视觉与模式识别 · 计算机科学 2026-04-03 Jona Ruthardt , Manu Gaur , Deva Ramanan , Makarand Tapaswi , Yuki M. Asano

The Vision Transformer (ViT) architecture has become widely recognized in computer vision, leveraging its self-attention mechanism to achieve remarkable success across various tasks. Despite its strengths, ViT's optimization remains…

计算机视觉与模式识别 · 计算机科学 2025-08-26 Haoyu Yun , Hamid Krim

We address the challenging problem of learning motion representations using deep models for video recognition. To this end, we make use of attention modules that learn to highlight regions in the video and aggregate features for…

计算机视觉与模式识别 · 计算机科学 2020-08-18 Miao Liu , Xin Chen , Yun Zhang , Yin Li , James M. Rehg

The features of self-supervised vision transformers (ViTs) contain strong semantic and positional information relevant to downstream tasks like object localization and segmentation. Recent works combine these features with traditional…

计算机视觉与模式识别 · 计算机科学 2025-08-07 Ronan Docherty , Antonis Vamvakeros , Samuel J. Cooper

In class incremental learning (CIL) setting, groups of classes are introduced to a model in each learning phase. The goal is to learn a unified model performant on all the classes observed so far. Given the recent popularity of Vision…

计算机视觉与模式识别 · 计算机科学 2023-06-06 Abdelrahman Mohamed , Rushali Grandhe , K J Joseph , Salman Khan , Fahad Khan

Recent advances in computer vision have made it possible to automatically assess from videos the manipulation skills of humans in performing a task, which breeds many important applications in domains such as health rehabilitation and…

计算机视觉与模式识别 · 计算机科学 2019-04-11 Zhenqiang Li , Yifei Huang , Minjie Cai , Yoichi Sato

Vision transformers using self-attention or its proposed alternatives have demonstrated promising results in many image related tasks. However, the underpinning inductive bias of attention is not well understood. To address this issue, this…

机器学习 · 计算机科学 2022-05-23 Arda Sahiner , Tolga Ergen , Batu Ozturkler , John Pauly , Morteza Mardani , Mert Pilanci

Driving in a complex urban environment is a difficult task that requires a complex decision policy. In order to make informed decisions, one needs to gain an understanding of the long-range context and the importance of other vehicles. In…

机器学习 · 计算机科学 2021-09-15 Eshagh Kargar , Ville Kyrki

Effectively utilizing multi-sensory data is important for robots to generalize across diverse tasks. However, the heterogeneous nature of these modalities makes fusion challenging. Existing methods propose strategies to obtain…

机器人学 · 计算机科学 2025-07-22 Jinzhou Li , Tianhao Wu , Jiyao Zhang , Zeyuan Chen , Haotian Jin , Mingdong Wu , Yujun Shen , Yaodong Yang , Hao Dong

Human vision achieves remarkable perceptual performance while operating under strict metabolic constraints. A key ingredient is the selective attention mechanism, driven by rapid saccadic eye movements that constantly reposition the…

计算机视觉与模式识别 · 计算机科学 2026-03-12 Matthis Dallain , Laurent Rodriguez , Laurent Udo Perrinet , Benoît Miramond

Vision Transformers (ViTs) have achieved state-of-the-art results on various computer vision tasks, including 3D object detection. However, their end-to-end implementation also makes ViTs less explainable, which can be a challenge for…

计算机视觉与模式识别 · 计算机科学 2023-12-25 Till Beemelmanns , Wassim Zahr , Lutz Eckstein

We present an open-source tool for visualizing multi-head self-attention in Transformer-based language representation models. The tool extends earlier work by visualizing attention at three levels of granularity: the attention-head level,…

人机交互 · 计算机科学 2019-06-12 Jesse Vig

Vision-language-action models (VLAs) trained on large-scale robotic datasets have demonstrated strong performance on manipulation tasks, including bimanual tasks. However, because most public datasets focus on single-arm demonstrations,…

机器人学 · 计算机科学 2026-02-24 Hokyun Im , Euijin Jeong , Andrey Kolobov , Jianlong Fu , Youngwoon Lee