English
Related papers

Related papers: GAP: Geometric Anchor Pre-training for Data-Effici…

200 papers

Robotic manipulation is essential for the widespread adoption of robots in industrial and home settings and has long been a focus within the robotics community. Advances in artificial intelligence have introduced promising learning-based…

Robotics · Computer Science 2025-03-04 Kelin Li , Shubham M Wagh , Nitish Sharma , Saksham Bhadani , Wei Chen , Chang Liu , Petar Kormushev

Prevailing Vision-Language-Action Models (VLAs) for robotic manipulation are built upon vision-language backbones pretrained on large-scale, but disconnected static web data. As a result, despite improved semantic generalization, the policy…

Robotics · Computer Science 2025-12-22 Jonas Pai , Liam Achenbach , Victoriano Montesinos , Benedek Forrai , Oier Mees , Elvis Nava

Model pre-training is essential in human-centric perception. In this paper, we first introduce masked image modeling (MIM) as a pre-training approach for this task. Upon revisiting the MIM training strategy, we reveal that human structure…

Computer Vision and Pattern Recognition · Computer Science 2023-11-01 Junkun Yuan , Xinyu Zhang , Hao Zhou , Jian Wang , Zhongwei Qiu , Zhiyin Shao , Shaofeng Zhang , Sifan Long , Kun Kuang , Kun Yao , Junyu Han , Errui Ding , Lanfen Lin , Fei Wu , Jingdong Wang

6D robotic grasping beyond top-down bin-picking scenarios is a challenging task. Previous solutions based on 6D grasp synthesis with robot motion planning usually operate in an open-loop setting, which are sensitive to grasp synthesis…

Robotics · Computer Science 2021-09-15 Lirui Wang , Yu Xiang , Wei Yang , Arsalan Mousavian , Dieter Fox

Since current Vision-Language-Action (VLA) systems suffer from limited spatial perception and the absence of memory throughout manipulation, we investigate visual anchors as a means to enhance spatial and temporal reasoning within VLA…

Robotics · Computer Science 2026-03-16 Juan Zhu , Zhanying Shao , Xiaoqi Li , Ethan Morgan , Jiadong Xu , Hongwei Fan , Hao Dong

Vision-Language Models (VLMs) such as CLIP learn a shared embedding space for images and text, yet their representations remain geometrically separated, a phenomenon known as the modality gap. This gap limits tasks requiring cross-modal…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Hongyuan Liu , Qinli Yang , Wen Li , Zhong Zhang , Jiaming Liu , Wei Han , Zhili Qin , Jinxia Guo , Junming Shao

Representation learning approaches for robotic manipulation have boomed in recent years. Due to the scarcity of in-domain robot data, prevailing methodologies tend to leverage large-scale human video datasets to extract generalizable…

Vision-language models like CLIP have demonstrated remarkable zero-shot transfer capabilities. However, their susceptibility to imperceptible adversarial perturbations remains a critical security concern. While test-time defenses offer a…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Zhiwei Li , Jiacheng Xue , Weining Wang , Ajian Liu , Xingyu Gao , Zhenan Sun , Qi Li

Grasping of diverse objects in unstructured environments remains a significant challenge. Open-loop grasping methods, effective in controlled settings, struggle in cluttered environments. Grasp prediction errors and object pose changes…

Brain-machine interfaces (BMIs) help the disabled restore body functions by translating neural activity into digital commands to control external devices. Neural adaptation, where the brain signals change in response to external stimuli or…

Signal Processing · Electrical Eng. & Systems 2021-07-28 Shuhang Chen , Xiang Zhang , Xiang Shen , Yifan Huang , Yiwen Wang

Searching for bindings of geometric parameters in task and motion planning (TAMP) is a finite-horizon stochastic planning problem with high-dimensional decision spaces. A robot manipulator can only move in a subspace of its whole range that…

Robotics · Computer Science 2022-01-25 Tianyu Ren , Alexander Imani Cowen-Rivers , Haitham Bou Ammar , Jan Peters

Grasping is one of the most fundamental challenging capabilities in robotic manipulation, especially in unstructured, cluttered, and semantically diverse environments. Recent researches have increasingly explored language-guided…

Robotics · Computer Science 2025-12-25 Zebin Jiang , Tianle Jin , Xiangtong Yao , Alois Knoll , Hu Cao

Manipulating objects without grasping them is an essential component of human dexterity, referred to as non-prehensile manipulation. Non-prehensile manipulation may enable more complex interactions with the objects, but also presents…

Robotics · Computer Science 2024-07-16 Wenxuan Zhou , Bowen Jiang , Fan Yang , Chris Paxton , David Held

In the realm of future home-assistant robots, 3D articulated object manipulation is essential for enabling robots to interact with their environment. Many existing studies make use of 3D point clouds as the primary input for manipulation…

Robotics · Computer Science 2023-10-16 Xiaoqi Li , Yanzi Wang , Yan Shen , Ponomarenko Iaroslav , Haoran Lu , Qianxu Wang , Boshi An , Jiaming Liu , Hao Dong

Continual learning methods usually preserve old behavior by regularizing parameters, matching old outputs, or replaying previous examples. These strategies can reduce forgetting, but they do not directly specify how the latent…

Machine Learning · Computer Science 2026-03-23 Henry J. Kobs

Generative pre-trained models have demonstrated remarkable effectiveness in language and vision domains by learning useful representations. In this paper, we extend the scope of this effectiveness by showing that visual robot manipulation…

Robotics · Computer Science 2023-12-22 Hongtao Wu , Ya Jing , Chilam Cheang , Guangzeng Chen , Jiafeng Xu , Xinghang Li , Minghuan Liu , Hang Li , Tao Kong

Effectively manipulating articulated objects in household scenarios is a crucial step toward achieving general embodied artificial intelligence. Mainstream research in 3D vision has primarily focused on manipulation through depth perception…

Robotics · Computer Science 2025-03-24 Wenbo Cui , Chengyang Zhao , Songlin Wei , Jiazhao Zhang , Haoran Geng , Yaran Chen , Haoran Li , He Wang

Robotic manipulation in dynamic environments often requires seamless transitions between different grasp types to maintain stability and efficiency. However, achieving smooth and adaptive grasp transitions remains a challenge, particularly…

Robotics · Computer Science 2025-09-24 Kuanqi Cai , Chunfeng Wang , Zeqi Li , Haowen Yao , Weinan Chen , Luis Figueredo , Aude Billard , Arash Ajoudani

Vision Transformers require significant computational resources and memory bandwidth, severely limiting their deployment on edge devices. While recent structured pruning methods successfully reduce theoretical FLOPs, they typically operate…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Andy Li , Aiden Durrant , Milan Markovic , Georgios Leontidis

Robotic grasping under uncertainty remains a fundamental challenge due to its uncertain and contact-rich nature. Traditional rigid robotic hands, with limited degrees of freedom and compliance, rely on complex model-based and heavy feedback…

Robotics · Computer Science 2026-04-06 Liudi Yang , Yang Bai , Yuhao Wang , Ibrahim Alsarraj , Gitta Kutyniok , Zhanchi Wang , Ke Wu