English
Related papers

Related papers: SAM2Grasp: Resolve Multi-modal Grasping via Prompt…

200 papers

Grasping in cluttered scenes has always been a great challenge for robots, due to the requirement of the ability to well understand the scene and object information. Previous works usually assume that the geometry information of the objects…

Robotics · Computer Science 2021-09-28 Yiming Li , Tao Kong , Ruihang Chu , Yifeng Li , Peng Wang , Lei Li

Promptable foundation models such as the Segment Anything Model (SAM) produce high-quality masks but remain semantically blind, relying on external prompts to specify categories. Existing vision-language approaches address this limitation…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Shayan Jalilian , Abdul Bais

Referring video object segmentation (RVOS) aims to segment objects in a video according to textual descriptions, which requires the integration of multimodal information and temporal dynamics perception. The Segment Anything Model 2 (SAM 2)…

Computer Vision and Pattern Recognition · Computer Science 2025-08-11 Fu Rong , Meng Lan , Qian Zhang , Lefei Zhang

Foundation models like the segment anything model require high-quality manual prompts for medical image segmentation, which is time-consuming and requires expertise. SAM and its variants often fail to segment structures in ultrasound (US)…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Assefa Seyoum Wahd , Banafshe Felfeliyan , Yuyue Zhou , Shrimanti Ghosh , Adam McArthur , Jiechen Zhang , Jacob L. Jaremko , Abhilash Hareendranathan

Goal-conditioned robotic grasping in cluttered environments remains a challenging problem due to occlusions caused by surrounding objects, which prevent direct access to the target object. A promising solution to mitigate this issue is…

Robotics · Computer Science 2025-04-07 Boce Hu , Heng Tian , Dian Wang , Haojie Huang , Xupeng Zhu , Robin Walters , Robert Platt

Robotic grasping is facing a variety of real-world uncertainties caused by non-static object states, unknown object properties, and cluttered object arrangements. The difficulty of grasping increases with the presence of more uncertainties,…

Robotics · Computer Science 2025-09-10 Hao Chen , Takuya Kiyokawa , Weiwei Wan , Kensuke Harada

Task-oriented object grasping and rearrangement are critical skills for robots to accomplish different real-world manipulation tasks. However, they remain challenging due to partial observations of the objects and shape variations in…

Robotics · Computer Science 2026-03-06 Yichen Cai , Jianfeng Gao , Christoph Pohl , Tamim Asfour

We consider the problem of closed-loop robotic grasping and present a novel planner which uses Visual Feedback and an uncertainty-aware Adaptive Sampling strategy (VFAS) to close the loop. At each iteration, our method VFAS-Grasp builds a…

Robotics · Computer Science 2023-10-31 Pedro Piacenza , Jiacheng Yuan , Jinwook Huh , Volkan Isler

Despite significant advancements in robotic manipulation, achieving consistent and stable grasping remains a fundamental challenge, often limiting the successful execution of complex tasks. Our analysis reveals that even state-of-the-art…

Artificial Intelligence · Computer Science 2025-03-20 Sungjae Lee , Yeonjoo Hong , Kwang In Kim

In this work, we tackle the problem of learning universal robotic dexterous grasping from a point cloud observation under a table-top setting. The goal is to grasp and lift up objects in high-quality and diverse ways and generalize across…

Robotic manipulation in dynamic environments often requires seamless transitions between different grasp types to maintain stability and efficiency. However, achieving smooth and adaptive grasp transitions remains a challenge, particularly…

Robotics · Computer Science 2025-09-24 Kuanqi Cai , Chunfeng Wang , Zeqi Li , Haowen Yao , Weinan Chen , Luis Figueredo , Aude Billard , Arash Ajoudani

We focus on the task of language-conditioned grasping in clutter, in which a robot is supposed to grasp the target object based on a language instruction. Previous works separately conduct visual grounding to localize the target object, and…

Robotics · Computer Science 2024-11-01 Kechun Xu , Shuqi Zhao , Zhongxiang Zhou , Zizhang Li , Huaijin Pi , Yue Wang , Rong Xiong

As the basis for prehensile manipulation, it is vital to enable robots to grasp as robustly as humans. Our innate grasping system is prompt, accurate, flexible, and continuous across spatial and temporal domains. Few existing methods cover…

Robotics · Computer Science 2023-06-07 Hao-Shu Fang , Chenxi Wang , Hongjie Fang , Minghao Gou , Jirong Liu , Hengxu Yan , Wenhai Liu , Yichen Xie , Cewu Lu

Objects we interact with and manipulate often share similar parts, such as handles, that allow us to transfer our actions flexibly due to their shared functionality. This work addresses the problem of transferring a grasp experience or a…

Robotics · Computer Science 2023-08-21 Ahmet Tekden , Marc Peter Deisenroth , Yasemin Bekiroglu

6-DoF grasp detection has been a fundamental and challenging problem in robotic vision. While previous works have focused on ensuring grasp stability, they often do not consider human intention conveyed through natural language, hindering…

Robotics · Computer Science 2024-07-26 Toan Nguyen , Minh Nhat Vu , Baoru Huang , An Vuong , Quan Vuong , Ngan Le , Thieu Vo , Anh Nguyen

The ability to grasp objects is an essential skill that enables many robotic manipulation tasks. Recent works have studied point cloud-based methods for object grasping by starting from simulated datasets and have shown promising…

Robotics · Computer Science 2022-06-07 Antonio Alliegro , Martin Rudorfer , Fabio Frattin , Aleš Leonardis , Tatiana Tommasi

Grasping objects with limited or no prior knowledge about them is a highly relevant skill in assistive robotics. Still, in this general setting, it has remained an open problem, especially when it comes to only partial observability and…

Robotics · Computer Science 2026-01-21 Matthias Humt , Dominik Winkelbauer , Ulrich Hillenbrand , Berthold Bäuml

We consider the problem of detecting robotic grasps in an RGB-D view of a scene containing objects. In this work, we apply a deep learning approach to solve this problem, which avoids time-consuming hand-design of features. This presents…

Machine Learning · Computer Science 2014-08-22 Ian Lenz , Honglak Lee , Ashutosh Saxena

Pre-trained on tremendous image-text pairs, vision-language models like CLIP have demonstrated promising zero-shot generalization across numerous image-based tasks. However, extending these capabilities to video tasks remains challenging…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Zichen Liu , Kunlun Xu , Bing Su , Xu Zou , Yuxin Peng , Jiahuan Zhou

Prompt tuning, a recently emerging paradigm, enables the powerful vision-language pre-training models to adapt to downstream tasks in a parameter -- and data -- efficient way, by learning the ``soft prompts'' to condition frozen…

Computer Vision and Pattern Recognition · Computer Science 2023-08-21 Juncheng Li , Minghe Gao , Longhui Wei , Siliang Tang , Wenqiao Zhang , Mengze Li , Wei Ji , Qi Tian , Tat-Seng Chua , Yueting Zhuang