English
Related papers

Related papers: Free-form language-based robotic reasoning and gra…

200 papers

Grasping large and flat objects (e.g. a book or a pan) is often regarded as an ungraspable task, which poses significant challenges due to the unreachable grasping poses. Previous works leverage Extrinsic Dexterity like walls or table edges…

Robotics · Computer Science 2024-12-12 Chengzhong Ma , Houxue Yang , Hanbo Zhang , Zeyang Liu , Chao Zhao , Jian Tang , Xuguang Lan , Nanning Zheng

Open-Vocabulary Object Detection (OVOD) aims to develop the capability to detect anything. Although myriads of large-scale pre-training efforts have built versatile foundation models that exhibit impressive zero-shot capabilities to…

Computer Vision and Pattern Recognition · Computer Science 2026-01-22 Guiying Zhu , Bowen Yang , Yin Zhuang , Tong Zhang , Guanqun Wang , Zhihao Che , He Chen , Lianlin Li

Generalist robots that can perform a range of different tasks in open-world settings must be able to not only reason about the steps needed to accomplish their goals, but also process complex instructions, prompts, and even feedback during…

In order to explore robotic grasping in unstructured and dynamic environments, this work addresses the visual perception phase involved in the task. This phase involves the processing of visual data to obtain the location of the object to…

Robotics · Computer Science 2021-03-02 Eduardo Godinho Ribeiro , Raul de Queiroz Mendes , Valdir Grassi

Advancing robotic grasping and manipulation requires the ability to test algorithms and/or train learning models on large numbers of grasps. Towards the goal of more advanced grasping, we present the Grasp Reset Mechanism (GRM), a fully…

Robotics · Computer Science 2024-03-01 Kyle DuFrene , Keegan Nave , Joshua Campbell , Ravi Balasubramanian , Cindy Grimm

A challenge in robot grasping is to achieve task-grasping which is to select a grasp that is advantageous to the success of tasks before and after grasps. One of the frameworks to address this difficulty is Learning-from-Observation (LfO),…

Robotics · Computer Science 2022-03-03 Daichi Saito , Kazuhiro Sasabuchi , Naoki Wake , Jun Takamatsu , Hideki Koike , Katsushi Ikeuchi

Remarkable progress has been made in recent years in the fields of vision, language, and robotics. We now have vision models capable of recognizing objects based on language queries, navigation systems that can effectively control mobile…

Robotics · Computer Science 2024-11-20 Peiqi Liu , Yaswanth Orru , Jay Vakil , Chris Paxton , Nur Muhammad Mahi Shafiullah , Lerrel Pinto

Task-oriented grasping (TOG) is crucial for robots to accomplish manipulation tasks, requiring the determination of TOG positions and directions. Existing methods either rely on costly manual TOG annotations or only extract coarse grasping…

Robotics · Computer Science 2024-09-25 Wenlong Dong , Dehao Huang , Jiangshan Liu , Chao Tang , Hong Zhang

Grasping has been a long-standing challenge in facilitating the final interface between a robot and the environment. As environments and tasks become complicated, the need to embed higher intelligence to infer from the surroundings and act…

Robotics · Computer Science 2025-08-14 Navin Sriram Ravie , Keerthi Vasan M , Asokan Thondiyath , Bijo Sebastian

In recent years, the field of computer vision has seen significant advancements thanks to the development of large language models (LLMs). These models have enabled more effective and sophisticated interactions between humans and machines,…

Computer Vision and Pattern Recognition · Computer Science 2023-05-25 Renjie Pi , Jiahui Gao , Shizhe Diao , Rui Pan , Hanze Dong , Jipeng Zhang , Lewei Yao , Jianhua Han , Hang Xu , Lingpeng Kong , Tong Zhang

As the basis for prehensile manipulation, it is vital to enable robots to grasp as robustly as humans. Our innate grasping system is prompt, accurate, flexible, and continuous across spatial and temporal domains. Few existing methods cover…

Robotics · Computer Science 2023-06-07 Hao-Shu Fang , Chenxi Wang , Hongjie Fang , Minghao Gou , Jirong Liu , Hengxu Yan , Wenhai Liu , Yichen Xie , Cewu Lu

Recent advancements in robotic grasping have led to its integration as a core module in many manipulation systems. For instance, language-driven semantic segmentation enables the grasping of any designated object or object part. However,…

Robotics · Computer Science 2025-07-09 Yun Du , Mengao Zhao , Tianwei Lin , Yiwei Jin , Chaodong Huang , Zhizhong Su

Inferring affordable (i.e., graspable) parts of arbitrary objects based on human specifications is essential for robots advancing toward open-vocabulary manipulation. Current grasp planners, however, are hindered by limited vision-language…

Robotics · Computer Science 2025-05-02 Teli Ma , Zifan Wang , Jiaming Zhou , Mengmeng Wang , Junwei Liang

The ability to robustly grasp a variety of objects is essential for dexterous robots. In this paper, we present a framework for zero-shot dynamic dexterous grasping using single-view visual inputs, designed to be resilient to various…

Robotics · Computer Science 2025-08-15 Hui Zhang , Zijian Wu , Linyi Huang , Sammy Christen , Jie Song

Geospatial reasoning requires models to resolve complex spatial semantics and user intent into precise target locations for Earth observation. Recent progress has liberated the reasoning path from manual curation, allowing models to…

Computer Vision and Pattern Recognition · Computer Science 2026-05-07 Liang Yao , Fan Liu , Shengxiang Xu , Chuanyi Zhang , Rui Min , Shimin Di , Yuhui Zheng

A significant challenge for real-world robotic manipulation is the effective 6DoF grasping of objects in cluttered scenes from any single viewpoint without the need for additional scene exploration. This work reinterprets grasping as…

Robotics · Computer Science 2024-05-30 Snehal Jauhri , Ishikaa Lunawat , Georgia Chalvatzaki

Interactive Object Grasping (IOG) is the task of identifying and grasping the desired object via human-robot natural language interaction. Current IOG systems assume that a human user initially specifies the target object's category (e.g.,…

Computation and Language · Computer Science 2024-04-08 Gi-Cheon Kang , Junghyun Kim , Jaein Kim , Byoung-Tak Zhang

Recent advances in training-free visual prompting, such as Set-of-Mark, have emerged as a promising direction for enhancing the grounding capabilities of multimodal language models (MLMs). These techniques operate by partitioning the input…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Giacomo Frisoni , Lorenzo Molfetta , Mattia Buzzoni , Gianluca Moro

3D visual grounding (3DVG) identifies objects in 3D scenes from language descriptions. Existing zero-shot approaches leverage 2D vision-language models (VLMs) by converting 3D spatial information (SI) into forms amenable to VLM processing,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-11 Yuanyuan Liu , Haiyang Mei , Dongyang Zhan , Jiayue Zhao , Dongsheng Zhou , Bo Dong , Xin Yang

Bridging the gap between natural language commands and autonomous execution in unstructured environments remains an open challenge for robotics. This requires robots to perceive and reason over the current task scene through multiple…

Robotics · Computer Science 2025-12-23 Jin Wang , Kim Tien Ly , Jacques Cloete , Nikos Tsagarakis , Ioannis Havoutis