English
Related papers

Related papers: Language-Driven Closed-Loop Grasping with Model-Pr…

200 papers

Robots often face manipulation tasks in environments where vision is inadequate due to clutter, occlusions, or poor lighting--for example, reaching a shutoff valve at the back of a sink cabinet or locating a light switch above a crowded…

Robotics · Computer Science 2025-10-24 Muhammad Suhail Saleem , Lai Yuan , Maxim Likhachev

Pre-trained vision-language models (e.g., CLIP) have shown promising zero-shot generalization in many downstream tasks with properly designed text prompts. Instead of relying on hand-engineered prompts, recent works learn prompts using the…

Computer Vision and Pattern Recognition · Computer Science 2022-09-16 Manli Shu , Weili Nie , De-An Huang , Zhiding Yu , Tom Goldstein , Anima Anandkumar , Chaowei Xiao

Comprehending natural language instructions is a critical skill for robots to cooperate effectively with humans. In this paper, we aim to learn 6D poses for roboticassembly by natural language instructions. For this purpose,…

Robotics · Computer Science 2023-10-24 Bowen Fu , Sek Kun Leong , Yan Di , Jiwen Tang , Xiangyang Ji

In this paper, we propose a training-free framework for vision-and-language navigation (VLN). Existing zero-shot VLN methods are mainly designed for discrete environments or involve unsupervised training in continuous simulator…

Robotics · Computer Science 2025-09-15 Hang Yin , Haoyu Wei , Xiuwei Xu , Wenxuan Guo , Jie Zhou , Jiwen Lu

One central goal of robotics is to enable robots to interact with the physical world. Traditional manipulation studies primarily focus on single robots and relatively small objects. However, factory and domestic environments often require…

Robotics · Computer Science 2026-05-26 Kun Song , Gaoming Chen , Shentao Ma , Ninglong Jin , Guangbao Zhao , Mingyu Ding , Zhenhua Xiong , Jia Pan

Pose estimation is a basic module in many robot manipulation pipelines. Estimating the pose of objects in the environment can be useful for grasping, motion planning, or manipulation. However, current state-of-the-art methods for pose…

Computer Vision and Pattern Recognition · Computer Science 2021-05-03 Brian Okorn , Qiao Gu , Martial Hebert , David Held

We present a fully autonomous real-world RL framework for mobile manipulation that can learn policies without extensive instrumentation or human supervision. This is enabled by 1) task-relevant autonomy, which guides exploration towards…

Robotics · Computer Science 2024-10-01 Russell Mendonca , Emmanuel Panov , Bernadette Bucher , Jiuguang Wang , Deepak Pathak

Language-guided robotic grasping is a rapidly advancing field where robots are instructed using human language to grasp specific objects. However, existing methods often depend on dense camera views and struggle to quickly update scenes,…

Robotics · Computer Science 2024-12-04 Junqiu Yu , Xinlin Ren , Yongchong Gu , Haitao Lin , Tianyu Wang , Yi Zhu , Hang Xu , Yu-Gang Jiang , Xiangyang Xue , Yanwei Fu

Large language models (LLMs) are shown to possess a wealth of actionable knowledge that can be extracted for robot manipulation in the form of reasoning and planning. Despite the progress, most still rely on pre-defined motion primitives to…

Robotics · Computer Science 2023-11-03 Wenlong Huang , Chen Wang , Ruohan Zhang , Yunzhu Li , Jiajun Wu , Li Fei-Fei

We propose VISO-Grasp, a novel vision-language-informed system designed to systematically address visibility constraints for grasping in severely occluded environments. By leveraging Foundation Models (FMs) for spatial reasoning and active…

Robotics · Computer Science 2025-08-07 Yitian Shi , Di Wen , Guanqi Chen , Edgar Welte , Sheng Liu , Kunyu Peng , Rainer Stiefelhagen , Rania Rayyes

This paper considers the final approach phase of visual-closed-loop grasping where the RGB-D camera is no longer able to provide valid depth information. Many current robotic grasping controllers are not closed-loop and therefore fail for…

Robotics · Computer Science 2020-03-02 Jesse Haviland , Feras Dayoub , Peter Corke

Intelligent manipulation benefits from the capacity to flexibly control an end-effector with high degrees of freedom (DoF) and dynamically react to the environment. However, due to the challenges of collecting effective training data and…

Computer Vision and Pattern Recognition · Computer Science 2020-06-19 Shuran Song , Andy Zeng , Johnny Lee , Thomas Funkhouser

We introduce a novel system for human-to-robot trajectory transfer that enables robots to manipulate objects by learning from human demonstration videos. The system consists of four modules. The first module is a data collection module that…

Robotics · Computer Science 2025-10-27 Sai Haneesh Allu , Jishnu Jaykumar P , Ninad Khargonkar , Tyler Summers , Jian Yao , Yu Xiang

In multi-robot systems where a central decision maker is specifying the movement of each individual robot, a communication failure can severely impair the performance of the system. This paper develops a motion strategy that allows robots…

Multiagent Systems · Computer Science 2017-02-14 Siddharth Mayya , Magnus Egerstedt

Robotic grasping in cluttered environments remains a significant challenge due to occlusions and complex object arrangements. We have developed ThinkGrasp, a plug-and-play vision-language grasping system that makes use of GPT-4o's advanced…

Robotics · Computer Science 2026-04-03 Yaoyao Qian , Xupeng Zhu , Ondrej Biza , Shuo Jiang , Linfeng Zhao , Haojie Huang , Yu Qi , Robert Platt

This paper presents a grounded language-image pre-training (GLIP) model for learning object-level, language-aware, and semantic-rich visual representations. GLIP unifies object detection and phrase grounding for pre-training. The…

Computer Vision and Pattern Recognition · Computer Science 2022-06-20 Liunian Harold Li , Pengchuan Zhang , Haotian Zhang , Jianwei Yang , Chunyuan Li , Yiwu Zhong , Lijuan Wang , Lu Yuan , Lei Zhang , Jenq-Neng Hwang , Kai-Wei Chang , Jianfeng Gao

Self-supervised and language-supervised image models contain rich knowledge of the world that is important for generalization. Many robotic tasks, however, require a detailed understanding of 3D geometry, which is often lacking in 2D image…

Computer Vision and Pattern Recognition · Computer Science 2024-01-02 William Shen , Ge Yang , Alan Yu , Jansen Wong , Leslie Pack Kaelbling , Phillip Isola

The ability to grasp objects in-the-wild from open-ended language instructions constitutes a fundamental challenge in robotics. An open-world grasping system should be able to combine high-level contextual with low-level physical-geometric…

Robotics · Computer Science 2024-10-15 Georgios Tziafas , Hamidreza Kasaei

Task-oriented dexterous grasping holds broad application prospects in robotic manipulation and human-object interaction. However, most existing methods still struggle to generalize across diverse objects and task instructions, as they…

Robotics · Computer Science 2025-11-18 Juntao Jian , Yi-Lin Wei , Chengjie Mou , Yuhao Lin , Xing Zhu , Yujun Shen , Wei-Shi Zheng , Ruizhen Hu

Large-scale contrastive vision-language pre-trained models provide the zero-shot model achieving competitive performance across a range of image classification tasks without requiring training on downstream data. Recent works have confirmed…

Machine Learning · Computer Science 2024-04-02 Giung Nam , Byeongho Heo , Juho Lee