English
Related papers

Related papers: Closed-Loop Open-Vocabulary Mobile Manipulation wi…

200 papers

The theoretical ability of modular robots to reconfigure in response to complex tasks in a priori unknown environments has frequently been cited as an advantage and remains a major motivator for work in the field. We present a modular robot…

Robotics · Computer Science 2018-12-14 Jonathan Daudelin , Gangyuan Jing , Tarik Tosun , Mark Yim , Hadas Kress-Gazit , Mark Campbell

This paper investigates the possibility of intuitive human-robot interaction through the application of Natural Language Processing (NLP) and Large Language Models (LLMs) in mobile robotics. This work aims to explore the feasibility of…

Due to the powerful vision-language reasoning and generalization abilities, multimodal large language models (MLLMs) have garnered significant attention in the field of end-to-end (E2E) autonomous driving. However, their application to…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Xueyi Liu , Zuodong Zhong , Yuxin Guo , Yun-Fu Liu , Zhiguo Su , Qichao Zhang , Junli Wang , Yinfeng Gao , Yupeng Zheng , Qiao Lin , Huiyong Chen , Dongbin Zhao

Classical robot navigation often relies on hardcoded state machines and purely geometric path planners, limiting a robot's ability to interpret high-level semantic instructions. In this paper, we first assess GPT-4's ability to act as a…

Robotics · Computer Science 2025-05-06 Jesse Barkley , Abraham George , Amir Barati Farimani

There is invariably a trade-off between safety and efficiency for collaborative robots (cobots) in human-robot collaborations. Robots that interact minimally with humans can work with high speed and accuracy but cannot adapt to new tasks or…

Robotics · Computer Science 2022-10-13 Xiangjie Yan , Yongpeng Jiang , Chen Chen , Leiliang Gong , Ming Ge , Tao Zhang , Xiang Li

Autonomous robotic exploration of unknown and hazardous environments, a long-standing challenge, can be significantly improved by leveraging the advanced reasoning of Vision-Language Models (VLMs). We introduce a novel exploration pipeline…

Robotics · Computer Science 2026-05-25 Aarush Aitha , Avideh Zakhor

Open-vocabulary mobile manipulation (OVMM) that involves the handling of novel and unseen objects across different workspaces remains a significant challenge for real-world robotic applications. In this paper, we propose a novel…

Robotics · Computer Science 2025-07-24 Shen Tan , Dong Zhou , Xiangyu Shao , Junqiao Wang , Guanghui Sun

Traditional rule-based conversational robots, constrained by predefined scripts and static response mappings, fundamentally lack adaptability for personalized, long-term human interaction. While Large Language Models (LLMs) like GPT-4 have…

Human-Computer Interaction · Computer Science 2025-03-24 Zhijin Meng , Mohammed Althubyani , Shengyuan Xie , Imran Razzak , Eduardo B. Sandoval , Mahdi Bamdad , Francisco Cruz

The deployment of autonomous robots in various domains has raised significant concerns about their trustworthiness and accountability. This study explores the potential of Large Language Models (LLMs) in analyzing ROS 2 logs generated by…

We present a vision and language model named MultiModal-GPT to conduct multi-round dialogue with humans. MultiModal-GPT can follow various instructions from humans, such as generating a detailed caption, counting the number of interested…

Computer Vision and Pattern Recognition · Computer Science 2023-06-14 Tao Gong , Chengqi Lyu , Shilong Zhang , Yudong Wang , Miao Zheng , Qian Zhao , Kuikun Liu , Wenwei Zhang , Ping Luo , Kai Chen

Human robot interaction is an exciting task, which aimed to guide robots following instructions from human. Since huge gap lies between human natural language and machine codes, end to end human robot interaction models is fair challenging.…

Robotics · Computer Science 2023-08-25 Zichao Dong , Weikun Zhang , Xufeng Huang , Hang Ji , Xin Zhan , Junbo Chen

Object navigation in open-world environments remains a formidable and pervasive challenge for robotic systems, particularly when it comes to executing long-horizon tasks that require both open-world object detection and high-level task…

Robotics · Computer Science 2025-07-10 Daojie Peng , Jiahang Cao , Qiang Zhang , Jun Ma

Multimodal Large Language Models (MLLMs) demonstrate remarkable image-language capabilities, but their widespread use faces challenges in cost-effective training and adaptation. Existing approaches often necessitate expensive language model…

Computer Vision and Pattern Recognition · Computer Science 2024-08-14 Sayna Ebrahimi , Sercan O. Arik , Tejas Nama , Tomas Pfister

Automatic detection and prevention of open-set failures are crucial in closed-loop robotic systems. Recent studies often struggle to simultaneously identify unexpected failures reactively after they occur and prevent foreseeable ones…

Robotics · Computer Science 2025-03-24 Enshen Zhou , Qi Su , Cheng Chi , Zhizheng Zhang , Zhongyuan Wang , Tiejun Huang , Lu Sheng , He Wang

For robots to seamlessly interact with humans, we first need to make sure that humans and robots understand one another. Diverse algorithms have been developed to enable robots to learn from humans (i.e., transferring information from…

Robotics · Computer Science 2023-12-05 Soheil Habibian , Antonio Alvarez Valdivia , Laura H. Blumenschein , Dylan P. Losey

This paper presents a novel approach to enhance autonomous robotic manipulation using the Large Language Model (LLM) for logical inference, converting high-level language commands into sequences of executable motion functions. The proposed…

Robotics · Computer Science 2023-08-30 Haokun Liu , Yaonan Zhu , Kenji Kato , Izumi Kondo , Tadayoshi Aoyama , Yasuhisa Hasegawa

Recent developments in multimodal methodologies have marked the beginning of an exciting era for models adept at processing diverse data types, encompassing text, audio, and visual content. Models like GPT-4V, which merge computer vision…

Computation and Language · Computer Science 2024-11-15 Xiang Zhang , Senyu Li , Ning Shi , Bradley Hauer , Zijun Wu , Grzegorz Kondrak , Muhammad Abdul-Mageed , Laks V. S. Lakshmanan

Vision-Language Models (VLMs) exhibit strong visual reasoning capabilities, yet they still struggle with 3D understanding. In particular, VLMs often fail to infer a text-consistent goal 6D pose of a target object in a 3D scene. However, we…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Sangwon Baik , Gunhee Kim , Mingi Choi , Hanbyul Joo

Service and assistive robots are increasingly being deployed in dynamic social environments; however, ensuring transparent and explainable interactions remains a significant challenge. This paper presents a multimodal explainability module…

Robotics · Computer Science 2026-04-09 Oluwadamilola Sotomi , Devika Kodi , Aliasghar Arab

Recent advances in open-vocabulary mobile manipulation have brought robots into real domestic environments. In such settings, reliable long-horizon execution under open-set object references and frequent disturbances becomes essential.…

Robotics · Computer Science 2026-04-29 Jinhao Jiang , Shengyu Fang , Sibo Zuo , Yujie Tang , Yirui Li
‹ Prev 1 4 5 6 7 8 10 Next ›