English
Related papers

Related papers: One Map to Find Them All: Real-time Open-Vocabular…

200 papers

Open vocabulary 3D object detection (OV3D) allows precise and extensible object recognition crucial for adapting to diverse environments encountered in assistive robotics. This paper presents OpenNav, a zero-shot 3D object detection…

Computer Vision and Pattern Recognition · Computer Science 2024-08-27 Muhammad Rameez ur Rahman , Piero Simonetto , Anna Polato , Francesco Pasti , Luca Tonin , Sebastiano Vascon

Object search is a fundamental task for robots deployed in indoor building environments, yet challenges arise due to observation instability, especially for open-vocabulary models. While foundation models (LLMs/VLMs) enable reasoning about…

Robotics · Computer Science 2025-03-05 Qianwei Wang , Yifan Xu , Vineet Kamat , Carol Menassa

We present LGX (Language-guided Exploration), a novel algorithm for Language-Driven Zero-Shot Object Goal Navigation (L-ZSON), where an embodied agent navigates to a uniquely described target object in a previously unseen environment. Our…

Robotics · Computer Science 2024-04-16 Vishnu Sashank Dorbala , James F. Mullen , Dinesh Manocha

Current vision-language navigation methods face substantial bottlenecks regarding heterogeneous robot compatibility, real-time performance, and navigation safety. Furthermore, they struggle to support open-vocabulary semantic generalization…

Robotics · Computer Science 2026-04-06 Mingao Tan , Yiyang Li , Shanze Wang , Xinming Zhang , Wei Zhang

The recent Segment Anything Model (SAM) has emerged as a new paradigmatic vision foundation model, showcasing potent zero-shot generalization and flexible prompting. Despite SAM finding applications and adaptations in various domains, its…

Computer Vision and Pattern Recognition · Computer Science 2025-02-14 Xumeng Han , Longhui Wei , Xuehui Yu , Zhiyang Dou , Xin He , Kuiran Wang , Yingfei Sun , Zhenjun Han , Qi Tian

Zero-shot object navigation (ZSON) in unseen environments remains a challenging problem for household robots, requiring strong perceptual understanding and decision-making capabilities. While recent methods leverage metric maps and Large…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Qunchao Jin , Yilin Wu , Changhao Chen

Zero-shot object navigation is a challenging task for home-assistance robots. This task emphasizes visual grounding, commonsense inference and locomotion abilities, where the first two are inherent in foundation models. But for the…

Robotics · Computer Science 2023-09-22 Wenzhe Cai , Siyuan Huang , Guangran Cheng , Yuxing Long , Peng Gao , Changyin Sun , Hao Dong

This paper addresses the problem of enabling a robot to search for a semantic object, i.e., an object with a semantic label, in an unknown and GPS-denied environment. For the robot in the unknown environment to detect and find the target…

Robotics · Computer Science 2023-11-22 Zhentian Qian , Jie Fu , Jing Xiao

Robotic search of people in human-centered environments, including healthcare settings, is challenging as autonomous robots need to locate people without complete or any prior knowledge of their schedules, plans or locations. Furthermore,…

Robotics · Computer Science 2024-12-03 Angus Fung , Aaron Hao Tan , Haitong Wang , Beno Benhabib , Goldie Nejat

Enabling robots to autonomously navigate unknown, complex, and dynamic real-world environments presents several challenges, including imperfect perception, partial observability, localization uncertainty, and safety constraints. Current…

Grounding language to the visual observations of a navigating agent can be performed using off-the-shelf visual-language models pretrained on Internet-scale data (e.g., image captions). While this is useful for matching images to natural…

Robotics · Computer Science 2023-03-09 Chenguang Huang , Oier Mees , Andy Zeng , Wolfram Burgard

Open-Vocabulary Object Navigation (OVON) requires an embodied agent to locate a language-specified target in unknown environments. Existing zero-shot methods often reason over dense frontier points under incomplete observations, causing…

Robotics · Computer Science 2026-03-31 Maoguo Gao , Zejun Zhu , Zhiming Sun , Zhengwei Ma , Longze Yuan , Zhongjing Ma , Zhigang Gao , Jinhui Zhang , Suli Zou

The ability to interpret and comprehend a 3D scene is essential for many vision and robotics systems. In numerous applications, this involves 3D object detection, i.e.~identifying the location and dimensions of objects belonging to a…

Computer Vision and Pattern Recognition · Computer Science 2025-09-22 Olivier Moliner , Viktor Larsson , Kalle Åström

Mapping and understanding complex 3D environments is fundamental to how autonomous systems perceive and interact with the physical world, requiring both precise geometric reconstruction and rich semantic comprehension. While existing 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-05-22 Naman Patel , Prashanth Krishnamurthy , Farshad Khorrami

Large language models (LLMs) have unlocked new capabilities of task planning from human instructions. However, prior attempts to apply LLMs to real-world robotic tasks are limited by the lack of grounding in the surrounding scene. In this…

Embodied navigation presents a core challenge for intelligent robots, requiring the comprehension of visual environments, natural language instructions, and autonomous exploration. Existing models often fall short in offering a unified…

Robotics · Computer Science 2026-01-08 Xinda Xue , Junjun Hu , Minghua Luo , Shichao Xie , Jintao Chen , Zixun Xie , Kuichen Quan , Wei Guo , Mu Xu , Zedong Chu

Modern intelligent and autonomous robotic applications often require robots to have more information about their environment than that provided by traditional occupancy grid maps. For example, a robot tasked to perform autonomous semantic…

Robotics · Computer Science 2021-03-30 Tobias Zaenker , Francesco Verdoja , Ville Kyrki

Navigating unknown environments to find a target object is a significant challenge. While semantic information is crucial for navigation, relying solely on it for decision-making may not always be efficient, especially in environments with…

Robotics · Computer Science 2025-09-08 Mingjie Zhang , Yuheng Du , Chengkai Wu , Jinni Zhou , Zhenchao Qi , Jun Ma , Boyu Zhou

Recent advances in vision-based navigation and exploration have shown impressive capabilities in photorealistic indoor environments. However, these methods still struggle with long-horizon tasks and require large amounts of data to…

Robotics · Computer Science 2022-08-25 Fabian Schmalstieg , Daniel Honerkamp , Tim Welschehold , Abhinav Valada

Recent advances in vision-language models have made zero-shot navigation feasible, enabling robots to follow natural language instructions without requiring labeling. However, existing methods that explicitly store language vectors in grid…

Robotics · Computer Science 2026-02-13 Sibaek Lee , Hyeonwoo Yu , Giseop Kim , Sunwook Choi