English
Related papers

Related papers: Exploring the Reliability of Foundation Model-Base…

200 papers

This paper investigates the automatic exploration problem under the unknown environment, which is the key point of applying the robotic system to some social tasks. The solution to this problem via stacking decision rules is impossible to…

Robotics · Computer Science 2020-07-24 Haoran Li , Qichao Zhang , Dongbin Zhao

Foundation Models (FMs) are increasingly integrated into remote sensing (RS) pipelines. These models include unimodal vision encoders and multimodal architectures. FMs are adapted to diverse perception tasks, such as image classification,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-12 Binger Chen , Tacettin Emre Bök , Behnood Rasti , Volker Markl , Begüm Demir

Zero-shot stance detection (ZSSD) seeks to determine the stance of text toward previously unseen targets, a task critical for analyzing dynamic and polarized online discourse with limited labeled data. While large language models (LLMs)…

Computation and Language · Computer Science 2026-01-27 Bowen Zhang , Jun Ma , Fuqiang Niu , Li Dong , Jinzhou Cao , Genan Dai

LLM-based agents have demonstrated impressive zero-shot performance in vision-language navigation (VLN) task. However, existing LLM-based methods often focus only on solving high-level task planning by selecting nodes in predefined…

Robotics · Computer Science 2024-08-21 Jiaqi Chen , Bingqian Lin , Xinmin Liu , Lin Ma , Xiaodan Liang , Kwan-Yee K. Wong

Efficient target localization and autonomous navigation in complex environments are fundamental to real-world embodied applications. While recent advances in multimodal foundation models have enabled zero-shot object goal navigation,…

Robotics · Computer Science 2026-04-02 Ming-Ming Yu , Yi Chen , Börje F. Karlsson , Wenjun Wu

We propose a modular framework that leverages the expertise of different foundation models over different modalities and domains in order to perform a single, complex, multi-modal task, without relying on prompt engineering or otherwise…

Computation and Language · Computer Science 2023-10-31 Daniela Ben-David , Tzuf Paz-Argaman , Reut Tsarfaty

Navigating toward specific objects in unknown environments without additional training, known as Zero-Shot object navigation, poses a significant challenge in the field of robotics, which demands high levels of auxiliary information and…

Robotics · Computer Science 2024-03-25 Lingfeng Zhang , Qiang Zhang , Hao Wang , Erjia Xiao , Zixuan Jiang , Honglei Chen , Renjing Xu

Vision-and-Language Navigation (VLN) in continuous environments requires agents to interpret natural language instructions while navigating unconstrained 3D spaces. Existing VLN-CE frameworks rely on a two-stage approach: a waypoint…

Robotics · Computer Science 2025-06-18 Xiangyu Shi , Zerui Li , Wenqi Lyu , Jiatong Xia , Feras Dayoub , Yanyuan Qiao , Qi Wu

Depth from Defocus (DfD) is the task of estimating a dense metric depth map from a focus stack. Unlike previous works overfitting to a certain dataset, this paper focuses on the challenging and practical setting of zero-shot generalization.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Yiming Zuo , Hongyu Wen , Venkat Subramanian , Patrick Chen , Karhan Kayan , Mario Bijelic , Felix Heide , Jia Deng

As autonomous robots move into complex, dynamic real-world environments, they must learn to navigate safely in real time, yet anticipating all possible behaviors is infeasible. We propose a composable, model-free reinforcement learning…

Robotics · Computer Science 2026-02-16 Xinhuan Sang , Abdelrahman Abdelgawad , Roberto Tron

Mobile robots exploring indoor environments increasingly rely on vision-language models to perceive high-level semantic cues in camera images, such as object categories. Such models offer the potential to substantially advance robot…

Robotics · Computer Science 2025-10-09 Utkarsh Bajpai , Julius Rückin , Cyrill Stachniss , Marija Popović

Recent open-vocabulary robot mapping methods enrich dense geometric maps with pre-trained visual-language features, achieving a high level of detail and guiding robots to find objects specified by open-vocabulary language queries. While the…

Robotics · Computer Science 2026-03-04 Fujing Xie , Sören Schwertfeger , Hermann Blum

Robust local navigation in unstructured and dynamic environments remains a significant challenge for humanoid robots, requiring a delicate balance between long-range navigation targets and immediate motion stability. In this paper, we…

Robotics · Computer Science 2026-01-21 Yang Zhang , Jianming Ma , Liyun Yan , Zhanxiang Cao , Yazhou Zhang , Haoyang Li , Yue Gao

Navigation Foundation Models (NFMs) trained on large cross-embodied datasets have demonstrated powerful generalizability in various scenarios. Adopting in-domain fine-tuning for an NFM efficiently calibrates the visuomotor policy, promising…

Robotics · Computer Science 2026-05-20 Shintaro Nakaoka , Takayuki Kanai , Kazuhito Tanaka

Autonomous ground robots operating in large-scale outdoor environments require both robust long-range navigation and fine-grained ''last-mile'' exploration. Current advances in visual-language navigation (VLN) work well at short-range…

Robotics · Computer Science 2026-05-26 Dongzhihan Wang , Yi Du , Jianan Sun , Yuan Xue , Yingchen Zhang , Bing Xiao , Chen Wang , Liang Xu

Development of navigation algorithms is essential for the successful deployment of robots in rapidly changing hazardous environments for which prior knowledge of configuration is often limited or unavailable. Use of traditional…

Robotics · Computer Science 2022-11-11 Paul Blum , Peter Crowley , George Lykotrafitis

Recent advancements in Generative AI, particularly in Large Language Models (LLMs) and Large Vision-Language Models (LVLMs), offer new possibilities for integrating cognitive planning into robotic systems. In this work, we present a novel…

Robotics · Computer Science 2024-11-06 Arjun P S , Andrew Melnik , Gora Chand Nandi

Semantic segmentation is critical to image content understanding and object localization. Recent development in fully-convolutional neural network (FCN) has enabled accurate pixel-level labeling. One issue in previous works is that the FCN…

Computer Vision and Pattern Recognition · Computer Science 2016-07-08 Qin Huang , Chunyang Xia , Wenchao Zheng , Yuhang Song , Hao Xu , C. -C. Jay Kuo

Training-free Vision-Language Navigation (VLN) agents powered by foundation models can follow instructions and explore 3D environments. However, existing approaches rely on greedy frontier selection and passive spatial memory, leading to…

Robotics · Computer Science 2026-04-03 Xueying Li , Feng Lyu , Hao Wu , Mingliu Liu , Jia-Nan Liu , Guozi Liu

Object-Goal Navigation (ObjectNav) requires an agent to autonomously explore an unknown environment and navigate toward target objects specified by a semantic label. While prior work has primarily studied zero-shot ObjectNav under 2D…

Robotics · Computer Science 2026-01-23 Zichen Yan , Yuchen Hou , Shenao Wang , Yichao Gao , Rui Huang , Lin Zhao