中文
相关论文

相关论文: Balancing Performance and Efficiency in Zero-shot …

200 篇论文

Open-world navigation requires robots to make decisions in complex everyday environments while adapting to flexible task requirements. Conventional navigation approaches often rely on dense 3D reconstruction and hand-crafted goal metrics,…

机器人学 · 计算机科学 2026-05-18 Esteban Padilla-Cerdio , Boyang Sun , Marc Pollefeys , Hermann Blum

Vision language models (VLMs) can simultaneously reason about images and texts to tackle many tasks, from visual question answering to image captioning. This paper focuses on map parsing, a novel task that is unexplored within the VLM…

机器人学 · 计算机科学 2025-11-26 David DeFazio , Hrudayangam Mehta , Meng Wang , Ping Yang , Jeremy Blackburn , Shiqi Zhang

The zero-shot object navigation (ZSON) in unknown open-ended environments coupled with semantically novel target often suffers from the significant decline in performance due to the neglect of high-dimensional implicit scene information and…

机器人学 · 计算机科学 2025-06-09 Chongshang Yan , Jiaxuan He , Delun Li , Yi Yang , Wenjie Song

Vision-Language Models (VLMs) have emerged as the dominant approach for zero-shot recognition, adept at handling diverse scenarios and significant distribution changes. However, their deployment in risk-sensitive areas requires a deeper…

计算机视觉与模式识别 · 计算机科学 2024-06-17 Weijie Tu , Weijian Deng , Dylan Campbell , Stephen Gould , Tom Gedeon

This paper investigates the potential of vision-language models (VLMs) to assist people with blindness and low vision (pBLV) in navigation tasks. We evaluate state-of-the-art closed-source models, including GPT-4V, GPT-4o, Gemini-1.5-Pro,…

计算机视觉与模式识别 · 计算机科学 2026-03-18 Yu Li , Yuchen Zheng , Giles Hamilton-Fletcher , Marco Mezzavilla , Yao Wang , Sundeep Rangan , Maurizio Porfiri , Zhou Yu , John-Ross Rizzo

Vision-Language Foundation Models (VLFMs) have made remarkable progress on various multimodal tasks, such as image captioning, image-text retrieval, visual question answering, and visual grounding. However, most methods rely on training…

计算机视觉与模式识别 · 计算机科学 2026-01-06 Yue Zhou , Zhihang Zhong , Xue Yang

This paper presents novel benchmarks for evaluating vision-language models (VLMs) in zero-shot recognition, focusing on granularity and specificity. Although VLMs excel in tasks like image captioning, they face challenges in open-world…

计算机视觉与模式识别 · 计算机科学 2024-06-19 Zhenlin Xu , Yi Zhu , Tiffany Deng , Abhay Mittal , Yanbei Chen , Manchen Wang , Paolo Favaro , Joseph Tighe , Davide Modolo

We present VLMnav, an embodied framework to transform a Vision-Language Model (VLM) into an end-to-end navigation policy. In contrast to prior work, we do not rely on a separation between perception, planning, and control; instead, we use a…

机器人学 · 计算机科学 2024-11-11 Dylan Goetting , Himanshu Gaurav Singh , Antonio Loquercio

The real world is messy and unstructured. Uncovering critical information often requires active, goal-driven exploration. It remains to be seen whether Vision-Language Models (VLMs), which recently emerged as a popular zero-shot tool in…

计算机视觉与模式识别 · 计算机科学 2025-10-22 Adam Pardyl , Dominik Matuszek , Mateusz Przebieracz , Marek Cygan , Bartosz Zieliński , Maciej Wołczyk

Autonomous navigation under natural language instructions represents a crucial step toward embodied intelligence, enabling complex task execution in environments ranging from industrial facilities to domestic spaces. However,…

机器人学 · 计算机科学 2026-03-05 Hongyu Song , Rishabh Dev Yadav , Cheng Guo , Wei Pan

While interacting in the world is a multi-sensory experience, many robots continue to predominantly rely on visual perception to map and navigate in their environments. In this work, we propose Audio-Visual-Language Maps (AVLMaps), a…

机器人学 · 计算机科学 2023-03-28 Chenguang Huang , Oier Mees , Andy Zeng , Wolfram Burgard

Commanding a robot to navigate with natural language instructions is a long-term goal for grounded language understanding and robotics. But the dominant language is English, according to previous studies on vision-language navigation (VLN).…

计算与语言 · 计算机科学 2020-12-08 An Yan , Xin Eric Wang , Jiangtao Feng , Lei Li , William Yang Wang

Analysis of state-of-the-art VO/VSLAM system exposes a gap in balancing performance (accuracy & robustness) and efficiency (latency). Feature-based systems exhibit good performance, yet have higher latency due to explicit data association;…

机器人学 · 计算机科学 2020-01-06 Yipu Zhao , Patricio A. Vela

This paper presents VisLingInstruct, a novel approach to advancing Multi-Modal Language Models (MMLMs) in zero-shot learning. Current MMLMs show impressive zero-shot abilities in multi-modal tasks, but their performance depends heavily on…

人工智能 · 计算机科学 2024-06-21 Dongsheng Zhu , Xunzhu Tang , Weidong Han , Jinghui Lu , Yukun Zhao , Guoliang Xing , Junfeng Wang , Dawei Yin

In open-vocabulary mobile manipulation (OVMM), task success often hinges on the selection of an appropriate base placement for the robot. Existing approaches typically navigate to proximity-based regions without considering affordances,…

机器人学 · 计算机科学 2026-01-06 Tzu-Jung Lin , Jia-Fong Yeh , Hung-Ting Su , Chung-Yi Lin , Yi-Ting Chen , Winston H. Hsu

Recent open-vocabulary robot mapping methods enrich dense geometric maps with pre-trained visual-language features, achieving a high level of detail and guiding robots to find objects specified by open-vocabulary language queries. While the…

机器人学 · 计算机科学 2026-03-04 Fujing Xie , Sören Schwertfeger , Hermann Blum

Recent advances in Vision-and-Language Navigation in Continuous Environments (VLN-CE) have leveraged multimodal large language models (MLLMs) to achieve zero-shot navigation. However, existing methods often rely on panoramic observations…

机器人学 · 计算机科学 2025-11-04 Xiangyu Shi , Zerui Li , Yanyuan Qiao , Qi Wu

Despite recent advances demonstrating vision-language models' (VLMs) abilities to describe complex relationships in images using natural language, their capability to quantitatively reason about object sizes and distances remains…

计算机视觉与模式识别 · 计算机科学 2024-09-17 Yuan-Hong Liao , Rafid Mahmood , Sanja Fidler , David Acuna

While specialized learning-based models have historically dominated image privacy prediction, the current literature increasingly favours adopting large Vision-Language Models (VLMs) designed for generic tasks. This trend risks overlooking…

计算机视觉与模式识别 · 计算机科学 2025-10-13 Alina Elena Baia , Alessio Xompero , Andrea Cavallaro

Remote sensing has become a vital tool across sectors such as urban planning, environmental monitoring, and disaster response. While the volume of data generated has increased significantly, traditional vision models are often constrained…

计算机视觉与模式识别 · 计算机科学 2025-10-17 Jia Yun Chua , Argyrios Zolotas , Miguel Arana-Catania