English
Related papers

Related papers: Balancing Performance and Efficiency in Zero-shot …

200 papers

Object navigation is a core capability of embodied intelligence, enabling an agent to locate target objects in unknown environments. Recent advances in vision-language models (VLMs) have facilitated zero-shot object navigation (ZSON).…

Robotics · Computer Science 2026-02-13 Wancai Zheng , Hao Chen , Xianlong Lu , Linlin Ou , Xinyi Yu

In this paper, we present a novel method for reliable frontier selection in Zero-Shot Object Goal Navigation (ZS-OGN), enhancing robotic navigation systems with foundation models to improve commonsense reasoning in indoor environments. Our…

Robotics · Computer Science 2024-10-29 Shuaihang Yuan , Halil Utku Unlu , Hao Huang , Congcong Wen , Anthony Tzes , Yi Fang

While Vision-Language Models (VLMs) enable high-level semantic reasoning for end-to-end autonomous driving, particularly in unstructured environments, existing off-road datasets suffer from language annotations that are weakly aligned with…

Robotics · Computer Science 2026-04-24 Byounggun Park , Soonmin Hwang

Vision-Language-Action models (VLAs) represent a significant frontier in embodied intelligence, aiming to bridge digital knowledge with physical-world interaction. Despite their remarkable performance, foundational VLAs are hindered by the…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Zhaoshu Yu , Bo Wang , Pengpeng Zeng , Haonan Zhang , Ji Zhang , Zheng Wang , Lianli Gao , Jingkuan Song , Nicu Sebe , Heng Tao Shen

Existing Vision-Language Navigation (VLN) agents based on Large Vision-Language Models (LVLMs) often suffer from perception errors, reasoning errors, and planning errors, which significantly hinder their navigation performance. To address…

Machine Learning · Computer Science 2025-12-03 Zhengcheng Wang , Zichuan Lin , Yijun Yang , Haobo Fu , Deheng Ye

Task-aware navigation continues to be a challenging area of research, especially in scenarios involving open vocabulary. Previous studies primarily focus on finding suitable locations for task completion, often overlooking the importance of…

Robotics · Computer Science 2024-09-18 Jun Zhu , Zihao Du , Haotian Xu , Fengbo Lan , Zilong Zheng , Bo Ma , Shengjie Wang , Tao Zhang

Grounding natural language instructions to visual observations is fundamental for embodied agents operating in open-world environments. Recent advances in visual-language mapping have enabled generalizable semantic representations by…

Robotics · Computer Science 2025-08-05 Danyang Li , Zenghui Yang , Guangpeng Qi , Songtao Pang , Guangyong Shang , Qiang Ma , Zheng Yang

Integration of Large Language Models (LLMs) into visual domain tasks, resulting in visual-LLMs (V-LLMs), has enabled exceptional performance in vision-language tasks, particularly for visual question answering (VQA). However, existing…

Computer Vision and Pattern Recognition · Computer Science 2024-04-12 Kanchana Ranasinghe , Satya Narayan Shukla , Omid Poursaeed , Michael S. Ryoo , Tsung-Yu Lin

Model-based control is a popular paradigm for robot navigation because it can leverage a known dynamics model to efficiently plan robust robot trajectories. However, it is challenging to use model-based methods in settings where the…

Robotics · Computer Science 2019-07-19 Somil Bansal , Varun Tolani , Saurabh Gupta , Jitendra Malik , Claire Tomlin

In the realm of household robotics, the Zero-Shot Object Navigation (ZSON) task empowers agents to adeptly traverse unfamiliar environments and locate objects from novel categories without prior explicit training. This paper introduces…

Robotics · Computer Science 2024-02-07 Pengying Wu , Yao Mu , Bingxian Wu , Yi Hou , Ji Ma , Shanghang Zhang , Chang Liu

Vision language models (VLMs) have shown impressive capabilities across a variety of tasks, from logical reasoning to visual understanding. This opens the door to richer interaction with the world, for example robotic control. However, VLMs…

Video generative models (VGMs) pretrained on large-scale internet data can produce temporally coherent rollout videos that capture rich object dynamics, offering a compelling foundation for zero-shot robotic manipulation. However, VGMs…

Robotics · Computer Science 2026-03-09 Gehao Zhang , Zhenyang Ni , Payal Mohapatra , Han Liu , Ruohan Zhang , Qi Zhu

In this paper, we propose a training-free framework for vision-and-language navigation (VLN). Existing zero-shot VLN methods are mainly designed for discrete environments or involve unsupervised training in continuous simulator…

Robotics · Computer Science 2025-09-15 Hang Yin , Haoyu Wei , Xiuwei Xu , Wenxuan Guo , Jie Zhou , Jiwen Lu

Vision-Language Models (VLMs) offer promising capabilities for mobile devices, but their deployment faces significant challenges due to computational limitations and energy inefficiency, especially for real-time applications. This study…

Machine Learning · Computer Science 2025-07-15 Pablo Robin Guerrero , Yueyang Pan , Sanidhya Kashyap

Mobile robots operating in human-centered environments must generate not only collision-free paths but also trajectories that follow local behavioral conventions. Conventional costmap-based navigation emphasizes geometric feasibility and…

Robotics · Computer Science 2026-05-19 Dongjie Huo , Junhui Wang , Chao Gao , Yan Qiao , Dong Zhang , Guyue Zhou

Estimating the 3D world from 2D monocular images is a fundamental yet challenging task due to the labour-intensive nature of 3D annotations. To simplify label acquisition, this work proposes a novel approach that bridges 2D vision…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Sihao Lin , Daqi Liu , Ruochong Fu , Dongrui Liu , Andy Song , Hongwei Xie , Zhihui Li , Bing Wang , Xiaojun Chang

While Vision-Language Models (VLMs) show significant promise for end-to-end autonomous driving by leveraging the common sense embedded in language models, their reliance on 2D image cues for complex scene understanding and decision-making…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Weijie Wei , Zhipeng Luo , Ling Feng , Venice Erin Liong

In this paper, we introduce Flash-VL 2B, a novel approach to optimizing Vision-Language Models (VLMs) for real-time applications, targeting ultra-low latency and high throughput without sacrificing accuracy. Leveraging advanced…

Computer Vision and Pattern Recognition · Computer Science 2025-05-15 Bo Zhang , Shuo Li , Runhe Tian , Yang Yang , Jixin Tang , Jinhao Zhou , Lin Ma

Vision Foundation Models (VFMs) have become a de facto choice for many downstream vision tasks, like image classification, image segmentation, and object localization. However, they can also provide significant utility for downstream 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-04-22 Johannes Spoecklberger , Wei Lin , Pedro Hermosilla , Sivan Doveh , Horst Possegger , M. Jehanzeb Mirza

Effective robot navigation in unseen environments is a challenging task that requires precise control actions at high frequencies. Recent advances have framed it as an image-goal-conditioned control problem, where the robot generates…