中文
相关论文

相关论文: NavigScene: Bridging Local Perception and Global N…

200 篇论文

Dataflow visualization systems enable flexible visual data exploration by allowing the user to construct a dataflow diagram that composes query and visualization modules to specify system functionality. However learning dataflow diagram…

人机交互 · 计算机科学 2019-10-08 Bowen Yu , Claudio T. Silva

Visual navigation tasks in real-world environments often require both self-motion and place recognition feedback. While deep reinforcement learning has shown success in solving these perception and decision-making problems in an end-to-end…

机器人学 · 计算机科学 2020-03-03 Marvin Chancán , Michael Milford

Autonomous driving requires accurate scene understanding, including road geometry, traffic agents, and their semantic relationships. In online HD map generation scenarios, raster-based representations are well-suited to vision models but…

计算机视觉与模式识别 · 计算机科学 2025-08-05 Zhigang Sun , Yiru Wang , Anqing Jiang , Shuo Wang , Yu Gao , Yuwen Heng , Shouyi Zhang , An He , Hao Jiang , Jinhao Chai , Zichong Gu , Wang Jijun , Shichen Tang , Lavdim Halilaj , Juergen Luettin , Hao Sun

The pursuit of autonomous driving technology hinges on the sophisticated integration of perception, decision-making, and control systems. Traditional approaches, both data-driven and rule-based, have been hindered by their inability to…

Autonomous driving systems depend on on models that can reason about high-level scene contexts and accurately predict the dynamics of their surrounding environment. Vision- Language Models (VLMs) have recently emerged as promising tools for…

计算机视觉与模式识别 · 计算机科学 2026-03-18 Stefan Englmeier , Katharina Winter , Fabian B. Flohr

Current end-to-end autonomous driving planners are fundamentally reactive: they condition on historical and present observations to predict future actions. We argue that autonomous agents should instead imagine future scenes before…

计算机视觉与模式识别 · 计算机科学 2026-05-11 Bozhou Zhang , Nan Song , Yuang Wang , Jiankang Deng , Xiatian Zhu , Li Zhang

Understanding and following natural language instructions while navigating through complex, real-world environments poses a significant challenge for general-purpose robots. These environments often include obstacles and pedestrians, making…

计算机视觉与模式识别 · 计算机科学 2024-03-15 Xiwen Liang , Liang Ma , Shanshan Guo , Jianhua Han , Hang Xu , Shikui Ma , Xiaodan Liang

Visual perception is an important component for autonomous navigation of unmanned surface vessels (USV), particularly for the tasks related to autonomous inspection and tracking. These tasks involve vision-based navigation techniques to…

计算机视觉与模式识别 · 计算机科学 2023-08-09 Muhayyuddin Ahmed , Ahsan Baidar Bakht , Taimur Hassan , Waseem Akram , Ahmed Humais , Lakmal Seneviratne , Shaoming He , Defu Lin , Irfan Hussain

In the field of autonomous driving, two important features of autonomous driving car systems are the explainability of decision logic and the accuracy of environmental perception. This paper introduces DME-Driver, a new autonomous driving…

机器人学 · 计算机科学 2024-01-09 Wencheng Han , Dongqian Guo , Cheng-Zhong Xu , Jianbing Shen

In recent years, predicting driver's focus of attention has been a very active area of research in the autonomous driving community. Unfortunately, existing state-of-the-art techniques achieve this by relying only on human gaze information,…

计算机视觉与模式识别 · 计算机科学 2020-04-02 Anwesan Pal , Sayan Mondal , Henrik I. Christensen

Goal-oriented navigation presents a fundamental challenge for autonomous systems, requiring agents to navigate complex environments to reach designated targets. This survey offers a comprehensive analysis of multimodal navigation approaches…

机器人学 · 计算机科学 2025-04-23 I-Tak Ieong , Hao Tang

We deal with the navigation problem where the agent follows natural language instructions while observing the environment. Focusing on language understanding, we show the importance of spatial semantics in grounding navigation instructions…

计算与语言 · 计算机科学 2021-05-17 Yue Zhang , Quan Guo , Parisa Kordjamshidi

Personalized driving refers to an autonomous vehicle's ability to adapt its driving behavior or control strategies to match individual users' preferences and driving styles while maintaining safety and comfort standards. However, existing…

Visual navigation is a fundamental capability for autonomous home-assistance robots, enabling long-horizon tasks such as object search. While recent methods have leveraged Large Language Models (LLMs) to incorporate commonsense reasoning…

机器人学 · 计算机科学 2026-05-01 Teng Wang , Xinxin Zhao , Wenzhe Cai , Changyin Sun

Vision-language models enable the understanding and reasoning of complex traffic scenarios through multi-source information fusion, establishing it as a core technology for autonomous driving. However, existing vision-language models are…

计算机视觉与模式识别 · 计算机科学 2025-12-17 Minghui Hou , Wei-Hsing Huang , Shaofeng Liang , Daizong Liu , Tai-Hao Wen , Gang Wang , Runwei Guan , Weiping Ding

We investigate the Vision-and-Language Navigation (VLN) problem in the context of autonomous driving in outdoor settings. We solve the problem by explicitly grounding the navigable regions corresponding to the textual command. At each…

计算机视觉与模式识别 · 计算机科学 2022-09-27 Kanishk Jain , Varun Chhangani , Amogh Tiwari , K. Madhava Krishna , Vineet Gandhi

Learning to navigate in a visual environment following natural-language instructions is a challenging task, because the multimodal inputs to the agent are highly variable, and the training data on a new task is often limited. In this paper,…

计算机视觉与模式识别 · 计算机科学 2020-04-07 Weituo Hao , Chunyuan Li , Xiujun Li , Lawrence Carin , Jianfeng Gao

Reliable autonomous driving requires scene understanding that is semantically consistent across heterogeneous sensors and verifiable at the reasoning stage. However, many recent LLM-driven driving systems attach the language model as a…

计算机视觉与模式识别 · 计算机科学 2026-05-07 Shuo Liu , Lei Shi , Haowen Liu , Jing Xu , Yufei Gao , Yucheng Shi

Language-guided navigation is a cornerstone of embodied AI, enabling agents to interpret language instructions and navigate complex environments. However, expert-provided instructions are limited in quantity, while synthesized annotations…

人工智能 · 计算机科学 2025-08-11 Zongtao He , Liuyi Wang , Lu Chen , Chengju Liu , Qijun Chen

Vision Language Action (VLA) models promise an open-vocabulary interface that can translate perceptual ambiguity into semantically grounded driving decisions, yet they still treat language as a static prior fixed at inference time. As a…

音频与语音处理 · 电气工程与系统科学 2026-01-30 Ziang Guo , Feng Yang , Xuefeng Zhang , Jiaqi Guo , Kun Zhao , Yixiao Zhou , Peng Lu , Sifa Zheng , Zufeng Zhang