English
Related papers

Related papers: RoadSceneVQA: Benchmarking Visual Question Answeri…

200 papers

The rise of Visual-Language Models (LVLMs) has unlocked new possibilities for seamlessly integrating visual and textual information. However, their ability to interpret cartographic maps remains largely unexplored. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2025-12-04 Huy Quang Ung , Guillaume Habault , Yasutaka Nishimura , Hao Niu , Roberto Legaspi , Tomoki Oya , Ryoichi Kojima , Masato Taya , Chihiro Ono , Atsunori Minamikawa , Yan Liu

If a Large Language Model (LLM) were to take a driving knowledge test today, would it pass? Beyond standard spatial and visual question-answering (QA) tasks on current autonomous driving benchmarks, driving knowledge tests require a…

Computer Vision and Pattern Recognition · Computer Science 2025-09-01 Maolin Wei , Wanzhou Liu , Eshed Ohn-Bar

Understanding novel situations in the traffic domain requires an intricate combination of domain-specific and causal commonsense knowledge. Prior work has provided sufficient perception-based modalities for traffic monitoring, in this…

Computation and Language · Computer Science 2022-12-16 Jiarui Zhang , Filip Ilievski , Aravinda Kollaa , Jonathan Francis , Kaixin Ma , Alessandro Oltramari

While autonomous navigation has achieved remarkable success in passive perception (e.g., object detection and segmentation), it remains fundamentally constrained by a void in knowledge-driven, interactive environmental cognition. In the…

Computer Vision and Pattern Recognition · Computer Science 2026-02-27 Runwei Guan , Shaofeng Liang , Ningwei Ouyang , Weichen Fei , Shanliang Yao , Wei Dai , Chenhao Ge , Penglei Sun , Xiaohui Zhu , Tao Huang , Ryan Wen Liu , Hui Xiong

While recent work has extended CoT to multimodal settings, achieving state-of-the-art results on science question answering benchmarks like ScienceQA, the generalizability of these approaches across diverse domains remains underexplored.…

Artificial Intelligence · Computer Science 2025-11-27 Nitya Tiwari , Parv Maheshwari , Vidisha Agarwal

Current autonomous driving vehicles rely mainly on their individual sensors to understand surrounding scenes and plan for future trajectories, which can be unreliable when the sensors are malfunctioning or occluded. To address this problem,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Hsu-kuang Chiu , Ryo Hachiuma , Chien-Yi Wang , Stephen F. Smith , Yu-Chiang Frank Wang , Min-Hung Chen

The ideal form of Visual Question Answering requires understanding, grounding and reasoning in the joint space of vision and language and serves as a proxy for the AI task of scene understanding. However, most existing VQA benchmarks are…

Computer Vision and Pattern Recognition · Computer Science 2023-03-07 Kang Chen , Xiangqian Wu

A fundamental challenge in artificial intelligence involves understanding the cognitive mechanisms underlying visual reasoning in sophisticated models like Vision-Language Models (VLMs). How do these models integrate visual perception with…

Computer Vision and Pattern Recognition · Computer Science 2025-05-07 Mohit Vaishnav , Tanel Tammet

Vision-Language Models (VLMs) have demonstrated remarkable progress in multimodal understanding, yet their capabilities for scientific reasoning remain inadequately assessed. Current multimodal benchmarks predominantly evaluate generic…

Computer Vision and Pattern Recognition · Computer Science 2025-06-18 Ai Jian , Weijie Qiu , Xiaokun Wang , Peiyu Wang , Yunzhuo Hao , Jiangbo Pei , Yichen Wei , Yi Peng , Xuchen Song

Visual Question Answering (VQA) has emerged as a pivotal task in the intersection of computer vision and natural language processing, requiring models to understand and reason about visual content in response to natural language questions.…

Computer Vision and Pattern Recognition · Computer Science 2025-03-05 Aiswarya Baby , Tintu Thankom Koshy

Large Vision Language Models (LVLMs) have shown strong capabilities in understanding and analyzing visual scenes across various domains. However, in the context of autonomous driving, their limited comprehension of 3D environments restricts…

Computer Vision and Pattern Recognition · Computer Science 2025-05-02 Jannik Lübberstedt , Esteban Rivera , Nico Uhlemann , Markus Lienkamp

We present a scalable, bottom-up and intrinsically diverse data collection scheme that can be used for high-level reasoning with long and medium horizons and that has 2.2x higher throughput compared to traditional narrow top-down…

Large vision-language models (LVLMs) are increasingly deployed in interactive applications such as virtual and augmented reality, where a first-person (egocentric) view captured by head-mounted cameras serves as key input. While this view…

Computer Vision and Pattern Recognition · Computer Science 2025-10-27 Insu Lee , Wooje Park , Jaeyun Jang , Minyoung Noh , Kyuhong Shim , Byonghyo Shim

Autonomous driving requires generating safe and reliable trajectories from complex multimodal inputs. Traditional modular pipelines separate perception, prediction, and planning, while recent end-to-end (E2E) systems learn them jointly.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Qihang Peng , Xuesong Chen , Chenye Yang , Shaoshuai Shi , Hongsheng Li

Learning contextual and spatial environmental representations enhances autonomous vehicle's hazard anticipation and decision-making in complex scenarios. Recent perception systems enhance spatial understanding with sensor fusion but often…

Robotics · Computer Science 2024-01-18 Shoaib Azam , Farzeen Munir , Ville Kyrki , Moongu Jeon , Witold Pedrycz

Integrating large language models (LLMs) into autonomous driving motion planning has recently emerged as a promising direction, offering enhanced interpretability, better controllability, and improved generalization in rare and long-tail…

Artificial Intelligence · Computer Science 2025-07-29 Zhipeng Tang , Sha Zhang , Jiajun Deng , Chenjie Wang , Guoliang You , Yuting Huang , Xinrui Lin , Yanyong Zhang

Multimodal large language models (MLLMs) have demonstrated powerful capabilities in general spatial understanding and reasoning. However, their fine-grained spatial understanding and reasoning capabilities in complex urban scenarios have…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Jun Zhang , Jie Feng , Long Chen , Junhui Wang , Zhicheng Liu , Depeng Jin , Yong Li

Reliable autonomous driving requires scene understanding that is semantically consistent across heterogeneous sensors and verifiable at the reasoning stage. However, many recent LLM-driven driving systems attach the language model as a…

Computer Vision and Pattern Recognition · Computer Science 2026-05-07 Shuo Liu , Lei Shi , Haowen Liu , Jing Xu , Yufei Gao , Yucheng Shi

Earth vision research typically focuses on extracting geospatial object locations and categories but neglects the exploration of relations between objects and comprehensive reasoning. Based on city planning needs, we develop a multi-modal…

Computer Vision and Pattern Recognition · Computer Science 2023-12-20 Junjue Wang , Zhuo Zheng , Zihang Chen , Ailong Ma , Yanfei Zhong

Vision-Language Models (VLMs) excel at reasoning in linguistic space but struggle with perceptual understanding that requires dense visual perception, e.g., spatial reasoning and geometric awareness. This limitation stems from the fact that…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Yiming Qin , Bomin Wei , Jiaxin Ge , Konstantinos Kallidromitis , Stephanie Fu , Trevor Darrell , XuDong Wang