English
Related papers

Related papers: ALIGN: A Vision-Language Framework for High-Accura…

200 papers

Multi-view 3D visual grounding is critical for autonomous driving vehicles to interpret natural languages and localize target objects in complex environments. However, existing datasets and methods suffer from coarse-grained language…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Fuhao Li , Huan Jin , Bin Gao , Liaoyuan Fan , Lihui Jiang , Long Zeng

Video-based spatial reasoning -- such as estimating distances, judging directions, or understanding layouts from multiple views -- requires selecting informative frames and, when needed, actively seeking additional viewpoints during…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Jiaxu Wan , Xu Wang , Mengwei Xie , Hang Zhang , Mu Xu , Yang Han , Hong Zhang , Ding Yuan , Yifan Yang

Localizing objects and parts from natural language in 3D space is essential for robotics, AR, and embodied AI, yet existing methods face a trade-off between the accuracy and geometric consistency of per-scene optimization and the efficiency…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Bryce Grant , Aryeh Rothenberg , Atri Banerjee , Peng Wang

Driver visual attention prediction is a critical task in autonomous driving and human-computer interaction (HCI) research. Most prior studies focus on estimating attention allocation at a single moment in time, typically using static RGB…

Computer Vision and Pattern Recognition · Computer Science 2025-08-11 Kaiser Hamid , Khandakar Ashrafi Akbar , Nade Liang

Large reasoning models (LRMs) show strong capabilities in complex reasoning, yet their marginal gains on evidence-dependent factual questions are limited. We find this limitation is partially attributable to a reasoning-answer hit gap,…

Computation and Language · Computer Science 2026-01-06 Xinming Wang , Jian Xu , Bin Yu , Sheng Lian , Hongzhu Yi , Yi Chen , Yingjian Zhu , Boran Wang , Hongming Yang , Han Hu , Xu-Yao Zhang , Cheng-Lin Liu

Integrating large language models (LLMs) into embodied AI models is becoming increasingly prevalent. However, existing zero-shot LLM-based Vision-and-Language Navigation (VLN) agents either encode images as textual scene descriptions,…

Artificial Intelligence · Computer Science 2025-09-30 Yue Zhang , Tianyi Ma , Zun Wang , Yanyuan Qiao , Parisa Kordjamshidi

Visual-language grounding aims to establish semantic correspondences between natural language and visual entities, enabling models to accurately identify and localize target objects based on textual instructions. Existing VLG approaches…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Linfei Li , Lin Zhang , Ying Shen

Recent advances in Multimodal Large Language Models (MLLMs) have achieved remarkable progress in general domains and demonstrated promise in multimodal mathematical reasoning. However, applying MLLMs to geometry problem solving (GPS)…

Computation and Language · Computer Science 2025-04-18 Yicheng Pan , Zhenrong Zhang , Pengfei Hu , Jiefeng Ma , Jun Du , Jianshu Zhang , Quan Liu , Jianqing Gao , Feng Ma

This paper introduces a multi-agent framework for comprehensive highway scene understanding, designed around a mixture-of-experts strategy. In this framework, a large generic vision-language model (VLM), such as GPT-4o, is contextualized…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Yunxiang Yang , Ningning Xu , Jidong J. Yang

Open-vocabulary grounding requires accurate vision-language alignment under weak supervision, yet existing methods either rely on global sentence embeddings that lack fine-grained expressiveness or introduce token-level alignment with…

Computer Vision and Pattern Recognition · Computer Science 2026-02-02 Junyi Hu , Tian Bai , Fengyi Wu , Wenyan Li , Zhenming Peng , Yi Zhang

Understanding and addressing corner cases is essential for ensuring the safety and reliability of autonomous driving systems. Vision-language models (VLMs) play a crucial role in enhancing scenario comprehension, yet they face significant…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Yujin Wang , Quanfeng Liu , Jiaqi Fan , Jinlong Hong , Hongqing Chu , Mengjian Tian , Bingzhao Gao , Hong Chen

Referential grounding in outdoor driving scenes is challenging due to large scene variability, many visually similar objects, and dynamic elements that complicate resolving natural-language references (e.g., "the black car on the right").…

Computer Vision and Pattern Recognition · Computer Science 2025-10-22 Pranav Saxena , Avigyan Bhattacharya , Ji Zhang , Wenshan Wang

End-to-end autonomous driving has emerged as a promising approach to unify perception, prediction, and planning within a single framework, reducing information loss and improving adaptability. However, existing methods often rely on fixed…

Robotics · Computer Science 2025-07-18 Yuhang Lu , Jiadong Tu , Yuexin Ma , Xinge Zhu

Rapid advances in multimodal models demand benchmarks that rigorously evaluate understanding and reasoning in safety-critical, dynamic real-world settings. We present AccidentBench, a large-scale benchmark that combines vehicle accident…

Traffic forecasting represents a crucial problem within intelligent transportation systems. In recent research, Large Language Models (LLMs) have emerged as a promising method, but their intrinsic design, tailored primarily for sequential…

Machine Learning · Computer Science 2025-09-18 Hyotaek Jeon , Hyunwook Lee , Juwon Kim , Sungahn Ko

Efficient target localization and autonomous navigation in complex environments are fundamental to real-world embodied applications. While recent advances in multimodal foundation models have enabled zero-shot object goal navigation,…

Robotics · Computer Science 2026-04-02 Ming-Ming Yu , Yi Chen , Börje F. Karlsson , Wenjun Wu

Despite the promise of Retrieval-Augmented Generation in grounding Multimodal Large Language Models with external knowledge, the transition to extensive contexts often leads to significant attention dilution and reasoning hallucinations.…

Computation and Language · Computer Science 2026-03-10 Junming Liu , Yuqi Li , Shiping Wen , Zhigang Zeng , Tingwen Huang

Safety on roads is of uttermost importance, especially in the context of autonomous vehicles. A critical need is to detect and communicate disruptive incidents early and effectively. In this paper we propose a system based on an…

Computer Vision and Pattern Recognition · Computer Science 2022-03-24 Alex Levering , Martin Tomko , Devis Tuia , Kourosh Khoshelham

Equitable urban transportation applications require high-fidelity digital representations of the built environment: not just streets and sidewalks, but bike lanes, marked and unmarked crossings, curb ramps and cuts, obstructions, traffic…

Computer Vision and Pattern Recognition · Computer Science 2024-08-05 Bin Han , Yiwei Yang , Anat Caspi , Bill Howe

Current Large Language Models have achieved Olympiad-level logic, yet Vision-Language Models paradoxically falter on elementary spatial tasks like block counting. This capability mismatch reveals a critical ``spatial intelligence gap,''…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Shaoxiong Zhan , Yanlin Lai , Zheng Liu , Hai Lin , Shen Li , Xiaodong Cai , Zijian Lin , Wen Huang , Hai-Tao Zheng
‹ Prev 1 3 4 5 6 7 10 Next ›