English
Related papers

Related papers: DriveXQA: Cross-modal Visual Question Answering fo…

200 papers

The visual world around us constantly evolves, from real-time news and social media trends to global infrastructure changes visible through satellite imagery and augmented reality enhancements. However, Multimodal Large Language Models…

Computer Vision and Pattern Recognition · Computer Science 2025-07-02 Mingyang Fu , Yuyang Peng , Dongping Chen , Zetong Zhou , Benlin Liu , Yao Wan , Zhou Zhao , Philip S. Yu , Ranjay Krishna

Multimodal Large Language Models are increasingly applied to biomedical imaging, yet scientific reasoning for microscopy remains limited by the scarcity of large-scale, high-quality training data. We introduce MicroVQA++, a three-stage,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-17 Manyu Li , Ruian He , Chenxi Ma , Weimin Tan , Bo Yan

Lately, researchers in artificial intelligence have been really interested in how language and vision come together, giving rise to the development of multimodal models that aim to seamlessly integrate textual and visual information.…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Rajat Chawla , Arkajit Datta , Tushar Verma , Adarsh Jha , Anmol Gautam , Ayush Vatsal , Sukrit Chaterjee , Mukunda NS , Ishaan Bhola

Is basic visual understanding really solved in state-of-the-art VLMs? We present VisualOverload, a slightly different visual question answering (VQA) benchmark comprising 2,720 question-answer pairs, with privately held ground-truth…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Paul Gavrikov , Wei Lin , M. Jehanzeb Mirza , Soumya Jahagirdar , Muhammad Huzaifa , Sivan Doveh , Serena Yeung-Levy , James Glass , Hilde Kuehne

Multimodal embedding models, built upon causal Vision Language Models (VLMs), have shown promise in various tasks. However, current approaches face three key limitations: the use of causal attention in VLM backbones is suboptimal for…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Haonan Chen , Hong Liu , Yuping Luo , Liang Wang , Nan Yang , Furu Wei , Zhicheng Dou

Modern autonomous driving perception systems utilize complementary multi-modal sensors, such as LiDAR and cameras. Although sensor fusion architectures enhance performance in challenging environments, they still suffer significant…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Konyul Park , Yecheol Kim , Daehun Kim , Jun Won Choi

Multimodal in-context learning (ICL) is becoming a key capability that allows large vision-language models (LVLMs) to adapt to novel tasks without parameter updates, which expands their usefulness in many real-world applications. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-11 Yanshu Li , Jianjiang Yang , Ziteng Yang , Bozheng Li , Ligong Han , Hongyang He , Zhengtao Yao , Yingjie Victor Chen , Songlin Fei , Dongfang Liu , Ruixiang Tang

Tourism and travel planning increasingly rely on digital assistance, yet existing multimodal AI systems often lack specialized knowledge and contextual understanding of urban environments. We present TraveLLaMA, a specialized multimodal…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Meng Chu , Yukang Chen , Haokun Gui , Shaozuo Yu , Yi Wang , Jiaya Jia

Accurate accident anticipation remains challenging when driver cognition and dynamic road conditions are underrepresented in predictive models. In this paper, we propose CAMERA (Context-Aware Multi-modal Enhanced Risk Anticipation), a…

Computational Engineering, Finance, and Science · Computer Science 2025-07-17 Jiaxun Zhang , Haicheng Liao , Yumu Xie , Chengyue Wang , Yanchen Guan , Bin Rao , Zhenning Li

Visual Question Answering (VQA) requires reasoning across visual and textual modalities, yet Large Vision-Language Models (LVLMs) often lack integrated commonsense knowledge, limiting their robustness in real-world scenarios. To address…

Computation and Language · Computer Science 2025-06-12 Shuo Yang , Siwen Luo , Soyeon Caren Han , Eduard Hovy

Urban transportation systems face growing safety challenges that require scalable intelligence for emerging smart mobility infrastructures. While recent advances in foundation models and large-scale multimodal datasets have strengthened…

Computer Vision and Pattern Recognition · Computer Science 2026-04-27 Wenhui Huang , Songyan Zhang , Collister Chua , Yang Liang , Zhiqi Mao , Heng Yang , Chen Lv

Vision-Language-Action (VLA) models for autonomous driving show promise but falter in unstructured corner case scenarios, largely due to a scarcity of targeted benchmarks. To address this, we introduce Impromptu VLA. Our core contribution…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Haohan Chi , Huan-ang Gao , Ziming Liu , Jianing Liu , Chenyu Liu , Jinwei Li , Kaisen Yang , Yangcheng Yu , Zeda Wang , Wenyi Li , Leichen Wang , Xingtao Hu , Hao Sun , Hang Zhao , Hao Zhao

Multimodal large language models (MLLMs), such as GPT-4o, Gemini, LLaVA, and Flamingo, have made significant progress in integrating visual and textual modalities, excelling in tasks like visual question answering (VQA), image captioning,…

Computer Vision and Pattern Recognition · Computer Science 2024-12-31 Junxiao Xue , Quan Deng , Fei Yu , Yanhao Wang , Jun Wang , Yuehua Li

Within the multimodal field, large vision-language models (LVLMs) have made significant progress due to their strong perception and reasoning capabilities in the visual and language systems. However, LVLMs are still plagued by the two…

Computer Vision and Pattern Recognition · Computer Science 2024-06-14 Sirui Cheng , Siyu Zhang , Jiayi Wu , Muchen Lan

Evaluating and Rethinking the current landscape of Large Multimodal Models (LMMs), we observe that widely-used visual-language projection approaches (e.g., Q-former or MLP) focus on the alignment of image-text descriptions yet ignore the…

Computation and Language · Computer Science 2024-06-27 Yunxin Li , Xinyu Chen , Baotian Hu , Haoyuan Shi , Min Zhang

Audio-visual quality assessment (AVQA) is essential for streaming, teleconferencing, and immersive media. In realistic streaming scenarios, distortions are often asymmetric, where one modality may be severely degraded while the other…

Multimedia · Computer Science 2026-05-05 Mayesha Maliha R. Mithila , Mylene C. Q. Farias

Multimodal Large Language Models (MLLMs) demonstrate remarkable fluency in understanding visual scenes, yet they exhibit a critical lack in a fundamental cognitive skill: object counting. This blind spot severely limits their reliability in…

Artificial Intelligence · Computer Science 2025-09-10 Jayant Sravan Tamarapalli , Rynaa Grover , Nilay Pande , Sahiti Yerramilli

In this technical report, we present CarLLaVA, a Vision Language Model (VLM) for autonomous driving, developed for the CARLA Autonomous Driving Challenge 2.0. CarLLaVA uses the vision encoder of the LLaVA VLM and the LLaMA architecture as…

Computer Vision and Pattern Recognition · Computer Science 2024-06-17 Katrin Renz , Long Chen , Ana-Maria Marcu , Jan Hünermann , Benoit Hanotte , Alice Karnsund , Jamie Shotton , Elahe Arani , Oleg Sinavski

Vision-Language Models (VLMs) have demonstrated notable promise in autonomous driving by offering the potential for multimodal reasoning through pretraining on extensive image-text pairs. However, adapting these models from broad web-scale…

Robotics · Computer Science 2025-06-18 Yupeng Zhou , Can Cui , Juntong Peng , Zichong Yang , Juanwu Lu , Jitesh H Panchal , Bin Yao , Ziran Wang

Autonomous driving is a popular research area within the computer vision research community. Since autonomous vehicles are highly safety-critical, ensuring robustness is essential for real-world deployment. While several public multimodal…