English
Related papers

Related papers: LingoQA: Visual Question Answering for Autonomous …

200 papers

Assessing the video comprehension capabilities of multimodal AI systems can effectively measure their understanding and reasoning abilities. Most video evaluation benchmarks are limited to a single language, typically English, and…

Computer Vision and Pattern Recognition · Computer Science 2025-05-21 Xinyu Chen , Yunxin Li , Haoyuan Shi , Baotian Hu , Wenhan Luo , Yaowei Wang , Min Zhang

We study the visual quality judgments of human subjects on digital human avatars (sometimes referred to as "holograms" in the parlance of virtual reality [VR] and augmented reality [AR] systems) that have been subjected to distortions. We…

Image and Video Processing · Electrical Eng. & Systems 2024-10-04 Yu-Chih Chen , Avinab Saha , Alexandre Chapiro , Christian Häne , Jean-Charles Bazin , Bo Qiu , Stefano Zanetti , Ioannis Katsavounidis , Alan C. Bovik

In this paper, we propose a new dataset, ReasonVQA, for the Visual Question Answering (VQA) task. Our dataset is automatically integrated with structured encyclopedic knowledge and constructed using a low-cost framework, which is capable of…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Duong T. Tran , Trung-Kien Tran , Manfred Hauswirth , Danh Le Phuoc

Existing video understanding datasets mostly focus on human interactions, with little attention being paid to the "in the wild" settings, where the videos are recorded outdoors. We propose WILDQA, a video understanding dataset of videos…

Computer Vision and Pattern Recognition · Computer Science 2022-09-15 Santiago Castro , Naihao Deng , Pingxuan Huang , Mihai Burzo , Rada Mihalcea

Social media imagery provides a low-latency source of situational information during natural and human-induced disasters, enabling rapid damage assessment and response. While Visual Question Answering (VQA) has shown strong performance in…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Aisha Al-Mohannadi , Ayisha Firoz , Yin Yang , Muhammad Imran , Ferda Ofli

\Ac{LFQA} aims to generate lengthy answers to complex questions. This scenario presents great flexibility as well as significant challenges for evaluation. Most evaluations rely on deterministic metrics that depend on string or n-gram…

Information Retrieval · Computer Science 2025-04-28 Ning Xian , Yixing Fan , Ruqing Zhang , Maarten de Rijke , Jiafeng Guo

In this technical report, we present CarLLaVA, a Vision Language Model (VLM) for autonomous driving, developed for the CARLA Autonomous Driving Challenge 2.0. CarLLaVA uses the vision encoder of the LLaVA VLM and the LLaMA architecture as…

Computer Vision and Pattern Recognition · Computer Science 2024-06-17 Katrin Renz , Long Chen , Ana-Maria Marcu , Jan Hünermann , Benoit Hanotte , Alice Karnsund , Jamie Shotton , Elahe Arani , Oleg Sinavski

In this paper, we introduce SecQA, a novel dataset tailored for evaluating the performance of Large Language Models (LLMs) in the domain of computer security. Utilizing multiple-choice questions generated by GPT-4 based on the "Computer…

Computation and Language · Computer Science 2023-12-27 Zefang Liu

State-of-the-art vision-language models (VLMs) score impressively on video benchmarks yet stumble on basic visual reasoning tasks involving spatial relations, navigation, and object selection that a preschooler solves easily. We hypothesize…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Bishoy Galoaa , Xiangyu Bai , Sarah Ostadabbas

NLP models today strive for supporting multiple languages and modalities, improving accessibility for diverse users. In this paper, we evaluate their multilingual, multimodal capabilities by testing on a visual reasoning task. We observe…

Computation and Language · Computer Science 2025-02-11 Yueqi Song , Simran Khanuja , Graham Neubig

In healthcare and medical diagnostics, Visual Question Answering (VQA) mayemergeasapivotal tool in scenarios where analysis of intricate medical images becomes critical for accurate diagnoses. Current text-based VQA systems limit their…

Computer Vision and Pattern Recognition · Computer Science 2024-07-17 Tonmoy Rajkhowa , Amartya Roy Chowdhury , Sankalp Nagaonkar , Achyut Mani Tripathi

Visual Question Answering (VQA) models play a critical role in enhancing the perception capabilities of autonomous driving systems by allowing vehicles to analyze visual inputs alongside textual queries, fostering natural interaction and…

Computer Vision and Pattern Recognition · Computer Science 2024-06-14 Kaavya Rekanar , Martin Hayes , Ganesh Sistu , Ciaran Eising

Humans apprehend the world through various sensory modalities, yet language is their predominant communication channel. Machine learning systems need to draw on the same multimodal richness to have informed discourses with humans in natural…

Computer Vision and Pattern Recognition · Computer Science 2022-08-25 Min Wang , Ata Mahjoubfar , Anupama Joshi

Multi-modality foundation models, as represented by GPT-4V, have brought a new paradigm for low-level visual perception and understanding tasks, that can respond to a broad range of natural human instructions in a model. While existing…

Computer Vision and Pattern Recognition · Computer Science 2023-11-14 Haoning Wu , Zicheng Zhang , Erli Zhang , Chaofeng Chen , Liang Liao , Annan Wang , Kaixin Xu , Chunyi Li , Jingwen Hou , Guangtao Zhai , Geng Xue , Wenxiu Sun , Qiong Yan , Weisi Lin

We present LLoVi, a language-based framework for long-range video question-answering (LVQA). Unlike prior long-range video understanding methods, which are often costly and require specialized long-range video modeling design (e.g., memory…

Computer Vision and Pattern Recognition · Computer Science 2024-10-11 Ce Zhang , Taixi Lu , Md Mohaiminul Islam , Ziyang Wang , Shoubin Yu , Mohit Bansal , Gedas Bertasius

Data visualizations are powerful tools for communicating patterns in quantitative data. Yet understanding any data visualization is no small feat -- succeeding requires jointly making sense of visual, numerical, and linguistic inputs…

Human-Computer Interaction · Computer Science 2025-05-26 Arnav Verma , Kushin Mukherjee , Christopher Potts , Elisa Kreiss , Judith E. Fan

Automated data visualization plays a crucial role in simplifying data interpretation, enhancing decision-making, and improving efficiency. While large language models (LLMs) have shown promise in generating visualizations from natural…

Computation and Language · Computer Science 2025-07-29 Mizanur Rahman , Md Tahmid Rahman Laskar , Shafiq Joty , Enamul Hoque

We introduce ScreenQA, a novel benchmarking dataset designed to advance screen content understanding through question answering. The existing screen datasets are focused either on low-level structural and component understanding, or on a…

Computation and Language · Computer Science 2025-02-11 Yu-Chung Hsiao , Fedir Zubach , Gilles Baechler , Srinivas Sunkara , Victor Carbune , Jason Lin , Maria Wang , Yun Zhu , Jindong Chen

In this paper, we critically evaluate the capabilities of the state-of-the-art multimodal large language model, i.e., GPT-4 with Vision (GPT-4V), on Visual Question Answering (VQA) task. Our experiments thoroughly assess GPT-4V's…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Zhiling Yan , Kai Zhang , Rong Zhou , Lifang He , Xiang Li , Lichao Sun

In this paper, we present a hierarchical question-answering (QA) approach for scene understanding in autonomous vehicles, balancing cost-efficiency with detailed visual interpretation. The method fine-tunes a compact vision-language model…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Safaa Abdullahi Moallim Mohamud , Minjin Baek , Dong Seog Han