中文
相关论文

相关论文: LingoQA: Visual Question Answering for Autonomous …

200 篇论文

Assessing the video comprehension capabilities of multimodal AI systems can effectively measure their understanding and reasoning abilities. Most video evaluation benchmarks are limited to a single language, typically English, and…

计算机视觉与模式识别 · 计算机科学 2025-05-21 Xinyu Chen , Yunxin Li , Haoyuan Shi , Baotian Hu , Wenhan Luo , Yaowei Wang , Min Zhang

We study the visual quality judgments of human subjects on digital human avatars (sometimes referred to as "holograms" in the parlance of virtual reality [VR] and augmented reality [AR] systems) that have been subjected to distortions. We…

图像与视频处理 · 电气工程与系统科学 2024-10-04 Yu-Chih Chen , Avinab Saha , Alexandre Chapiro , Christian Häne , Jean-Charles Bazin , Bo Qiu , Stefano Zanetti , Ioannis Katsavounidis , Alan C. Bovik

In this paper, we propose a new dataset, ReasonVQA, for the Visual Question Answering (VQA) task. Our dataset is automatically integrated with structured encyclopedic knowledge and constructed using a low-cost framework, which is capable of…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Duong T. Tran , Trung-Kien Tran , Manfred Hauswirth , Danh Le Phuoc

Existing video understanding datasets mostly focus on human interactions, with little attention being paid to the "in the wild" settings, where the videos are recorded outdoors. We propose WILDQA, a video understanding dataset of videos…

计算机视觉与模式识别 · 计算机科学 2022-09-15 Santiago Castro , Naihao Deng , Pingxuan Huang , Mihai Burzo , Rada Mihalcea

Social media imagery provides a low-latency source of situational information during natural and human-induced disasters, enabling rapid damage assessment and response. While Visual Question Answering (VQA) has shown strong performance in…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Aisha Al-Mohannadi , Ayisha Firoz , Yin Yang , Muhammad Imran , Ferda Ofli

\Ac{LFQA} aims to generate lengthy answers to complex questions. This scenario presents great flexibility as well as significant challenges for evaluation. Most evaluations rely on deterministic metrics that depend on string or n-gram…

信息检索 · 计算机科学 2025-04-28 Ning Xian , Yixing Fan , Ruqing Zhang , Maarten de Rijke , Jiafeng Guo

In this technical report, we present CarLLaVA, a Vision Language Model (VLM) for autonomous driving, developed for the CARLA Autonomous Driving Challenge 2.0. CarLLaVA uses the vision encoder of the LLaVA VLM and the LLaMA architecture as…

计算机视觉与模式识别 · 计算机科学 2024-06-17 Katrin Renz , Long Chen , Ana-Maria Marcu , Jan Hünermann , Benoit Hanotte , Alice Karnsund , Jamie Shotton , Elahe Arani , Oleg Sinavski

In this paper, we introduce SecQA, a novel dataset tailored for evaluating the performance of Large Language Models (LLMs) in the domain of computer security. Utilizing multiple-choice questions generated by GPT-4 based on the "Computer…

计算与语言 · 计算机科学 2023-12-27 Zefang Liu

State-of-the-art vision-language models (VLMs) score impressively on video benchmarks yet stumble on basic visual reasoning tasks involving spatial relations, navigation, and object selection that a preschooler solves easily. We hypothesize…

计算机视觉与模式识别 · 计算机科学 2026-05-11 Bishoy Galoaa , Xiangyu Bai , Sarah Ostadabbas

NLP models today strive for supporting multiple languages and modalities, improving accessibility for diverse users. In this paper, we evaluate their multilingual, multimodal capabilities by testing on a visual reasoning task. We observe…

计算与语言 · 计算机科学 2025-02-11 Yueqi Song , Simran Khanuja , Graham Neubig

In healthcare and medical diagnostics, Visual Question Answering (VQA) mayemergeasapivotal tool in scenarios where analysis of intricate medical images becomes critical for accurate diagnoses. Current text-based VQA systems limit their…

计算机视觉与模式识别 · 计算机科学 2024-07-17 Tonmoy Rajkhowa , Amartya Roy Chowdhury , Sankalp Nagaonkar , Achyut Mani Tripathi

Visual Question Answering (VQA) models play a critical role in enhancing the perception capabilities of autonomous driving systems by allowing vehicles to analyze visual inputs alongside textual queries, fostering natural interaction and…

计算机视觉与模式识别 · 计算机科学 2024-06-14 Kaavya Rekanar , Martin Hayes , Ganesh Sistu , Ciaran Eising

Humans apprehend the world through various sensory modalities, yet language is their predominant communication channel. Machine learning systems need to draw on the same multimodal richness to have informed discourses with humans in natural…

计算机视觉与模式识别 · 计算机科学 2022-08-25 Min Wang , Ata Mahjoubfar , Anupama Joshi

Multi-modality foundation models, as represented by GPT-4V, have brought a new paradigm for low-level visual perception and understanding tasks, that can respond to a broad range of natural human instructions in a model. While existing…

计算机视觉与模式识别 · 计算机科学 2023-11-14 Haoning Wu , Zicheng Zhang , Erli Zhang , Chaofeng Chen , Liang Liao , Annan Wang , Kaixin Xu , Chunyi Li , Jingwen Hou , Guangtao Zhai , Geng Xue , Wenxiu Sun , Qiong Yan , Weisi Lin

We present LLoVi, a language-based framework for long-range video question-answering (LVQA). Unlike prior long-range video understanding methods, which are often costly and require specialized long-range video modeling design (e.g., memory…

计算机视觉与模式识别 · 计算机科学 2024-10-11 Ce Zhang , Taixi Lu , Md Mohaiminul Islam , Ziyang Wang , Shoubin Yu , Mohit Bansal , Gedas Bertasius

Data visualizations are powerful tools for communicating patterns in quantitative data. Yet understanding any data visualization is no small feat -- succeeding requires jointly making sense of visual, numerical, and linguistic inputs…

人机交互 · 计算机科学 2025-05-26 Arnav Verma , Kushin Mukherjee , Christopher Potts , Elisa Kreiss , Judith E. Fan

Automated data visualization plays a crucial role in simplifying data interpretation, enhancing decision-making, and improving efficiency. While large language models (LLMs) have shown promise in generating visualizations from natural…

计算与语言 · 计算机科学 2025-07-29 Mizanur Rahman , Md Tahmid Rahman Laskar , Shafiq Joty , Enamul Hoque

We introduce ScreenQA, a novel benchmarking dataset designed to advance screen content understanding through question answering. The existing screen datasets are focused either on low-level structural and component understanding, or on a…

In this paper, we critically evaluate the capabilities of the state-of-the-art multimodal large language model, i.e., GPT-4 with Vision (GPT-4V), on Visual Question Answering (VQA) task. Our experiments thoroughly assess GPT-4V's…

计算机视觉与模式识别 · 计算机科学 2023-10-31 Zhiling Yan , Kai Zhang , Rong Zhou , Lifang He , Xiang Li , Lichao Sun

In this paper, we present a hierarchical question-answering (QA) approach for scene understanding in autonomous vehicles, balancing cost-efficiency with detailed visual interpretation. The method fine-tunes a compact vision-language model…

计算机视觉与模式识别 · 计算机科学 2025-06-04 Safaa Abdullahi Moallim Mohamud , Minjin Baek , Dong Seog Han