English
Related papers

Related papers: MTVQA: Benchmarking Multilingual Text-Centric Visu…

200 papers

Charts are a universally adopted medium for data communication, yet existing chart understanding benchmarks are overwhelmingly English-centric, limiting their accessibility and relevance to global audiences. To address this limitation, we…

Computation and Language · Computer Science 2026-01-09 Yichen Xu , Liangyu Chen , Liang Zhang , Jianzhe Ma , Wenxuan Wang , Qin Jin

The current research direction in generative models, such as the recently developed GPT4, aims to find relevant knowledge information for multimodal and multilingual inputs to provide answers. Under these research circumstances, the demand…

Computation and Language · Computer Science 2024-08-01 Minjun Kim , Seungwoo Song , Youhan Lee , Haneol Jang , Kyungtae Lim

Spectra are a prevalent yet highly information-dense form of scientific imagery, presenting substantial challenges to multimodal large language models (MLLMs) due to their unstructured and domain-specific characteristics. Here we introduce…

Artificial Intelligence · Computer Science 2026-05-01 Jialu Shen , Han Lyu , Suyang Zhong , Hanzheng Li , Haoyi Tao , Nan Wang , Changhong Chen , Xi Fang

We present TUMTraffic-VideoQA, a novel dataset and benchmark designed for spatio-temporal video understanding in complex roadside traffic scenarios. The dataset comprises 1,000 videos, featuring 85,000 multiple-choice QA pairs, 2,300 object…

Computer Vision and Pattern Recognition · Computer Science 2025-02-05 Xingcheng Zhou , Konstantinos Larintzakis , Hao Guo , Walter Zimmer , Mingyu Liu , Hu Cao , Jiajie Zhang , Venkatnarayanan Lakshminarasimhan , Leah Strand , Alois C. Knoll

Visual Question Answering(VQA) is a highly complex problem set, relying on many sub-problems to produce reasonable answers. In this paper, we present the hypothesis that Visual Question Answering should be viewed as a multi-task problem,…

Computer Vision and Pattern Recognition · Computer Science 2020-07-06 Amelia Elizabeth Pollard , Jonathan L. Shapiro

Real-world clinical practice demands multi-image comparative reasoning, yet current medical benchmarks remain limited to single-frame interpretation. We present MedFrameQA, the first benchmark explicitly designed to test multi-image medical…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Suhao Yu , Haojin Wang , Juncheng Wu , Luyang Luo , Jingshen Wang , Cihang Xie , Pranav Rajpurkar , Carl Yang , Yang Yang , Kang Wang , Yannan Yu , Yuyin Zhou

Visual Question Answering (VQA) has emerged as a Visual Turing Test to validate the reasoning ability of AI agents. The pivot to existing VQA models is the joint embedding that is learned by combining the visual features from an image and…

Computer Vision and Pattern Recognition · Computer Science 2020-01-22 Moshiur R. Farazi , Salman H. Khan , Nick Barnes

Building a reliable visual question answering~(VQA) system across different languages is a challenging problem, primarily due to the lack of abundant samples for training. To address this challenge, recent studies have employed machine…

Computation and Language · Computer Science 2024-06-05 ChaeHun Park , Koanho Lee , Hyesu Lim , Jaeseok Kim , Junmo Park , Yu-Jung Heo , Du-Seong Chang , Jaegul Choo

Performance on the most commonly used Visual Question Answering dataset (VQA v2) is starting to approach human accuracy. However, in interacting with state-of-the-art VQA models, it is clear that the problem is far from being solved. In…

Computer Vision and Pattern Recognition · Computer Science 2021-06-07 Sasha Sheng , Amanpreet Singh , Vedanuj Goswami , Jose Alberto Lopez Magana , Wojciech Galuba , Devi Parikh , Douwe Kiela

As Large Language Models (LLMs) are increasingly popularized in the multilingual world, ensuring hallucination-free factuality becomes markedly crucial. However, existing benchmarks for evaluating the reliability of Multimodal Large…

Computation and Language · Computer Science 2026-01-28 Yexing Du , Kaiyuan Liu , Youcheng Pan , Zheng Chu , Bo Yang , Xiaocheng Feng , Ming Liu , Yang Xiang

Visual Question Answering (VQA) is increasingly used in diverse applications ranging from general visual reasoning to safety-critical domains such as medical imaging and autonomous systems, where models must provide not only accurate…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Xingjian Diao , Weiyi Wu , Keyi Kong , Peijun Qing , Xinwen Xu , Ming Cheng , Soroush Vosoughi , Jiang Gui

The predominant approach to Visual Question Answering (VQA) demands that the model represents within its weights all of the information required to answer any question about any image. Learning this information from any real training set…

Computer Vision and Pattern Recognition · Computer Science 2017-11-23 Damien Teney , Anton van den Hengel

Text-to-image generative models excel in creating images from text but struggle with ensuring alignment and consistency between outputs and prompts. This paper introduces TextMatch, a novel framework that leverages multimodal optimization…

Computer Vision and Pattern Recognition · Computer Science 2025-01-28 Yucong Luo , Mingyue Cheng , Jie Ouyang , Xiaoyu Tao , Qi Liu

Vision Language Models (VLMs) have recently shown significant advancements in video understanding, especially in feature alignment, event reasoning, and instruction-following tasks. However, their capability for counterfactual reasoning,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Yuefei Chen , Jiang Liu , Xiaodong Lin , Ruixiang Tang

Visual Question Answering (VQA) with multiple choice questions enables a vision-centric evaluation of Multimodal Large Language Models (MLLMs). Although it reliably checks the existence of specific visual abilities, it is easier for the…

Computer Vision and Pattern Recognition · Computer Science 2024-09-24 Manu Gaur , Darshan Singh S , Makarand Tapaswi

Optical Character Recognition - Visual Question Answering (OCR-VQA) is the task of answering text information contained in images that have just been significantly developed in the English language in recent years. However, there are…

Computer Vision and Pattern Recognition · Computer Science 2024-04-30 Huy Quang Pham , Thang Kien-Bao Nguyen , Quan Van Nguyen , Dan Quang Tran , Nghia Hieu Nguyen , Kiet Van Nguyen , Ngan Luu-Thuy Nguyen

Multi-modal Large Language Models (MLLMs) are gaining significant attention for their ability to process multi-modal data, providing enhanced contextual understanding of complex problems. MLLMs have demonstrated exceptional capabilities in…

Computer Vision and Pattern Recognition · Computer Science 2025-09-09 Pragati Shuddhodhan Meshram , Swetha Karthikeyan , Bhavya Bhavya , Suma Bhat

Video Question Answering (VideoQA) is a challenging task that requires understanding complex visual and temporal relationships within videos to answer questions accurately. In this work, we introduce \textbf{ReasVQA} (Reasoning-enhanced…

Computer Vision and Pattern Recognition · Computer Science 2025-01-24 Jianxin Liang , Xiaojun Meng , Huishuai Zhang , Yueqian Wang , Jiansheng Wei , Dongyan Zhao

A hierarchical cross-modal fusion model is proposed for vision-language question answering (VLQA) in industrial robotics, targeting the challenges of semantic ambiguity, complex environmental layouts, and domain-specific terminology common…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Ping Li , Bartlomiej Brzozka

Visual Question Answering (VQA) within the surgical domain, utilizing Large Language Models (LLMs), offers a distinct opportunity to improve intra-operative decision-making and facilitate intuitive surgeon-AI interaction. However, the…

Computer Vision and Pattern Recognition · Computer Science 2024-05-24 Runlong He , Mengya Xu , Adrito Das , Danyal Z. Khan , Sophia Bano , Hani J. Marcus , Danail Stoyanov , Matthew J. Clarkson , Mobarakol Islam
‹ Prev 1 8 9 10 Next ›