English
Related papers

Related papers: Decoding Scientific Experimental Images: The SPUR …

200 papers

Seeking answers to questions within long scientific research articles is a crucial area of study that aids readers in quickly addressing their inquiries. However, existing question-answering (QA) datasets based on scientific papers are…

Computation and Language · Computer Science 2025-01-14 Shraman Pramanick , Rama Chellappa , Subhashini Venugopalan

Visualization, a domain-specific yet widely used form of imagery, is an effective way to turn complex datasets into intuitive insights, and its value depends on whether data are faithfully represented, clearly communicated, and…

Computation and Language · Computer Science 2026-03-03 Yupeng Xie , Zhiyang Zhang , Yifan Wu , Sirong Lu , Jiayi Zhang , Zhaoyang Yu , Jinlin Wang , Sirui Hong , Bang Liu , Chenglin Wu , Yuyu Luo

Multimodal Large Language Models (MLLMs) have shown impressive capabilities in image understanding and generation. However, current benchmarks fail to accurately evaluate the chart comprehension of MLLMs due to limited chart types and…

Computer Vision and Pattern Recognition · Computer Science 2024-06-21 Zhengzhuo Xu , Sinan Du , Yiyan Qi , Chengjin Xu , Chun Yuan , Jian Guo

The ability to process information from multiple modalities and to reason through it step-by-step remains a critical challenge in advancing artificial intelligence. However, existing reasoning benchmarks focus on text-only reasoning, or…

Artificial Intelligence · Computer Science 2025-07-01 Yulun Jiang , Yekun Chai , Maria Brbić , Michael Moor

This paper presents GRASP, a novel benchmark to evaluate the language grounding and physical understanding capabilities of video-based multimodal large language models (LLMs). This evaluation is accomplished via a two-tier approach…

Computation and Language · Computer Science 2024-06-07 Serwan Jassim , Mario Holubar , Annika Richter , Cornelius Wolff , Xenia Ohmer , Elia Bruni

Multimodal Large Language Models (MLLMs) have shown promising potential in diverse understanding tasks, e.g., image and video analysis, math and physics olympiads. However, they remain blank and unexplored for Small Object Understanding…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Fujun Han , Junan Chen , Xintong Zhu , Jingqi Ye , Xuanjie Mao , Tao Chen , Peng Ye

In this paper, the solution of HYU MLLAB KT Team to the Multimodal Algorithmic Reasoning Task: SMART-101 CVPR 2024 Challenge is presented. Beyond conventional visual question-answering problems, the SMART-101 challenge aims to achieve…

Computer Vision and Pattern Recognition · Computer Science 2024-06-11 Jinwoo Ahn , Junhyeok Park , Min-Jun Kim , Kang-Hyeon Kim , So-Yeong Sohn , Yun-Ji Lee , Du-Seong Chang , Yu-Jung Heo , Eun-Sol Kim

Current state of the art measures like BLEU, CIDEr, VQA score, SigLIP-2 and CLIPScore are often unable to capture semantic or structural accuracy, especially for domain-specific or context-dependent scenarios. For this, this paper proposes…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Kishor Datta Gupta , Marufa Kamal , Md. Mahfuzur Rahman , Fahad Rahman , Mohd Ariful Haque , Sunzida Siddique

Multi-modal Large Language Models (MLLMs) exhibit impressive problem-solving abilities in various domains, but their visual comprehension and abstract reasoning skills remain under-evaluated. To this end, we present PolyMATH, a challenging…

Artificial Intelligence · Computer Science 2026-05-12 Himanshu Gupta , Shreyas Verma , Ujjwala Anantheswaran , Kevin Scaria , Mihir Parmar , Swaroop Mishra , Chitta Baral

Multimodal reasoning, which integrates language and visual cues into problem solving and decision making, is a fundamental aspect of human intelligence and a crucial step toward artificial general intelligence. However, the evaluation of…

Unified multimodal models integrate the reasoning capacity of large language models with both image understanding and generation, showing great promise for advanced multimodal intelligence. However, the community still lacks a rigorous…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Hongxiang Li , Yaowei Li , Bin Lin , Yuwei Niu , Yuhang Yang , Xiaoshuang Huang , Jiayin Cai , Xiaolong Jiang , Yao Hu , Long Chen

Recent advancements in Large Vision-Language Models (LVLMs) have significantly enhanced their ability to integrate visual and linguistic information, achieving near-human proficiency in tasks like object recognition, captioning, and visual…

Computer Vision and Pattern Recognition · Computer Science 2025-05-14 Zhikai Wang , Jiashuo Sun , Wenqi Zhang , Zhiqiang Hu , Xin Li , Fan Wang , Deli Zhao

Current multimodal large language models (MLLMs) are mainly focused on the understanding and processing of perceptual modalities such as images and videos, while their capability for scientific data understanding remains insufficient. To…

Artificial Intelligence · Computer Science 2026-05-14 Yanjie Li , Lina Yu , Weijun Li , Min Wu , Liping Zhang , Jingyi Liu , Yusong Deng , Mingzhu Wan , Xin Ning

Referring Expression Comprehension (REC) is a popular multimodal task that aims to accurately detect target objects within a single image based on a given textual expression. However, due to the limitations of earlier models, traditional…

Machine Learning · Computer Science 2025-08-21 Guanghao Jin , Jingpei Wu , Tianpei Guo , Yiyi Niu , Weidong Zhou , Guoyang Liu

The rapid advancement of educational applications, artistic creation, and AI-generated content (AIGC) technologies has substantially increased practical requirements for comprehensive Image Aesthetics Assessment (IAA), particularly…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Shuo Cao , Nan Ma , Jiayang Li , Xiaohui Li , Lihao Shao , Kaiwen Zhu , Yu Zhou , Yuandong Pu , Jiarui Wu , Jiaquan Wang , Bo Qu , Wenhai Wang , Yu Qiao , Dajuin Yao , Yihao Liu

Solving expert-level multimodal tasks is a key milestone towards general intelligence. As the capabilities of multimodal large language models (MLLMs) continue to improve, evaluation of such advanced multimodal intelligence becomes…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Yan Yang , Dongxu Li , Haoning Wu , Bei Chen , Liu Liu , Liyuan Pan , Junnan Li

Millions of people take surveys every day, from market polls and academic studies to medical questionnaires and customer feedback forms. These datasets capture valuable insights, but their scale and structure present a unique challenge for…

Artificial Intelligence · Computer Science 2025-10-31 Duc-Hai Nguyen , Vijayakumar Nanjappan , Barry O'Sullivan , Hoang D. Nguyen

Recent advances in Large Language Models (LLMs) and Large Multimodal Models (LMMs) have improved Document Layout Analysis (DLA), yet structural errors such as region merging, splitting, and omission remain persistent. Conventional…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Inbum Heo , Taewook Hwang , Jeesu Jung , Sangkeun Jung

Fine-grained visual understanding and high-level reasoning in real-world open-water environments remain under-explored due to the lack of dedicated benchmarks. We introduce MARINER, a comprehensive benchmark built under the novel…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Xingming Liao , Ning Chen , Muying Shu , Yunpeng Yin , Peijian Zeng , Zhuowei Wang , Nankai Lin , Lianglun Cheng

Establishing a clear link between model predictions and the visual evidence that supports them is critical for transparency and reliability in multimodal reasoning, yet current multimodal large language model (MLLM) evaluations do not…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Mozhgan Nasr Azadani , Yimu Wang , Yongpeng Zhu , Lihong Chen , Milan Ganai , Sean Sedwards , Marco Pavone , Krzysztof Czarnecki