English
Related papers

Related papers: Multi-Modal Vision vs. Text-Based Parsing: Benchma…

200 papers

Multimodal Large Language Models (MLLMs) have become a powerful tool for integrating visual and textual information. Despite their exceptional performance on visual understanding benchmarks, measuring their ability to reason abstractly…

Computer Vision and Pattern Recognition · Computer Science 2026-02-26 Nilay Yilmaz , Maitreya Patel , Yiran Lawrence Luo , Tejas Gokhale , Chitta Baral , Suren Jayasuriya , Yezhou Yang

Large language models are effective at few-shot in-context learning (ICL). Recent advancements in multimodal foundation models have enabled unprecedentedly long context windows, presenting an opportunity to explore their capability to…

Machine Learning · Computer Science 2024-10-08 Yixing Jiang , Jeremy Irvin , Ji Hun Wang , Muhammad Ahmed Chaudhry , Jonathan H. Chen , Andrew Y. Ng

This paper introduces Code-Vision, a benchmark designed to evaluate the logical understanding and code generation capabilities of Multimodal Large Language Models (MLLMs). It challenges MLLMs to generate a correct program that fulfills…

Computation and Language · Computer Science 2025-02-18 Hanbin Wang , Xiaoxuan Zhou , Zhipeng Xu , Keyuan Cheng , Yuxin Zuo , Kai Tian , Jingwei Song , Junting Lu , Wenhui Hu , Xueyang Liu

Recent advancements in Retrieval-Augmented Generation (RAG) have enabled Large Language Models (LLMs) to access multimodal knowledge bases containing both text and visual information such as charts, diagrams, and tables in financial…

In this report, we introduce InternVL 1.5, an open-source multimodal large language model (MLLM) to bridge the capability gap between open-source and proprietary commercial models in multimodal understanding. We introduce three simple…

Multimodal Large Language Models (MLLMs) have demonstrated strong performance across a wide range of vision-language tasks, yet their internal processing dynamics remain underexplored. In this work, we introduce a probing framework to…

Computer Vision and Pattern Recognition · Computer Science 2025-08-29 Zhuoran Yu , Yong Jae Lee

This study introduces a benchmark framework for evaluating the financial decision-making capabilities of large language models (LLMs) through portfolio optimization problems with mathematically explicit solutions. Unlike existing financial…

Portfolio Management · Quantitative Finance 2026-05-28 Hanyong Cho , Jang Ho Kim

Vision-language models (VLMs) are widely assumed to exhibit in-context learning (ICL), a property similar to that of their language-only counterparts. While recent work suggests VLMs can perform multimodal ICL (MM-ICL), studies show they…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Chengyue Huang , Yuchen Zhu , Sichen Zhu , Jingyun Xiao , Moises Andrade , Shivang Chopra , Zsolt Kira

Multimodal foundation models (MFMs), such as GPT-4o, have recently made remarkable progress. However, their detailed visual understanding beyond question answering remains unclear. In this paper, we benchmark popular MFMs (GPT-4o, o4-mini,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-04 Rahul Ramachandran , Ali Garjani , Roman Bachmann , Andrei Atanov , Oğuzhan Fatih Kar , Amir Zamir

Large Language Models (LLMs) are trained on massive amounts of data, enabling their application across diverse domains and tasks. Despite their remarkable performance, most LLMs are developed and evaluated primarily in English. Recently, a…

Computation and Language · Computer Science 2024-10-18 Krishno Dey , Prerona Tarannum , Md. Arid Hasan , Imran Razzak , Usman Naseem

Vision-language models (VLMs) are increasingly proposed as general-purpose solutions for visual recognition tasks, yet their reliability for agricultural decision support remains poorly understood. We benchmark a diverse set of open-source…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Earl Ranario , Mason J. Earles

This study evaluates the capabilities of Multimodal Large Language Models (LLMs) and Vision Language Models (VLMs) in the task of single-label classification of Christian Iconography. The goal was to assess whether general-purpose VLMs…

Computer Vision and Pattern Recognition · Computer Science 2025-09-24 Gianmarco Spinaci , Lukas Klic , Giovanni Colavizza

Since the release of ChatGPT, the field of Natural Language Processing has experienced rapid advancements, particularly in Large Language Models (LLMs) and their multimodal counterparts, Large Multimodal Models (LMMs). Despite their…

Computation and Language · Computer Science 2024-08-27 Florian Schneider , Sunayana Sitaram

Recent multimodal image generators such as GPT-4o, Gemini 2.0 Flash, and Gemini 2.5 Pro excel at following complex instructions, editing images and maintaining concept consistency. However, they are still evaluated by disjoint toolkits:…

Computer Vision and Pattern Recognition · Computer Science 2025-05-29 Hang Hua , Ziyun Zeng , Yizhi Song , Yunlong Tang , Liu He , Daniel Aliaga , Wei Xiong , Jiebo Luo

In contemporary workplaces, meetings are essential for exchanging ideas and ensuring team alignment but often face challenges such as time consumption, scheduling conflicts, and inefficient participation. Recent advancements in Large…

Computation and Language · Computer Science 2025-02-10 Lingxiang Hu , Shurun Yuan , Xiaoting Qin , Jue Zhang , Qingwei Lin , Dongmei Zhang , Saravan Rajmohan , Qi Zhang

Large multimodal models (LMMs) have demonstrated significant potential as generalists in vision-language (VL) tasks. However, adoption of LMMs in real-world tasks is hindered by their poor performance in tasks that require a combination of…

Computation and Language · Computer Science 2025-12-15 Zhoutong Ye , Mingze Sun , Huan-ang Gao , Xutong Wang , Xiangyang Wang , Yu Mei , Chang Liu , Qinwei Li , Chengwen Zhang , Qinghuan Lan , Chun Yu , Yuanchun Shi

Multimodal large language models (MLLMs) have shown success in vision-language tasks, but their ability to reason over complex educational materials remains largely untested. This work presents the first evaluation of state-of-the-art…

Computation and Language · Computer Science 2025-07-16 Hessa A. Alawwad , Anas Zafar , Areej Alhothali , Usman Naseem , Ali Alkhathlan , Amani Jamal

Scene understanding is critical for various downstream tasks in autonomous driving, including facilitating driver-agent communication and enhancing human-centered explainability of autonomous vehicle (AV) decisions. This paper evaluates the…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 Mohammed Elhenawy , Shadi Jaradat , Taqwa I. Alhadidi , Huthaifa I. Ashqar , Ahmed Jaber , Andry Rakotonirainy , Mohammad Abu Tami

A recent advancement in Multimodal Large Language Models (MLLMs) research is the emergence of "reasoning MLLMs" that offer explicit control over their internal thinking processes (normally referred as the "thinking mode") alongside the…

Computation and Language · Computer Science 2025-11-06 Jindong Hong , Tianjie Chen , Lingjie Luo , Chuanyang Zheng , Ting Xu , Haibao Yu , Jianing Qiu , Qianzhong Chen , Suning Huang , Yan Xu , Yong Gui , Yijun He , Jiankai Sun

Multimodal Large Language Models (MLLMs) promise advanced vision language capabilities, yet their effectiveness in visually presented mathematics remains underexplored. This paper analyzes the development and evaluation of MLLMs for…