English
Related papers

Related papers: WildTableBench: Benchmarking Multimodal Foundation…

200 papers

Multimodal tables i.e. tabular layouts interleaved with charts, maps, icons, and color encodings are ubiquitous in real applications yet remain difficult for Multimodal Large Language Models (MLLMs). Despite advances in text and image…

Computer Vision and Pattern Recognition · Computer Science 2026-05-28 Prasham Titiya , Jainil Trivedi , Chitta Baral , Vivek Gupta

We introduce TableVista, a comprehensive benchmark for evaluating foundation models in multimodal table reasoning under visual and structural complexity. TableVista consists of 3,000 high-quality table reasoning problems, where each…

Computation and Language · Computer Science 2026-05-08 Zheyuan Yang , Liqiang Shang , Junjie Chen , Xun Yang , Chenglong Xu , Bo Yuan , Chenyuan Jiao , Yaoru Sun , Yilun Zhao

Learning multimodal representations involves integrating information from multiple heterogeneous sources of data. It is a challenging yet crucial area with numerous real-world applications in multimedia, affective computing, robotics,…

We introduce MuirBench, a comprehensive benchmark that focuses on robust multi-image understanding capabilities of multimodal LLMs. MuirBench consists of 12 diverse multi-image tasks (e.g., scene understanding, ordering) that involve 10…

This article introduces a benchmark designed to evaluate the capabilities of multimodal models in analyzing and interpreting images. The benchmark focuses on seven key visual aspects: main object, additional objects, background, detail,…

Computer Vision and Pattern Recognition · Computer Science 2025-01-15 Evgenii Evstafev

While various multimodal multi-image evaluation datasets have been emerged, but these datasets are primarily based on English, and there has yet to be a Chinese multi-image dataset. To fill this gap, we introduce RealBench, the first…

Computation and Language · Computer Science 2025-09-23 Fei Zhao , Chengqiang Lu , Yufan Shen , Qimeng Wang , Yicheng Qian , Haoxin Zhang , Yan Gao , Yi Wu , Yao Hu , Zhen Wu , Shangyu Xing , Xinyu Dai

Existing vision-language understanding benchmarks largely consist of images of objects in their usual contexts. As a consequence, recent multimodal large language models can perform well with only a shallow visual understanding by relying…

The ability to distinguish whether an image is generated by artificial intelligence (AI) is a crucial ingredient in human intelligence, usually accompanied by a complex and dialectical forensic and reasoning process. However, current fake…

Computer Vision and Pattern Recognition · Computer Science 2024-09-10 Yixuan Li , Xuelin Liu , Xiaoyang Wang , Bu Sung Lee , Shiqi Wang , Anderson Rocha , Weisi Lin

We present GeoGrid-Bench, a benchmark designed to evaluate the ability of foundation models to understand geo-spatial data in the grid structure. Geo-spatial datasets pose distinct challenges due to their dense numerical values, strong…

Computation and Language · Computer Science 2025-05-27 Bowen Jiang , Yangxinyu Xie , Xiaomeng Wang , Jiashu He , Joshua Bergerson , John K Hutchison , Jordan Branham , Camillo J Taylor , Tanwi Mallick

Multimodal Large Language Models (MLLMs) demonstrate impressive problem-solving abilities across a wide range of tasks and domains. However, their capacity for face understanding has not been systematically studied. To address this gap, we…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Kartik Narayan , Vibashan VS , Vishal M. Patel

Large Multimodal Models (LMMs) exhibit major shortfalls when interpreting images and, by some measures, have poorer spatial cognition than small children or animals. Despite this, they attain high scores on many popular visual benchmarks,…

Remote sensing lithology interpretation is fundamental to geological surveys, mineral exploration, and regional geological mapping. Unlike general land-cover recognition, lithology interpretation is a knowledge-intensive task that requires…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Jun Wang , Fengpeng Li , Hang Dong , Tianjin Huang , Wei Han

Multimodal Large Language Models (MLLM) have made significant progress in the field of document analysis. Despite this, existing benchmarks typically focus only on extracting text and simple layout information, neglecting the complex…

Computer Vision and Pattern Recognition · Computer Science 2024-07-04 Lei Chen , Feng Yan , Yujie Zhong , Shaoxiang Chen , Zequn Jie , Lin Ma

While modern visual generation models excel at creating aesthetically pleasing natural images, they struggle with producing or editing structured visuals like charts, diagrams, and mathematical figures, which demand composition planning,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-05 Le Zhuo , Songhao Han , Yuandong Pu , Boxiang Qiu , Sayak Paul , Yue Liao , Yihao Liu , Jie Shao , Xi Chen , Si Liu , Hongsheng Li

Solving expert-level multimodal tasks is a key milestone towards general intelligence. As the capabilities of multimodal large language models (MLLMs) continue to improve, evaluation of such advanced multimodal intelligence becomes…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Yan Yang , Dongxu Li , Haoning Wu , Bei Chen , Liu Liu , Liyuan Pan , Junnan Li

The rapid advancements in Multimodal Large Language Models (MLLMs) have significantly enhanced capabilities in Document Understanding. However, prevailing benchmarks like DocVQA and ChartQA predominantly comprise \textit{scanned or digital}…

Computer Vision and Pattern Recognition · Computer Science 2025-05-28 An-Lan Wang , Jingqun Tang , Liao Lei , Hao Feng , Qi Liu , Xiang Fei , Jinghui Lu , Han Wang , Weiwei Liu , Hao Liu , Yuliang Liu , Xiang Bai , Can Huang

We present TableBank, a new image-based table detection and recognition dataset built with novel weak supervision from Word and Latex documents on the internet. Existing research for image-based table detection and recognition usually…

Computer Vision and Pattern Recognition · Computer Science 2020-07-07 Minghao Li , Lei Cui , Shaohan Huang , Furu Wei , Ming Zhou , Zhoujun Li

We present MaterialFigBench, a benchmark dataset designed to evaluate the ability of multimodal large language models (LLMs) to solve university-level materials science problems that require accurate interpretation of figures. Unlike…

Computation and Language · Computer Science 2026-03-13 Michiko Yoshitake , Yuta Suzuki , Ryo Igarashi , Yoshitaka Ushiku , Keisuke Nagato

Large multimodal models (LMMs) have proven flexible and generalisable across many tasks and fields. Although they have strong potential to aid scientific research, their capabilities in this domain are not well characterised. A key aspect…

Computer Vision and Pattern Recognition · Computer Science 2024-12-06 Jonathan Roberts , Kai Han , Neil Houlsby , Samuel Albanie

Top-down images play an important role in safety-critical settings such as autonomous navigation and aerial surveillance, where they provide holistic spatial information that front-view images cannot capture. Despite this, Vision Language…

Machine Learning · Computer Science 2025-10-02 Kaiyuan Hou , Minghui Zhao , Lilin Xu , Yuang Fan , Xiaofan Jiang
‹ Prev 1 2 3 10 Next ›