English
Related papers

Related papers: LENS: Multi-level Evaluation of Multimodal Reasoni…

200 papers

Reasoning is a fundamental capability for solving complex multi-step problems, particularly in visual contexts where sequential step-wise understanding is essential. Existing approaches lack a comprehensive framework for evaluating visual…

Multimodal large language models (MLLMs), building upon the foundation of powerful large language models (LLMs), have recently demonstrated exceptional capabilities in generating not only texts but also images given interleaved multimodal…

Computer Vision and Pattern Recognition · Computer Science 2023-11-30 Bohao Li , Yuying Ge , Yixiao Ge , Guangzhi Wang , Rui Wang , Ruimao Zhang , Ying Shan

Multimodal large language models (MLLMs) have enabled a wide range of advanced vision-language applications, including fine-grained object recognition and contextual understanding. When querying specific regions or objects in an image,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-17 Mingjie Xu , Jinpeng Chen , Yuzhi Zhao , Jason Chun Lok Li , Yue Qiu , Zekang Du , Mengyang Wu , Pingping Zhang , Kun Li , Hongzheng Yang , Wenao Ma , Jiaheng Wei , Qinbin Li , Kangcheng Liu , Wenqiang Lei

The ability to understand and reason about spatial relationships between objects in images is an important component of visual reasoning. This skill rests on the ability to recognize and localize objects of interest and determine their…

Computation and Language · Computer Science 2024-10-14 Navid Rajabi , Jana Kosecka

Large Vision-Language Models (LVLMs), despite their recent success, are hardly comprehensively tested for their cognitive abilities. Inspired by the prevalent use of the Cookie Theft task in human cognitive tests, we propose a novel…

Artificial Intelligence · Computer Science 2025-02-14 Xiujie Song , Mengyue Wu , Kenny Q. Zhu , Chunhao Zhang , Yanyi Chen

Since the release of ChatGPT, the field of Natural Language Processing has experienced rapid advancements, particularly in Large Language Models (LLMs) and their multimodal counterparts, Large Multimodal Models (LMMs). Despite their…

Computation and Language · Computer Science 2024-08-27 Florian Schneider , Sunayana Sitaram

Multimodal large language models (MLLMs) have achieved impressive progress on vision language benchmarks, yet their capacity for visual cognitive and visuospatial reasoning remains less understood. We introduce "Mind's Eye", a…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Rohit Sinha , Aditya Kanade , Sai Srinivas Kancheti , Vineeth N Balasubramanian , Tanuja Ganu

Multimodal large language models (MLLMs), which integrate language and visual cues for problem-solving, are crucial for advancing artificial general intelligence (AGI). However, current benchmarks for measuring the intelligence of MLLMs…

The sequential structure of videos poses a challenge to the ability of multimodal large language models (MLLMs) to locate multi-frame evidence and conduct multimodal reasoning. However, existing video benchmarks mainly focus on…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Kejian Zhu , Zhuoran Jin , Hongbang Yuan , Jiachun Li , Shangqing Tu , Pengfei Cao , Yubo Chen , Kang Liu , Jun Zhao

Recent large language models (LLMs) have advanced table understanding capabilities but rely on converting tables into text sequences. While multimodal large language models (MLLMs) enable direct visual processing, they face limitations in…

Computation and Language · Computer Science 2025-02-26 Bohao Yang , Yingji Zhang , Dong Liu , André Freitas , Chenghua Lin

Multimodal large language models (MLLMs) have shown strong capabilities across a broad range of benchmarks. However, most existing evaluations focus on passive inference, where models perform step-by-step reasoning under complete…

Computation and Language · Computer Science 2025-10-20 Hongcheng Liu , Pingjie Wang , Yuhao Wang , Siqu Ou , Yanfeng Wang , Yu Wang

Recent multimodal large language models (MLLMs) show strong capabilities in visual-language reasoning, yet their performance on ultra-high-resolution imagery remains largely unexplored. Existing visual question answering (VQA) benchmarks…

Computer Vision and Pattern Recognition · Computer Science 2026-01-14 Siqi Li , Xinyu Cai , Jianbiao Mei , Nianchen Deng , Pinlong Cai , Licheng Wen , Yufan Shen , Xuemeng Yang , Botian Shi , Yong Liu

End-to-end text-image machine translation (TIMT), which directly translates textual content in images across languages, is crucial for real-world multilingual scene understanding. Despite advances in vision-language large models (VLLMs),…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Gengluo Li , Chengquan Zhang , Yupu Liang , Huawen Shen , Yaping Zhang , Pengyuan Lyu , Weinong Wang , Xingyu Wan , Gangyan Zeng , Han Hu , Can Ma , Yu Zhou

This paper reviews the MARS2 2025 Challenge on Multimodal Reasoning. We aim to bring together different approaches in multimodal machine learning and LLMs via a large benchmark. We hope it better allows researchers to follow the…

Computer Vision and Pattern Recognition · Computer Science 2025-09-18 Peng Xu , Shengwu Xiong , Jiajun Zhang , Yaxiong Chen , Bowen Zhou , Chen Change Loy , David A. Clifton , Kyoung Mu Lee , Luc Van Gool , Ruiming He , Ruilin Yao , Xinwei Long , Jirui Huang , Kai Tian , Sa Yang , Yihua Shao , Jin Feng , Yue Zhong , Jiakai Zhou , Cheng Tang , Tianyu Zou , Yifang Zhang , Junming Liang , Guoyou Li , Zhaoxiang Wang , Qiang Zhou , Yichen Zhao , Shili Xiong , Hyeongjin Nam , Jaerin Lee , Jaeyoung Chung , JoonKyu Park , Junghun Oh , Kanggeon Lee , Wooseok Lee , Juneyoung Ro , Turghun Osman , Can Hu , Chaoyang Liao , Cheng Chen , Chengcheng Han , Chenhao Qiu , Chong Peng , Cong Xu , Dailin Li , Feiyu Wang , Feng Gao , Guibo Zhu , Guopeng Tang , Haibo Lu , Han Fang , Han Qi , Hanxiao Wu , Haobo Cheng , Hongbo Sun , Hongyao Chen , Huayong Hu , Hui Li , Jiaheng Ma , Jiang Yu , Jianing Wang , Jie Yang , Jing He , Jinglin Zhou , Jingxuan Li , Josef Kittler , Lihao Zheng , Linnan Zhao , Mengxi Jia , Muyang Yan , Nguyen Thanh Thien , Pu Luo , Qi Li , Shien Song , Shijie Dong , Shuai Shao , Shutao Li , Taofeng Xue , Tianyang Xu , Tianyi Gao , Tingting Li , Wei Zhang , Weiyang Su , Xiaodong Dong , Xiao-Jun Wu , Xiaopeng Zhou , Xin Chen , Xin Wei , Xinyi You , Xudong Kang , Xujie Zhou , Xusheng Liu , Yanan Wang , Yanbin Huang , Yang Liu , Yang Yang , Yanglin Deng , Yashu Kang , Ye Yuan , Yi Wen , Yicen Tian , Yilin Tao , Yin Tang , Yipeng Lin , Yiqing Wang , Yiting Xi , Yongkang Yu , Yumei Li , Yuxin Qin , Yuying Chen , Yuzhe Cen , Zhaofan Zou , Zhaohong Liu , Zhehao Shen , Zhenglin Du , Zhengyang Li , Zhenni Huang , Zhenwei Shao , Zhilong Song , Zhiyong Feng , Zhiyu Wang , Zhou Yu , Ziang Li , Zihan Zhai , Zijian Zhang , Ziyang Peng , Ziyun Xiao , Zongshu Li

Multimodal Large Language Models (MLLMs) have made rapid progress in perception, understanding, and reasoning, yet existing benchmarks fall short in evaluating these abilities under continuous and dynamic real-world video streams. Such…

Computer Vision and Pattern Recognition · Computer Science 2026-01-16 Shuhang Xun , Sicheng Tao , Jungang Li , Yibo Shi , Zhixin Lin , Zhanhui Zhu , Yibo Yan , Hanqian Li , Linghao Zhang , Shikang Wang , Yixin Liu , Hanbo Zhang , Ying Ma , Xuming Hu

Recent multimodal large language models (MLLMs) achieve strong performance on visual reasoning benchmarks, yet it remains unclear to what extent such performance reflects reasoning directly grounded in visual evidence. We introduce…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Longteng Guo , Yifan Wang , Pengkang Huo , Tailai Chen , Yuze Wu , Jing Liu , Xinxin Zhu

Large vision-language models (LVLMs) have significantly improved multimodal reasoning tasks, such as visual question answering and image captioning. These models embed multimodal facts within their parameters, rather than relying on…

Computation and Language · Computer Science 2025-02-18 Shengkang Wang , Hongzhan Lin , Ziyang Luo , Zhen Ye , Guang Chen , Jing Ma

Large Language Models (LLMs) demonstrate enhanced capabilities and reliability by reasoning more, evolving from Chain-of-Thought prompting to product-level solutions like OpenAI o1. Despite various efforts to improve LLM reasoning,…

Computer Vision and Pattern Recognition · Computer Science 2025-05-05 Yuhao Dong , Zuyan Liu , Hai-Long Sun , Jingkang Yang , Winston Hu , Yongming Rao , Ziwei Liu

Vision-Language Models (VLMs) have achieved remarkable progress across tasks such as visual question answering and image captioning. Yet, the extent to which these models perform visual reasoning as opposed to relying on linguistic priors…

Computer Vision and Pattern Recognition · Computer Science 2025-12-25 Brigitta Malagurski Törtei , Yasser Dahou , Ngoc Dung Huynh , Wamiq Reyaz Para , Phúc H. Lê Khac , Ankit Singh , Sofian Chaybouti , Sanath Narayan

Multimodal Large Language Models (MLLMs) have shown promising capabilities in mathematical reasoning within visual contexts across various datasets. However, most existing multimodal math benchmarks are limited to single-visual contexts,…

Artificial Intelligence · Computer Science 2025-08-04 Peijie Wang , Zhong-Zhi Li , Fei Yin , Xin Yang , Dekang Ran , Cheng-Lin Liu