中文
相关论文

相关论文: Explain with Visual Keypoints Like a Real Mentor! …

200 篇论文

The recent development of Multimodal Large Language Models (MLLMs) has significantly advanced AI's ability to understand visual modalities. However, existing evaluation benchmarks remain limited to single-turn question answering,…

The sequential structure of videos poses a challenge to the ability of multimodal large language models (MLLMs) to locate multi-frame evidence and conduct multimodal reasoning. However, existing video benchmarks mainly focus on…

计算机视觉与模式识别 · 计算机科学 2025-06-05 Kejian Zhu , Zhuoran Jin , Hongbang Yuan , Jiachun Li , Shangqing Tu , Pengfei Cao , Yubo Chen , Kang Liu , Jun Zhao

Comprehending text-rich visual content is paramount for the practical application of Multimodal Large Language Models (MLLMs), since text-rich scenarios are ubiquitous in the real world, which are characterized by the presence of extensive…

计算机视觉与模式识别 · 计算机科学 2024-04-26 Bohao Li , Yuying Ge , Yi Chen , Yixiao Ge , Ruimao Zhang , Ying Shan

Large language models are increasingly used as educational assistants, yet evaluation of their educational capabilities remains concentrated on question-answering and tutoring tasks. A critical gap exists for multimedia instructional…

计算机与社会 · 计算机科学 2026-04-14 Shuzhen Bi , Mingzi Zhang , Zhuoxuan Li , Xiaolong Wang , Keqian Li , Aimin Zhou

Multi-modal Large Language Models (MLLMs) exhibit impressive problem-solving abilities in various domains, but their visual comprehension and abstract reasoning skills remain under-evaluated. To this end, we present PolyMATH, a challenging…

Multimodal Large Language Models (MLLMs) have gained significant attention recently, showing remarkable potential in artificial general intelligence. However, assessing the utility of MLLMs presents considerable challenges, primarily due to…

计算与语言 · 计算机科学 2024-06-12 Dongping Chen , Ruoxi Chen , Shilin Zhang , Yinuo Liu , Yaochen Wang , Huichi Zhou , Qihui Zhang , Yao Wan , Pan Zhou , Lichao Sun

Despite strong performance on vision-language tasks, Multimodal Large Language Models (MLLMs) struggle with mathematical problem-solving, with both open-source and state-of-the-art models falling short of human performance on visual-math…

计算机视觉与模式识别 · 计算机科学 2025-08-26 William Rudman , Michal Golovanevsky , Amir Bar , Vedant Palit , Yann LeCun , Carsten Eickhoff , Ritambhara Singh

Large Multimodal Models (LMMs) exhibit impressive cross-modal understanding and reasoning abilities, often assessed through multiple-choice questions (MCQs) that include an image, a question, and several options. However, many benchmarks…

计算机视觉与模式识别 · 计算机科学 2025-02-28 Jinsheng Huang , Liang Chen , Taian Guo , Fu Zeng , Yusheng Zhao , Bohan Wu , Ye Yuan , Haozhe Zhao , Zhihui Guo , Yichi Zhang , Jingyang Yuan , Wei Ju , Luchen Liu , Tianyu Liu , Baobao Chang , Ming Zhang

Interleaved multimodal comprehension and generation, enabling models to produce and interpret both images and text in arbitrary sequences, have become a pivotal area in multimodal learning. Despite significant advancements, the evaluation…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Peng Xia , Siwei Han , Shi Qiu , Yiyang Zhou , Zhaoyang Wang , Wenhao Zheng , Zhaorun Chen , Chenhang Cui , Mingyu Ding , Linjie Li , Lijuan Wang , Huaxiu Yao

Visual mathematical reasoning, as a fundamental visual reasoning ability, has received widespread attention from the Large Multimodal Models (LMMs) community. Existing benchmarks, such as MathVista and MathVerse, focus more on the…

Large language models (LLMs) have shown remarkable ability in various language tasks, especially with their emergent in-context learning capability. Extending LLMs to incorporate visual inputs, large vision-language models (LVLMs) have…

机器学习 · 计算机科学 2025-10-13 Aneesh Komanduri , Karuna Bhaila , Xintao Wu

Unified multimodal models aim to jointly enable visual understanding and generation, yet current benchmarks rarely examine their true integration. Existing evaluations either treat the two abilities in isolation or overlook tasks that…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Kai Zou , Ziqi Huang , Yuhao Dong , Shulin Tian , Dian Zheng , Hongbo Liu , Jingwen He , Bin Liu , Yu Qiao , Ziwei Liu

Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated remarkable progress in visual understanding. This impressive leap raises a compelling question: how can language models, initially trained solely on…

计算机视觉与模式识别 · 计算机科学 2025-11-12 Jing Bi , Junjia Guo , Yunlong Tang , Lianggong Bruce Wen , Zhang Liu , Chenliang Xu

Large Language Models (LLMs) are increasingly being used in education, yet their correctness alone does not capture the quality, reliability, or pedagogical validity of their problem-solving behavior, especially in mathematics, where…

计算机与社会 · 计算机科学 2025-10-22 Sagnik Dakshit , Sushmita Sinha Roy

Multimodal large language models (MLLMs) perform well on many vision-language tasks but often struggle with vision-centric problems that require fine-grained visual reasoning. Recent evidence suggests that this limitation arises not from…

计算机视觉与模式识别 · 计算机科学 2026-04-15 Sophia Sirko-Galouchenko , Monika Wysoczanska , Andrei Bursuc , Nicolas Thome , Spyros Gidaris

We propose MM-Vet, an evaluation benchmark that examines large multimodal models (LMMs) on complicated multimodal tasks. Recent LMMs have shown various intriguing abilities, such as solving math problems written on the blackboard, reasoning…

人工智能 · 计算机科学 2024-12-03 Weihao Yu , Zhengyuan Yang , Linjie Li , Jianfeng Wang , Kevin Lin , Zicheng Liu , Xinchao Wang , Lijuan Wang

Evaluating the performance of Multi-modal Large Language Models (MLLMs), integrating both point cloud and language, presents significant challenges. The lack of a comprehensive assessment hampers determining whether these models truly…

计算机视觉与模式识别 · 计算机科学 2024-04-24 Junjie Zhang , Tianci Hu , Xiaoshui Huang , Yongshun Gong , Dan Zeng

Multimodal Large Language Models (MLLMs) are evolving from passive observers into active agents, solving problems through Visual Expansion (invoking visual tools) and Knowledge Expansion (open-web search). However, existing evaluations fall…

This paper reviews the MARS2 2025 Challenge on Multimodal Reasoning. We aim to bring together different approaches in multimodal machine learning and LLMs via a large benchmark. We hope it better allows researchers to follow the…

计算机视觉与模式识别 · 计算机科学 2025-09-18 Peng Xu , Shengwu Xiong , Jiajun Zhang , Yaxiong Chen , Bowen Zhou , Chen Change Loy , David A. Clifton , Kyoung Mu Lee , Luc Van Gool , Ruiming He , Ruilin Yao , Xinwei Long , Jirui Huang , Kai Tian , Sa Yang , Yihua Shao , Jin Feng , Yue Zhong , Jiakai Zhou , Cheng Tang , Tianyu Zou , Yifang Zhang , Junming Liang , Guoyou Li , Zhaoxiang Wang , Qiang Zhou , Yichen Zhao , Shili Xiong , Hyeongjin Nam , Jaerin Lee , Jaeyoung Chung , JoonKyu Park , Junghun Oh , Kanggeon Lee , Wooseok Lee , Juneyoung Ro , Turghun Osman , Can Hu , Chaoyang Liao , Cheng Chen , Chengcheng Han , Chenhao Qiu , Chong Peng , Cong Xu , Dailin Li , Feiyu Wang , Feng Gao , Guibo Zhu , Guopeng Tang , Haibo Lu , Han Fang , Han Qi , Hanxiao Wu , Haobo Cheng , Hongbo Sun , Hongyao Chen , Huayong Hu , Hui Li , Jiaheng Ma , Jiang Yu , Jianing Wang , Jie Yang , Jing He , Jinglin Zhou , Jingxuan Li , Josef Kittler , Lihao Zheng , Linnan Zhao , Mengxi Jia , Muyang Yan , Nguyen Thanh Thien , Pu Luo , Qi Li , Shien Song , Shijie Dong , Shuai Shao , Shutao Li , Taofeng Xue , Tianyang Xu , Tianyi Gao , Tingting Li , Wei Zhang , Weiyang Su , Xiaodong Dong , Xiao-Jun Wu , Xiaopeng Zhou , Xin Chen , Xin Wei , Xinyi You , Xudong Kang , Xujie Zhou , Xusheng Liu , Yanan Wang , Yanbin Huang , Yang Liu , Yang Yang , Yanglin Deng , Yashu Kang , Ye Yuan , Yi Wen , Yicen Tian , Yilin Tao , Yin Tang , Yipeng Lin , Yiqing Wang , Yiting Xi , Yongkang Yu , Yumei Li , Yuxin Qin , Yuying Chen , Yuzhe Cen , Zhaofan Zou , Zhaohong Liu , Zhehao Shen , Zhenglin Du , Zhengyang Li , Zhenni Huang , Zhenwei Shao , Zhilong Song , Zhiyong Feng , Zhiyu Wang , Zhou Yu , Ziang Li , Zihan Zhai , Zijian Zhang , Ziyang Peng , Ziyun Xiao , Zongshu Li

Multimodal large language models (MLLMs) have broadened the scope of AI applications. Existing automatic evaluation methodologies for MLLMs are mainly limited in evaluating queries without considering user experiences, inadequately…