中文
相关论文

相关论文: MMMG: a Comprehensive and Reliable Evaluation Suit…

200 篇论文

Recently, Multimodal Large Language Models (MLLMs) have achieved exceptional performance across diverse tasks, continually surpassing previous expectations regarding their capabilities. Nevertheless, their proficiency in perceiving emotions…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Daiqing Wu , Dongbao Yang , Sicheng Zhao , Can Ma , Yu Zhou

Multimodal LLMs can accurately perceive numerical content across modalities yet fail to perform exact multi-digit multiplication when the identical underlying arithmetic problem is presented as numerals, number words, images, or in audio…

计算与语言 · 计算机科学 2026-04-21 Samuel G. Balter , Ethan Jerzak , Connor T. Jerzak

The capability to process multiple images is crucial for Large Vision-Language Models (LVLMs) to develop a more thorough and nuanced understanding of a scene. Recent multi-image LVLMs have begun to address this need. However, their…

计算机视觉与模式识别 · 计算机科学 2024-08-07 Fanqing Meng , Jin Wang , Chuanhao Li , Quanfeng Lu , Hao Tian , Jiaqi Liao , Xizhou Zhu , Jifeng Dai , Yu Qiao , Ping Luo , Kaipeng Zhang , Wenqi Shao

Recent advances in deep learning have significantly enhanced generative AI capabilities across text, images, and audio. However, automatically evaluating the quality of these generated outputs presents ongoing challenges. Although numerous…

计算与语言 · 计算机科学 2025-06-13 Tian Lan , Yang-Hao Zhou , Zi-Ao Ma , Fanshu Sun , Rui-Qing Sun , Junyu Luo , Rong-Cheng Tu , Heyan Huang , Chen Xu , Zhijing Wu , Xian-Ling Mao

Reasoning plays a crucial role in advancing Multimodal Large Language Models (MLLMs) toward Artificial General Intelligence. However, existing MLLM benchmarks often fall short in precisely and comprehensively evaluating long-chain reasoning…

Artificial Intelligence (AI) has demonstrated significant potential in healthcare, particularly in disease diagnosis and treatment planning. Recent progress in Medical Large Vision-Language Models (Med-LVLMs) has opened up new possibilities…

机器学习 · 计算机科学 2025-03-04 Peng Xia , Kangyu Zhu , Haoran Li , Tianze Wang , Weijia Shi , Sheng Wang , Linjun Zhang , James Zou , Huaxiu Yao

Interleaved text-and-image generation has been an intriguing research direction, where the models are required to generate both images and text pieces in an arbitrary order. Despite the emerging advancements in interleaved generation, the…

计算机视觉与模式识别 · 计算机科学 2024-10-10 Minqian Liu , Zhiyang Xu , Zihao Lin , Trevor Ashby , Joy Rimchala , Jiaxin Zhang , Lifu Huang

We introduce a new benchmark, ChartMimic, aimed at assessing the visually-grounded code generation capabilities of large multimodal models (LMMs). ChartMimic utilizes information-intensive visual charts and textual instructions as inputs,…

Multimodal Large Language Models (MLLMs) are gaining increasing popularity in both academia and industry due to their remarkable performance in various applications such as visual question answering, visual perception, understanding, and…

计算与语言 · 计算机科学 2024-09-09 Jian Li , Weiheng Lu , Hao Fei , Meng Luo , Ming Dai , Min Xia , Yizhang Jin , Zhenye Gan , Ding Qi , Chaoyou Fu , Ying Tai , Wankou Yang , Yabiao Wang , Chengjie Wang

In different multimodal scenarios, it needs to integrate and utilize information across modalities in a specific way based on the demands of the task. Different integration ways between modalities are referred to as "multimodal…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Yu Miao , Zequn Yang , Yake Wei , Ziheng Chen , Haotian Ni , Haodong Duan , Kai Chen , Di Hu

Multimodal Foundation Models (MMFMs) have demonstrated strong performance in both computer vision and natural language processing tasks. However, their performance diminishes in tasks that require a high degree of integration between these…

计算机视觉与模式识别 · 计算机科学 2025-02-06 Franz Louis Cesista

Medical imaging provides essential visual insights for diagnosis, and multimodal large language models (MLLMs) are increasingly utilized for its analysis due to their strong generalization capabilities; however, the underlying factors…

计算机视觉与模式识别 · 计算机科学 2025-06-03 Zhenyang Cai , Junying Chen , Rongsheng Wang , Weihong Wang , Yonglin Deng , Dingjie Song , Yize Chen , Zixu Zhang , Benyou Wang

Retrieval-augmented generation (RAG) improves large language models (LLMs) by using external knowledge to guide response generation, reducing hallucinations. However, RAG, particularly multi-modal RAG, can introduce new hallucination…

机器学习 · 计算机科学 2025-01-08 Matin Mortaheb , Mohammad A. Amir Khojastepour , Srimat T. Chakradhar , Sennur Ulukus

The ability to comprehend audio--which includes speech, non-speech sounds, and music--is crucial for AI agents to interact effectively with the world. We present MMAU, a novel benchmark designed to evaluate multimodal audio understanding…

音频与语音处理 · 电气工程与系统科学 2024-10-28 S Sakshi , Utkarsh Tyagi , Sonal Kumar , Ashish Seth , Ramaneswaran Selvakumar , Oriol Nieto , Ramani Duraiswami , Sreyan Ghosh , Dinesh Manocha

Existing multimodal retrieval benchmarks primarily focus on evaluating whether models can retrieve and utilize external textual knowledge for question answering. However, there are scenarios where retrieving visual information is either…

计算机视觉与模式识别 · 计算机科学 2025-03-21 Wenbo Hu , Jia-Chen Gu , Zi-Yi Dou , Mohsen Fayyaz , Pan Lu , Kai-Wei Chang , Nanyun Peng

Multimodal Retrieval-Augmented Generation (MRAG) enables Multimodal Large Language Models (MLLMs) to generate responses with external multimodal evidence, and numerous video-based MRAG benchmarks have been proposed to evaluate model…

Context: Machine learning (ML) may enable effective automated test generation. Objective: We characterize emerging research, examining testing practices, researcher goals, ML techniques applied, evaluation, and challenges. Methods: We…

软件工程 · 计算机科学 2023-04-18 Afonso Fontes , Gregory Gay

Adaptive multimodal reasoning has emerged as a promising frontier in Vision-Language Models (VLMs), aiming to dynamically modulate between tool-augmented visual reasoning and text reasoning to enhance both effectiveness and efficiency.…

计算机视觉与模式识别 · 计算机科学 2026-04-09 Xintong Zhang , Xiaowen Zhang , Jingrong Wu , Zhi Gao , Shilin Yan , Zhenxin Diao , Kunpeng Gao , Xuanyan Chen , Yuwei Wu , Yunde Jia , Qing Li

Neural topic models can successfully find coherent and diverse topics in textual data. However, they are limited in dealing with multimodal datasets (e.g., images and text). This paper presents the first systematic and comprehensive…

计算与语言 · 计算机科学 2024-03-27 Felipe González-Pizarro , Giuseppe Carenini

Large Vision-Language Models (LVLMs) are capable of handling diverse data types such as imaging, text, and physiological signals, and can be applied in various fields. In the medical field, LVLMs have a high potential to offer substantial…