中文
相关论文

相关论文: Physics-Based Benchmarking Metrics for Multimodal …

200 篇论文

Multimodal Large Language Models (MLLMs) are gaining increasing popularity in both academia and industry due to their remarkable performance in various applications such as visual question answering, visual perception, understanding, and…

计算与语言 · 计算机科学 2024-09-09 Jian Li , Weiheng Lu , Hao Fei , Meng Luo , Ming Dai , Min Xia , Yizhang Jin , Zhenye Gan , Ding Qi , Chaoyou Fu , Ying Tai , Wankou Yang , Yabiao Wang , Chengjie Wang

Existing multimodal reasoning approaches predominantly follow two paradigms: converting visual inputs into text prior to reasoning, or performing end-to-end reasoning within a unified vision-language representation space. Despite their…

人工智能 · 计算机科学 2026-05-28 Yang Zhang , Xiaoshuai Sun , Rui Zhao , Wujin Sun , Yidong Chen , Jiayi Ji , Qian Chen , Rongrong Ji

Multi-modal information retrieval (MMIR) is a rapidly evolving field, where significant progress, particularly in image-text pairing, has been made through advanced representation learning and cross-modality alignment research. However,…

Knowledge editing techniques have emerged as essential tools for updating the factual knowledge of large language models (LLMs) and multimodal models (LMMs), allowing them to correct outdated or inaccurate information without retraining…

计算与语言 · 计算机科学 2025-03-04 Yuntao Du , Kailin Jiang , Zhi Gao , Chenrui Shi , Zilong Zheng , Siyuan Qi , Qing Li

Multimodal large language models (MLLMs) perform strongly on natural images, yet their ability to understand discrete visual symbols remains unclear. We present a multi-domain benchmark spanning language, culture, mathematics, physics and…

Multilook coherent imaging is a widely used technique in applications such as digital holography, ultrasound imaging, and synthetic aperture radar. A central challenge in these systems is the presence of multiplicative noise, commonly known…

机器学习 · 统计学 2025-05-30 Xi Chen , Soham Jana , Christopher A. Metzler , Arian Maleki , Shirin Jalali

While recent Vision-Language Models (VLMs) have achieved impressive progress, it remains difficult to determine why they succeed or fail on complex reasoning tasks. Traditional benchmarks evaluate what models can answer correctly, not why…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Ieva Bagdonaviciute , Vibhav Vineet

Multi-image Interleaved Reasoning aims to improve Multi-modal Large Language Models (MLLMs) ability to jointly comprehend and reason across multiple images and their associated textual contexts, introducing unique challenges beyond…

计算机视觉与模式识别 · 计算机科学 2025-10-17 Hang Du , Jiayang Zhang , Guoshun Nan , Wendi Deng , Zhenyan Chen , Chenyang Zhang , Wang Xiao , Shan Huang , Yuqi Pan , Tao Qi , Sicong Leng

Most existing CLIP-style medical vision--language pretraining methods rely on global or local alignment with substantial paired data. However, global alignment is easily dominated by non-diagnostic information, while local alignment fails…

计算机视觉与模式识别 · 计算机科学 2026-02-10 Huimin Yan , Liang Bai , Xian Yang , Long Chen

Large multimodal models (LMMs) have demonstrated impressive capabilities in understanding various types of image, including text-rich images. Most existing text-rich image benchmarks are simple extraction-based question answering, and many…

计算机视觉与模式识别 · 计算机科学 2024-08-28 Jian Chen , Ruiyi Zhang , Yufan Zhou , Ryan Rossi , Jiuxiang Gu , Changyou Chen

Physics-aware symbolic simulation of 3D scenes is critical for robotics, embodied AI, and scientific computing, requiring models to understand natural language descriptions of physical phenomena and translate them into executable simulation…

机器人学 · 计算机科学 2026-04-28 Tianyidan Xie , Peiyu Wang , Yuyi Qian , Yuxuan Wang , Rui Ma , Ying Tai , Song Wu , Qian Wang , Lanjun Wang , Zili Yi

The ability to compare objects, scenes, or situations is crucial for effective decision-making and problem-solving in everyday life. For instance, comparing the freshness of apples enables better choices during grocery shopping while…

计算机视觉与模式识别 · 计算机科学 2025-01-14 Jihyung Kil , Zheda Mai , Justin Lee , Zihe Wang , Kerrie Cheng , Lemeng Wang , Ye Liu , Arpita Chowdhury , Wei-Lun Chao

Current metrics for evaluating factuality for abstractive document summarization have achieved high correlations with human judgment, but they do not account for the vision modality and thus are not adequate for vision-and-language…

计算与语言 · 计算机科学 2022-11-07 David Wan , Mohit Bansal

To advance the evaluation of multimodal math reasoning in large multimodal models (LMMs), this paper introduces a novel benchmark, MM-MATH. MM-MATH consists of 5,929 open-ended middle school math problems with visual contexts, with…

计算与语言 · 计算机科学 2024-07-03 Kai Sun , Yushi Bai , Ji Qi , Lei Hou , Juanzi Li

Evaluating vision-language models (VLMs) in scientific domains like mathematics and physics poses unique challenges that go far beyond predicting final answers. These domains demand conceptual understanding, symbolic reasoning, and…

人工智能 · 计算机科学 2025-12-08 Shima Imani , Seungwhan Moon , Adel Ahmadyan , Lu Zhang , Kirmani Ahmed , Babak Damavandi

Multimodal Large Language Models are primarily trained and evaluated on aligned image-text pairs, which leaves their ability to detect and resolve real-world inconsistencies largely unexplored. In open-domain applications visual and textual…

Multi-image spatial reasoning remains challenging for current multimodal large language models (MLLMs). While single-view perception is inherently 2D, reasoning over multiple views requires building a coherent scene understanding across…

计算机视觉与模式识别 · 计算机科学 2026-02-09 Xuejun Zhang , Aditi Tiwari , Zhenhailong Wang , Heng Ji

Multimodal Large Language Models (MLLMs) have emerged to tackle the challenges of Visual Question Answering (VQA), sparking a new research focus on conducting objective evaluations of these models. Existing evaluation methods face…

计算机视觉与模式识别 · 计算机科学 2025-03-20 Qihui Zhang , Munan Ning , Zheyuan Liu , Yanbo Wang , Jiayi Ye , Yue Huang , Shuo Yang , Xiao Chen , Yibing Song , Li Yuan

Medical decision-making requires integrating diverse medical information, from imaging to clinical narratives. These medical modalities are often acquired in a many-to-many manner. However, current medical vision-language pretraining models…

计算机视觉与模式识别 · 计算机科学 2025-10-01 Yuan Gao , Sangwook Kim , Jianzhong You , Chris McIntosh

Spatial consistency is a fundamental property of the visual world and a key requirement for models that aim to understand physical reality. Despite recent advances, multimodal large language models (MLLMs) often struggle to reason about 3D…

计算机视觉与模式识别 · 计算机科学 2026-04-06 Om Khangaonkar , Hadi J. Rad , Hamed Pirsiavash