中文
相关论文

相关论文: DynamicVL: Benchmarking Multimodal Large Language …

200 篇论文

Multimodal Large Language Models (MLLMs) have demonstrated remarkable effectiveness in various general-domain scenarios, such as visual question answering and image captioning. Recently, researchers have increasingly focused on empowering…

计算机视觉与模式识别 · 计算机科学 2025-11-04 Yan Shu , Chi Liu , Robin Chen , Derek Li , Bryan Dai

Multimodal language analysis is a rapidly evolving field that leverages multiple modalities to enhance the understanding of high-level semantics underlying human conversational utterances. Despite its significance, little research has…

计算与语言 · 计算机科学 2025-04-25 Hanlei Zhang , Zhuohang Li , Yeshuang Zhu , Hua Xu , Peiwu Wang , Haige Zhu , Jie Zhou , Jinchao Zhang

Text-rich images, where text serves as the central visual element guiding the overall understanding, are prevalent in real-world applications, such as presentation slides, scanned documents, and webpage snapshots. Tasks involving multiple…

计算机视觉与模式识别 · 计算机科学 2025-06-09 Mengzhao Jia , Wenhao Yu , Kaixin Ma , Tianqing Fang , Zhihan Zhang , Siru Ouyang , Hongming Zhang , Dong Yu , Meng Jiang

Medical image analysis is essential in modern healthcare. Deep learning has redirected research focus toward complex medical multimodal tasks, including report generation and visual question answering. Traditional task-specific models often…

计算机视觉与模式识别 · 计算机科学 2025-09-15 Yiming Shi , Shaoshuai Yang , Xun Zhu , Haoyu Wang , Xiangling Fu , Miao Li , Ji Wu

The evaluation of Long Video Understanding (LVU) performance poses an important but challenging research problem. Despite previous efforts, the existing video understanding benchmarks are severely constrained by several issues, especially…

计算机视觉与模式识别 · 计算机科学 2025-01-03 Junjie Zhou , Yan Shu , Bo Zhao , Boya Wu , Zhengyang Liang , Shitao Xiao , Minghao Qin , Xi Yang , Yongping Xiong , Bo Zhang , Tiejun Huang , Zheng Liu

Over the past few years, the advancement of Multimodal Large Language Models (MLLMs) has captured the wide interest of researchers, leading to numerous innovations to enhance MLLMs' comprehension. In this paper, we present AdaptVision, a…

计算机视觉与模式识别 · 计算机科学 2024-09-02 Yonghui Wang , Wengang Zhou , Hao Feng , Houqiang Li

Vision-and-Language Navigation (VLN) is a core task where embodied agents leverage their spatial mobility to navigate in 3D environments toward designated destinations based on natural language instructions. Recently, video-language large…

计算机视觉与模式识别 · 计算机科学 2025-05-19 Zihan Wang , Seungjun Lee , Gim Hee Lee

Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in various tasks. However, effectively evaluating these MLLMs on face perception remains largely unexplored. To address this gap, we introduce FaceBench, a…

计算机视觉与模式识别 · 计算机科学 2025-03-28 Xiaoqin Wang , Xusen Ma , Xianxu Hou , Meidan Ding , Yudong Li , Junliang Chen , Wenting Chen , Xiaoyang Peng , Linlin Shen

Multimodal large language models (MLLMs) have achieved impressive performance on visual perception and reasoning tasks with RGB imagery, yet they remain fragile under common degradations, such as fog, blur, or low-light conditions. Infrared…

计算机视觉与模式识别 · 计算机科学 2026-04-22 Abrar Majeedi , Zhiyuan Ruan , Ziyi Zhao , Hongcheng Wang , Jianglin Lu , Yin Li

Recent advances in multimodal large language models (MLLMs) have demonstrated substantial potential in video understanding. However, existing benchmarks fail to comprehensively evaluate synergistic reasoning capabilities across audio and…

Recent progress in multimodal large language models has markedly enhanced the understanding of short videos (typically under one minute), and several evaluation datasets have emerged accordingly. However, these advancements fall short of…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Weihan Wang , Zehai He , Wenyi Hong , Yean Cheng , Xiaohan Zhang , Ji Qi , Xiaotao Gu , Shiyu Huang , Bin Xu , Yuxiao Dong , Ming Ding , Jie Tang

Large Vision and Language Models (LVLMs) have shown strong performance across various vision-language tasks in natural image domains. However, their application to remote sensing (RS) remains underexplored due to significant domain…

计算机视觉与模式识别 · 计算机科学 2025-06-30 Sungjune Park , Yeongyun Kim , Se Yeon Kim , Yong Man Ro

Image-text matching (ITM) aims to address the fundamental challenge of aligning visual and textual modalities, which inherently differ in their representations, continuous, high-dimensional image features vs. discrete, structured text. We…

多媒体 · 计算机科学 2025-07-14 Junyu Chen , Yihua Gao , Mingyong Li

Multi-modal Large Language Models (MLLMs) exhibit impressive capabilities in 2D tasks, yet encounter challenges in discerning the spatial positions, interrelations, and causal logic in scenes when transitioning from 2D to 3D…

计算机视觉与模式识别 · 计算机科学 2025-01-15 Haomiao Xiong , Yunzhi Zhuge , Jiawen Zhu , Lu Zhang , Huchuan Lu

Large multimodal models (LMMs) have recently gained attention due to their effectiveness to understand and generate descriptions of visual content. Most existing LMMs are in English language. While few recent works explore multilingual…

Multimodal Large Language Models (MLLMs) show impressive vision-language benchmark performance, yet growing concerns about data contamination (test set exposure during training) risk masking true generalization. This concern extends to…

人工智能 · 计算机科学 2025-06-10 Ming Liu , Wensheng Zhang

Multi-modal Large Language Models (MLLMs) have demonstrated remarkable capabilities in executing instructions for a variety of single-image tasks. Despite this progress, significant challenges remain in modeling long image sequences. In…

计算机视觉与模式识别 · 计算机科学 2024-08-14 Jiabo Ye , Haiyang Xu , Haowei Liu , Anwen Hu , Ming Yan , Qi Qian , Ji Zhang , Fei Huang , Jingren Zhou

Large multimodal models (LMMs) show strong visual-linguistic reasoning but their capacity for spatial decision-making and action remains unclear. In this work, we investigate whether LMMs can achieve embodied spatial action like human…

Despite significant progress, existing research on Multimodal Large Language Models (MLLMs) mainly focuses on general visual understanding, overlooking the ability to integrate textual context associated with objects for a more…

计算机视觉与模式识别 · 计算机科学 2025-09-01 Hongliang Wei , Xianqi Zhang , Xingtao Wang , Xiaopeng Fan , Debin Zhao

The popularity of multimodal large language models (MLLMs) has triggered a recent surge in research efforts dedicated to evaluating these models. Nevertheless, existing evaluation studies of MLLMs primarily focus on the comprehension and…

计算与语言 · 计算机科学 2023-10-16 Xiaocui Yang , Wenfang Wu , Shi Feng , Ming Wang , Daling Wang , Yang Li , Qi Sun , Yifei Zhang , Xiaoming Fu , Soujanya Poria