中文
相关论文

相关论文: MetaphorVU: Towards Metaphorical Video Understandi…

200 篇论文

Long videos, ranging from minutes to hours, present significant challenges for current Multi-modal Large Language Models (MLLMs) due to their complex events, diverse scenes, and long-range dependencies. Direct encoding of such videos is…

计算机视觉与模式识别 · 计算机科学 2026-01-12 Zizhong Li , Haopeng Zhang , Jiawei Zhang

Faces and humans are crucial elements in social interaction and are widely included in everyday photos and videos. Therefore, a deep understanding of faces and humans will enable multi-modal assistants to achieve improved response quality…

计算机视觉与模式识别 · 计算机科学 2025-10-24 Lixiong Qin , Shilong Ou , Miaoxuan Zhang , Jiangning Wei , Yuhang Zhang , Xiaoshuai Song , Yuchen Liu , Mei Wang , Weiran Xu

Vision-Language Models (VLMs) have demonstrated remarkable capabilities in interpreting visual layouts and text. However, a significant challenge remains in their ability to interpret robustly and reason over multi-tabular data presented as…

计算机视觉与模式识别 · 计算机科学 2025-06-16 Anshul Singh , Chris Biemann , Jan Strich

Video Large Language Models (VideoLLMs) exhibit various types of hallucinations. Existing research has primarily focused on hallucinations involving the presence of events, objects, and scenes in videos, while largely neglecting event…

计算机视觉与模式识别 · 计算机科学 2026-01-16 Zefan Zhang , Kehua Zhu , Shijie Jiang , Hongyuan Lu , Shengkai Sun , Tian Bai

Recently, multimodal large language models have made significant advancements in video understanding tasks. However, their ability to understand unprocessed long videos is very limited, primarily due to the difficulty in supporting the…

计算机视觉与模式识别 · 计算机科学 2024-06-18 Yiwei Sun , Zhihang Liu , Chuanbin Liu , Bowei Pu , Zhihan Zhang , Hongtao Xie

Large Vision-Language Models (LVLMs) have achieved remarkable performance in many vision-language tasks, yet their capabilities in fine-grained visual understanding remain insufficiently evaluated. Existing benchmarks either contain limited…

计算机视觉与模式识别 · 计算机科学 2024-10-30 Fengbin Zhu , Ziyang Liu , Xiang Yao Ng , Haohui Wu , Wenjie Wang , Fuli Feng , Chao Wang , Huanbo Luan , Tat Seng Chua

Detecting AI-generated images with multimodal large language models (MLLMs) has gained increasing attention, due to their rich world knowledge, common-sense reasoning, and potential for explainability. However, naively applying those MLLMs…

计算机视觉与模式识别 · 计算机科学 2025-10-01 Kaiqing Lin , Zhiyuan Yan , Ruoxin Chen , Junyan Ye , Ke-Yue Zhang , Yue Zhou , Peng Jin , Bin Li , Taiping Yao , Shouhong Ding

Extracting structured information from visual documents (Visual Information Extraction, VIE) is a cornerstone of business automation. While recent Multimodal Large Language Models (MLLMs) have shown promising capabilities, existing…

计算机视觉与模式识别 · 计算机科学 2026-05-22 Yandi Wang , Libin Zhan , Ziwei Huang , Tiancheng Luo , Yuxuan Jiang , Wang Dong , Leilei Gan , Jun Chen

The rapid development of Multi-modality Large Language Models (MLLMs) has significantly influenced various aspects of industry and daily life, showcasing impressive capabilities in visual perception and understanding. However, these models…

计算机视觉与模式识别 · 计算机科学 2024-09-17 Yinan Sun , Zicheng Zhang , Haoning Wu , Xiaohong Liu , Weisi Lin , Guangtao Zhai , Xiongkuo Min

Multimodal large language models (MLLMs) have demonstrated powerful capabilities in general spatial understanding and reasoning. However, their fine-grained spatial understanding and reasoning capabilities in complex urban scenarios have…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Jun Zhang , Jie Feng , Long Chen , Junhui Wang , Zhicheng Liu , Depeng Jin , Yong Li

Multimodal reward models (MRMs) play a crucial role in the training, inference, and evaluation of Large Vision Language Models (LVLMs) by assessing response quality. However, existing benchmarks for evaluating MRMs in the video domain…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Zhihong Zhang , Xiaojian Huang , Jin Xu , Zhuodong Luo , Xinzhi Wang , Jiansheng Wei , Xuejin Chen

Recently, Multimodal Large Language Models (MLLMs) have demonstrated significant potential in complex visual tasks through the integration of Chain-of-Thought (CoT) reasoning. However, in Video Question Answering, extended thinking…

计算机视觉与模式识别 · 计算机科学 2026-03-18 Xiaokun Sun , Yubo Wang , Haoyu Cao , Linli Xu

Generating natural and meaningful responses to communicate with multi-modal human inputs is a fundamental capability of Large Vision-Language Models(LVLMs). While current open-source LVLMs demonstrate promising performance in simplified…

计算机视觉与模式识别 · 计算机科学 2024-10-30 Ziyu Liu , Tao Chu , Yuhang Zang , Xilin Wei , Xiaoyi Dong , Pan Zhang , Zijian Liang , Yuanjun Xiong , Yu Qiao , Dahua Lin , Jiaqi Wang

Video Anomaly Understanding (VAU) is essential for applications such as smart cities, security surveillance, and disaster alert systems, yet remains challenging due to its demand for fine-grained spatio-temporal perception and robust…

计算机视觉与模式识别 · 计算机科学 2025-05-30 Liyun Zhu , Qixiang Chen , Xi Shen , Xiaodong Cun

Visual document understanding (VDU) is a challenging task for large vision language models (LVLMs), requiring the integration of visual perception, text recognition, and reasoning over structured layouts. Although recent LVLMs have shown…

计算与语言 · 计算机科学 2026-04-07 Haruka Kawasaki , Ryota Tanaka , Kyosuke Nishida

Empowered by Large Language Models (LLMs), recent advancements in Video-based LLMs (VideoLLMs) have driven progress in various video understanding tasks. These models encode video representations through pooling or query aggregation over a…

计算机视觉与模式识别 · 计算机科学 2024-07-23 Yuetian Weng , Mingfei Han , Haoyu He , Xiaojun Chang , Bohan Zhuang

Video reasoning has emerged as a critical capability for multimodal large language models (MLLMs), requiring models to move beyond static perception toward coherent understanding of temporal dynamics in complex scenes. Yet existing MLLMs…

计算机视觉与模式识别 · 计算机科学 2025-09-29 Sicheng Tao , Jungang Li , Yibo Yan , Junyan Zhang , Yubo Gao , Hanqian Li , ShuHang Xun , Yuxuan Fan , Hong Chen , Jianxiang He , Xuming Hu

Multimodal Large Language Models (MLLMs) have shown promising progress in understanding and analyzing video content. However, processing long videos remains a significant challenge constrained by LLM's context size. To address this…

There has been growing sentiment recently that modern large multimodal models (LMMs) have addressed most of the key challenges related to short video comprehension. As a result, both academia and industry are gradually shifting their…

计算机视觉与模式识别 · 计算机科学 2024-10-04 Jianrui Zhang , Mu Cai , Yong Jae Lee

Understanding long-form videos, such as movies and TV episodes ranging from tens of minutes to two hours, remains a significant challenge for multi-modal models. Existing benchmarks often fail to test the full range of cognitive skills…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Kirolos Ataallah , Eslam Abdelrahman , Mahmoud Ahmed , Chenhui Gou , Khushbu Pahwa , Jian Ding , Mohamed Elhoseiny
‹ 上一页 1 8 9 10 下一页 ›