中文
相关论文

相关论文: SOVABench: A Vehicle Surveillance Action Retrieval…

200 篇论文

Existing vision-language understanding benchmarks largely consist of images of objects in their usual contexts. As a consequence, recent multimodal large language models can perform well with only a shallow visual understanding by relying…

Large Vision-Language Models (LVLMs) increasingly rely on retrieval to answer knowledge-intensive multimodal questions. Existing benchmarks overlook conflicts between visual and textual evidence and the importance of generating deflections…

计算与语言 · 计算机科学 2026-04-15 Nicholas Moratelli , Christopher Davis , Leonardo F. R. Ribeiro , Bill Byrne , Gonzalo Iglesias

Large-scale Vision Language Models (LVLMs) exhibit advanced capabilities in tasks that require visual information, including object detection. These capabilities have promising applications in various industrial domains, such as autonomous…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Haruki Sakajo , Hiroshi Takato , Hiroshi Tsutsui , Komei Soda , Hidetaka Kamigaito , Taro Watanabe

We present Rodent-Bench, a novel benchmark designed to evaluate the ability of Multimodal Large Language Models (MLLMs) to annotate rodent behaviour footage. We evaluate state-of-the-art MLLMs, including Gemini-2.5-Pro, Gemini-2.5-Flash and…

计算机视觉与模式识别 · 计算机科学 2026-02-24 Thomas Heap , Laurence Aitchison , Emma Cahill , Adriana Casado Rodriguez

Pre-trained vision-language models (VLMs) have enabled significant progress in open vocabulary computer vision tasks such as image classification, object detection and image segmentation. Some recent works have focused on extending VLMs to…

计算机视觉与模式识别 · 计算机科学 2025-10-14 Rohit Gupta , Mamshad Nayeem Rizve , Jayakrishnan Unnikrishnan , Ashish Tawari , Son Tran , Mubarak Shah , Benjamin Yao , Trishul Chilimbi

Long videos contain a vast amount of information, making video-text retrieval an essential and challenging task in multimodal learning. However, existing benchmarks suffer from limited video duration, low-quality captions, and coarse…

计算机视觉与模式识别 · 计算机科学 2026-02-05 Qifeng Cai , Hao Liang , Zhaoyang Han , Hejun Dong , Meiyi Qiang , Ruichuan An , Quanqing Xu , Bin Cui , Wentao Zhang

Large Vision-Language Models (LVLMs) show significant strides in general-purpose multimodal applications such as visual dialogue and embodied navigation. However, existing multimodal evaluation benchmarks cover a limited number of…

The rise of Omni-modal Large Language Models (OLLMs), which integrate visual and auditory processing with text, necessitates robust safety evaluations to mitigate harmful outputs. However, no dedicated benchmarks currently exist for OLLMs,…

Multimodal Large Language Models are primarily trained and evaluated on aligned image-text pairs, which leaves their ability to detect and resolve real-world inconsistencies largely unexplored. In open-domain applications visual and textual…

Video foundation models aim to integrate video understanding, generation, editing, and instruction following within a single framework, making them a central direction for next-generation multimodal systems. However, existing evaluation…

计算机视觉与模式识别 · 计算机科学 2026-03-09 Jianhui Wei , Xiaotian Zhang , Yichen Li , Yuan Wang , Yan Zhang , Ziyi Chen , Zhihang Tang , Wei Xu , Zuozhu Liu

Multi-model routing has evolved from an engineering technique into essential infrastructure, yet existing work lacks a systematic, reproducible benchmark for evaluating vision-language models (VLMs). We present VL-RouterBench to assess the…

机器学习 · 计算机科学 2026-03-19 Zhehao Huang , Baijiong Lin , Jingyuan Zhang , Jingying Wang , Yuhang Liu , Ning Lu , Tao Li , Xiaolin Huang

General-purposed embodied agents are designed to understand the users' natural instructions or intentions and act precisely to complete universal tasks. Recently, methods based on foundation models especially Vision-Language-Action models…

机器人学 · 计算机科学 2024-12-25 Shiduo Zhang , Zhe Xu , Peiju Liu , Xiaopeng Yu , Yuan Li , Qinghui Gao , Zhaoye Fei , Zhangyue Yin , Zuxuan Wu , Yu-Gang Jiang , Xipeng Qiu

The emergence of Large Vision-Language Models (LVLMs) marks significant strides towards achieving general artificial intelligence. However, these advancements are accompanied by concerns about biased outputs, a challenge that has yet to be…

计算机视觉与模式识别 · 计算机科学 2026-04-07 Sibo Wang , Xiangkui Cao , Jie Zhang , Zheng Yuan , Shiguang Shan , Xilin Chen , Wen Gao

The generative capabilities of Large Language Models (LLMs) are rapidly expanding from static code to dynamic, interactive visual artifacts. This progress is bottlenecked by a critical evaluation gap: established benchmarks focus on…

Autonomous driving systems must operate reliably in safety-critical scenarios, particularly those involving unusual or complex behavior by Vulnerable Road Users (VRUs). Identifying these edge cases in driving datasets is essential for…

计算机视觉与模式识别 · 计算机科学 2025-08-13 Stefan Englmeier , Max A. Büttner , Katharina Winter , Fabian B. Flohr

Evaluating the robustness of Large Vision-Language Models (LVLMs) is essential for their continued development and responsible deployment in real-world applications. However, existing robustness benchmarks typically focus on hallucination…

计算机视觉与模式识别 · 计算机科学 2026-05-28 Huiyi Chen , Jiawei Peng , Dehai Min , Changchang Sun , Kaijie Chen , Yan Yan , Xu Yang , Lu Cheng

Large language models have demonstrated impressive performance when integrated with vision models even enabling video understanding. However, evaluating video models presents its own unique challenges, for which several benchmarks have been…

计算机视觉与模式识别 · 计算机科学 2025-03-26 Daniel Cores , Michael Dorkenwald , Manuel Mucientes , Cees G. M. Snoek , Yuki M. Asano

The robust safety of Vision-Language Large Models (VLLMs) against joint multilingual and multimodal threats remains severely underexplored. Current benchmarks typically isolate these dimensions, being either multilingual but text-only, or…

计算机视觉与模式识别 · 计算机科学 2026-04-22 Enyi Shi , Pengyang Shao , Yanxin Zhang , Chenhang Cui , Jiayi Lyu , Xiaobo Xia , Fei Shen , Tat-Seng Chua

Despite recent advances in video understanding, the capabilities of Large Video Language Models (LVLMs) to perform video-based causal reasoning remains underexplored, largely due to the absence of relevant and dedicated benchmarks for…

计算机视觉与模式识别 · 计算机科学 2025-05-14 Pritam Sarkar , Ali Etemad

The recent development of Multimodal Large Language Models (MLLMs) has significantly advanced AI's ability to understand visual modalities. However, existing evaluation benchmarks remain limited to single-turn question answering,…