中文
相关论文

相关论文: OR-VSKC: Resolving Visual-Semantic Knowledge Confl…

200 篇论文

Vision-language models (VLMs) are essential to Embodied AI, enabling robots to perceive, reason, and act in complex environments. They also serve as the foundation for the recent Vision-Language-Action (VLA) models. Yet most evaluations of…

Multimodal Large Language Models (MLLMs) are widely used in various fields due to their powerful cross-modal comprehension and generation capabilities. However, more modalities bring more vulnerabilities to being utilized for jailbreak…

计算机视觉与模式识别 · 计算机科学 2025-12-09 Shiji Zhao , Shukun Xiong , Yao Huang , Yan Jin , Zhenyu Wu , Jiyang Guan , Ranjie Duan , Jialing Tao , Hui Xue , Xingxing Wei

Fully comprehending scientific papers by machines reflects a high level of Artificial General Intelligence, requiring the ability to reason across fragmented and heterogeneous sources of information, presenting a complex and practically…

计算与语言 · 计算机科学 2025-06-30 Yang Tian , Zheng Lu , Mingqi Gao , Zheng Liu , Bo Zhao

In recent years, Multi-modal Large Language Models (MLLMs) have achieved strong performance in OCR-centric Visual Question Answering (VQA) tasks, illustrating their capability to process heterogeneous data and exhibit adaptability across…

计算机视觉与模式识别 · 计算机科学 2026-02-24 Chen Duan , Zhentao Guo , Pei Fu , Zining Wang , Kai Zhou , Pengfei Yan

We introduce two new benchmarks REST and REST+ (Render-Equivalence Stress Tests) to enable systematic evaluation of cross-modal inconsistency in multimodal large language models (MLLMs). MLLMs are trained to represent vision and language in…

人工智能 · 计算机科学 2026-04-23 Angela van Sprang , Laurens Samson , Ana Lucic , Erman Acar , Sennay Ghebreab , Yuki M. Asano

Misalignment in Large Language Models (LLMs) arises when model behavior diverges from human expectations and fails to simultaneously satisfy safety, value, and cultural dimensions, which must co-occur in real-world settings to solve a…

Cooperative autonomous driving requires traffic scene understanding from both vehicle and infrastructure perspectives. While vision-language models (VLMs) show strong general reasoning capabilities, their performance in safety-critical…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Rui Gan , Junyi Ma , Pei Li , Xingyou Yang , Kai Chen , Sikai Chen , Bin Ran

Medical artificial intelligence (AI) systems, particularly multimodal vision-language models (VLM), often exhibit intersectional biases where models are systematically less confident in diagnosing marginalised patient subgroups. Such bias…

计算机视觉与模式识别 · 计算机科学 2025-12-25 Yupeng Zhang , Adam G. Dunn , Usman Naseem , Jinman Kim

We introduce V-SONAR, a vision-language embedding space extended from the text-only embedding space SONAR (Omnilingual Embeddings Team et al., 2026), which supports 1500 text languages and 177 speech languages. To construct V-SONAR, we…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Yifu Qiu , Paul-Ambroise Duquenne , Holger Schwenk

Open-vocabulary semantic segmentation (OVSS) in remote sensing images is a promising task that employs textual descriptions for identifying undefined land cover categories. Despite notable advances, existing methods typically employ a…

计算机视觉与模式识别 · 计算机科学 2026-04-30 Guanchun Wang , Chenxiao Wu , Xiangrong Zhang , Zelin Peng , Jianxun Lai , Tianyang Zhang , Xu Tang

Vision Large Language Models (VLLMs) represent a significant advancement in artificial intelligence by integrating image-processing capabilities with textual understanding, thereby enhancing user interactions and expanding application…

计算与语言 · 计算机科学 2025-05-09 Madhur Jindal , Saurabh Deshpande

Despite the existing evolution of Multimodal Large Language Models (MLLMs), a non-neglectable limitation remains in their struggle with visual text grounding, especially in text-rich images of documents. Document images, such as scanned…

计算机视觉与模式识别 · 计算机科学 2025-09-25 Ming Li , Ruiyi Zhang , Jian Chen , Chenguang Wang , Jiuxiang Gu , Yufan Zhou , Franck Dernoncourt , Wanrong Zhu , Tianyi Zhou , Tong Sun

Recent advances in vision-language models (VLMs) have accelerated their application to indoor safety hazards assessment. However, existing benchmarks suffer from three fundamental limitations: (1) heavy reliance on synthetic datasets…

计算机视觉与模式识别 · 计算机科学 2026-04-01 Qiucheng Yu , Ruijie Xu , Mingang Chen , Xuequan Lu , Jianfeng Dong , Chaochao Lu , Xin Tan

Recent advances in multimodal large language models (LLMs) have highlighted their potential for medical and surgical applications. However, existing surgical datasets predominantly adopt a Visual Question Answering (VQA) format with…

计算机视觉与模式识别 · 计算机科学 2025-11-27 Tae-Min Choi , Tae Kyeong Jeong , Garam Kim , Jaemin Lee , Yeongyoon Koh , In Cheul Choi , Jae-Ho Chung , Jong Woong Park , Juyoun Park

Open-source Vision-Language Models show immense promise for enterprise applications, yet a critical disconnect exists between academic evaluation and enterprise deployment requirements. Current benchmarks rely heavily on multiple-choice…

计算机视觉与模式识别 · 计算机科学 2025-09-10 Srihari Bandraupalli , Anupam Purwar

Large Language Models (LLMs) require careful safety alignment to prevent malicious outputs. While significant research focuses on mitigating harmful content generation, the enhanced safety often come with the side effect of over-refusal,…

计算与语言 · 计算机科学 2025-06-17 Justin Cui , Wei-Lin Chiang , Ion Stoica , Cho-Jui Hsieh

Vision Language Models (VLMs) are increasingly used for tasks like medical report generation and visual question answering. However, fluent diagnostic text does not guarantee safe visual understanding. In clinical practice, interpretation…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Ufaq Khan , Umair Nawaz , L D M S S Teja , Numaan Saeed , Muhammad Bilal , Yutong Xie , Mohammad Yaqub , Muhammad Haris Khan

Multi-modal large language models (MLLMs), such as GPT-4o, excel at integrating text and visual data but face systematic challenges when interpreting ambiguous or incomplete visual stimuli. This study leverages statistical modeling to…

机器学习 · 计算机科学 2024-12-09 Ching-Yi Wang

Multi-contrast Magnetic Resonance Imaging super-resolution (MC-MRI SR) aims to enhance low-resolution (LR) contrasts leveraging high-resolution (HR) references, shortening acquisition time and improving imaging efficiency while preserving…

计算机视觉与模式识别 · 计算机科学 2025-09-24 Xiaoman Wu , Lubin Gan , Siying Wu , Jing Zhang , Yunwei Ou , Xiaoyan Sun

With the increasing adoption of vision-language models (VLMs) in critical decision-making systems such as healthcare or autonomous driving, the calibration of their uncertainty estimates becomes paramount. Yet, this dimension has been…

计算机视觉与模式识别 · 计算机科学 2026-05-28 Leo Fillioux , Omprakash Chakraborty , Ismail Ben Ayed , Paul-Henry Cournède , Stergios Christodoulidis , Maria Vakalopoulou , Jose Dolz