中文
相关论文

相关论文: Insight Over Sight: Exploring the Vision-Knowledge…

200 篇论文

Large multimodal models (LMMs) excel in adhering to human instructions. However, self-contradictory instructions may arise due to the increasing trend of multimodal interaction and context length, which is challenging for language beginners…

人工智能 · 计算机科学 2024-08-06 Jin Gao , Lei Gan , Yuankai Li , Yixin Ye , Dequan Wang

The rapid evolution of multimodal large language models (MLLMs) has significantly enhanced their real-world applications. However, achieving consistent performance across languages, especially when integrating cultural knowledge, remains a…

计算与语言 · 计算机科学 2025-08-26 Hao Wang , Pinzhi Huang , Jihan Yang , Saining Xie , Daisuke Kawahara

Is basic visual understanding really solved in state-of-the-art VLMs? We present VisualOverload, a slightly different visual question answering (VQA) benchmark comprising 2,720 question-answer pairs, with privately held ground-truth…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Paul Gavrikov , Wei Lin , M. Jehanzeb Mirza , Soumya Jahagirdar , Muhammad Huzaifa , Sivan Doveh , Serena Yeung-Levy , James Glass , Hilde Kuehne

The recent development of Multimodal Large Language Models (MLLMs) has significantly advanced AI's ability to understand visual modalities. However, existing evaluation benchmarks remain limited to single-turn question answering,…

Vision-language models (VLMs) have demonstrated impressive generalization across multimodal tasks, yet most evaluation benchmarks remain Western-centric, leaving open questions about their performance in culturally diverse and multilingual…

计算机视觉与模式识别 · 计算机科学 2025-11-10 Ali Faraz , Akash , Shaharukh Khan , Raja Kolla , Akshat Patidar , Suranjan Goswami , Abhinav Ravi , Chandra Khatri , Shubham Agarwal

The remarkable progress of Multi-modal Large Language Models (MLLMs) has garnered unparalleled attention, due to their superior performance in visual contexts. However, their capabilities in visual math problem-solving remain insufficiently…

计算机视觉与模式识别 · 计算机科学 2024-08-20 Renrui Zhang , Dongzhi Jiang , Yichi Zhang , Haokun Lin , Ziyu Guo , Pengshuo Qiu , Aojun Zhou , Pan Lu , Kai-Wei Chang , Peng Gao , Hongsheng Li

Recently, Multimodal Large Language Models (MLLMs) have made rapid progress, particularly in enhancing their reasoning capabilities. However, existing reasoning benchmarks still primarily assess language-based reasoning, often treating…

计算机视觉与模式识别 · 计算机科学 2025-10-13 Junyan Ye , Dongzhi Jiang , Jun He , Baichuan Zhou , Zilong Huang , Zhiyuan Yan , Hongsheng Li , Conghui He , Weijia Li

Recent multimodal large language models (MLLMs) show strong capabilities in visual-language reasoning, yet their performance on ultra-high-resolution imagery remains largely unexplored. Existing visual question answering (VQA) benchmarks…

计算机视觉与模式识别 · 计算机科学 2026-01-14 Siqi Li , Xinyu Cai , Jianbiao Mei , Nianchen Deng , Pinlong Cai , Licheng Wen , Yufan Shen , Xuemeng Yang , Botian Shi , Yong Liu

The capability to process multiple images is crucial for Large Vision-Language Models (LVLMs) to develop a more thorough and nuanced understanding of a scene. Recent multi-image LVLMs have begun to address this need. However, their…

计算机视觉与模式识别 · 计算机科学 2024-08-07 Fanqing Meng , Jin Wang , Chuanhao Li , Quanfeng Lu , Hao Tian , Jiaqi Liao , Xizhou Zhu , Jifeng Dai , Yu Qiao , Ping Luo , Kaipeng Zhang , Wenqi Shao

Multimodal large language models (MLLMs) have broadened the scope of AI applications. Existing automatic evaluation methodologies for MLLMs are mainly limited in evaluating queries without considering user experiences, inadequately…

The prevalence of vision-threatening eye diseases is a significant global burden, with many cases remaining undiagnosed or diagnosed too late for effective treatment. Large vision-language models (LVLMs) have the potential to assist in…

计算机视觉与模式识别 · 计算机科学 2025-02-06 Zhenyue Qin , Yu Yin , Dylan Campbell , Xuansheng Wu , Ke Zou , Yih-Chung Tham , Ninghao Liu , Xiuzhen Zhang , Qingyu Chen

Visual mathematical reasoning, as a fundamental visual reasoning ability, has received widespread attention from the Large Multimodal Models (LMMs) community. Existing benchmarks, such as MathVista and MathVerse, focus more on the…

We introduce Blink, a new benchmark for multimodal language models (LLMs) that focuses on core visual perception abilities not found in other evaluations. Most of the Blink tasks can be solved by humans "within a blink" (e.g., relative…

计算机视觉与模式识别 · 计算机科学 2024-07-04 Xingyu Fu , Yushi Hu , Bangzheng Li , Yu Feng , Haoyu Wang , Xudong Lin , Dan Roth , Noah A. Smith , Wei-Chiu Ma , Ranjay Krishna

We analyze knowledge-based visual question answering, for which given a question, the models need to ground it into the visual modality and retrieve the relevant knowledge from a given large knowledge base (KB) to be able to answer. Our…

人工智能 · 计算机科学 2024-04-17 Elham J. Barezi , Parisa Kordjamshidi

Large Vision-Language Models (LVLMs) show significant strides in general-purpose multimodal applications such as visual dialogue and embodied navigation. However, existing multimodal evaluation benchmarks cover a limited number of…

Large language models (LLMs) famously exhibit emergent in-context learning (ICL) -- the ability to rapidly adapt to new tasks using few-shot examples provided as a prompt, without updating the model's weights. Built on top of LLMs, vision…

机器学习 · 计算机科学 2025-04-02 Yongshuo Zong , Ondrej Bohdal , Timothy Hospedales

Multimodal Large Language Models (MLLMs) often struggle with fine-grained perception, such as identifying small objects in high-resolution images or detecting key moments in long videos. Existing methods typically rely on complex,…

计算机视觉与模式识别 · 计算机科学 2026-03-16 Sanghwan Kim , Rui Xiao , Stephan Alaniz , Yongqin Xian , Zeynep Akata

Large language models (LLMs) and multimodal large language models (MLLMs) have significantly advanced artificial intelligence. However, visual reasoning, reasoning involving both visual and textual inputs, remains underexplored. Recent…

计算机视觉与模式识别 · 计算机科学 2025-04-18 I-Sheng Fang , Jun-Cheng Chen

As vision-language models (VLMs) are deployed globally, their ability to understand culturally situated knowledge becomes essential. Yet, existing evaluations largely assess static recall or isolated visual grounding, leaving unanswered…

计算机视觉与模式识别 · 计算机科学 2026-01-27 Bryan Chen Zhengyu Tan , Zheng Weihua , Zhengyuan Liu , Nancy F. Chen , Hwaran Lee , Kenny Tsu Wei Choo , Roy Ka-Wei Lee