English
Related papers

Related papers: Just Noticeable Difference for Large Multimodal Mo…

200 papers

The remarkable advancements in Multimodal Large Language Models (MLLMs) have not rendered them immune to challenges, particularly in the context of handling deceptive information in prompts, thus producing hallucinated responses under such…

Computer Vision and Pattern Recognition · Computer Science 2024-07-24 Yusu Qian , Haotian Zhang , Yinfei Yang , Zhe Gan

Recent advancements in Large Vision-Language Models (LVLMs) have demonstrated remarkable multimodal perception capabilities, garnering significant attention. While numerous evaluation studies have emerged, assessing LVLMs both holistically…

Computer Vision and Pattern Recognition · Computer Science 2026-05-01 Hong-Tao Yu , Yuxin Peng , Serge Belongie , Xiu-Shen Wei

Multimodal Large Language Models (MLLMs) have shown significant potential in medical image analysis. However, their capabilities in interpreting fundus images, a critical skill for ophthalmology, remain under-evaluated. Existing benchmarks…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Qijie Wei , Kaiheng Qian , Xirong Li

Comparing two images in terms of Commonalities and Differences (CaD) is a fundamental human capability that forms the basis of advanced visual reasoning and interpretation. It is essential for the generation of detailed and contextually…

Computer Vision and Pattern Recognition · Computer Science 2024-06-14 Wei Lin , Muhammad Jehanzeb Mirza , Sivan Doveh , Rogerio Feris , Raja Giryes , Sepp Hochreiter , Leonid Karlinsky

Multimodal Large Reasoning Models (MLRMs) have achieved remarkable strides in visual reasoning through test time compute scaling, yet long chain reasoning remains prone to hallucinations. We identify a concerning phenomenon termed the…

Artificial Intelligence · Computer Science 2026-05-29 Zhe Qian , Yanbiao Ma , Zhuohan Ouyang , Zhonghua Wang , Zhongxing Xu , Fei Luo , Xinyu Liu , Zongyuan Ge , Yike Guo , Jungong Han

Multimodal Large Language Models (MLLMs), which couple pre-trained vision encoders and language models, have shown remarkable capabilities. However, their reliance on the ubiquitous Pre-Norm architecture introduces a subtle yet critical…

Computer Vision and Pattern Recognition · Computer Science 2025-12-10 Bozhou Li , Xinda Xue , Sihan Yang , Yang Shi , Xinlong Chen , Yushuo Guan , Yuanxing Zhang , Wentao Zhang

A central question in computational vision is whether human-like visual representations are better explained by discriminative or generative learning. Existing comparisons, however, often confound the learning objective with architecture,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Jorge Chang Ortega , Bastien Le Lan , Thomas Serre , Victor Boutin

Writing is a universal cultural technology that reuses vision for symbolic communication. Humans display striking resilience: we readily recognize words even when characters are fragmented, fused, or partially occluded. This paper…

Computer Vision and Pattern Recognition · Computer Science 2025-12-03 Jie Zhang , Ting Xu , Gelei Deng , Runyi Hu , Han Qiu , Tianwei Zhang , Qing Guo , Ivor Tsang

Leveraging vast training data, multimodal large language models (MLLMs) have demonstrated formidable general visual comprehension capabilities and achieved remarkable performance across various tasks. However, their performance in visual…

Computer Vision and Pattern Recognition · Computer Science 2024-04-11 Chaohu Liu , Kun Yin , Haoyu Cao , Xinghua Jiang , Xin Li , Yinsong Liu , Deqiang Jiang , Xing Sun , Linli Xu

Large Vision-Language Models (VLMs) are increasingly used to evaluate outputs of other models, for image-to-text (I2T) tasks such as visual question answering, and text-to-image (T2I) generation tasks. Despite this growing reliance, the…

Computer Vision and Pattern Recognition · Computer Science 2026-04-24 Mohammed Safi Ur Rahman Khan , Sanjay Suryanarayanan , Tushar Anand , Mitesh M. Khapra

Instruction tuned Large Vision Language Models (LVLMs) have significantly advanced in generalizing across a diverse set of multi-modal tasks, especially for Visual Question Answering (VQA). However, generating detailed responses that are…

Computer Vision and Pattern Recognition · Computer Science 2024-02-13 Anisha Gunjal , Jihan Yin , Erhan Bas

Multimodal Machine Translation (MMT) focuses on enhancing text-only translation with visual features, which has attracted considerable attention from both natural language processing and computer vision communities. Recent advances still…

Computation and Language · Computer Science 2022-11-29 Hongcheng Guo , Jiaheng Liu , Haoyang Huang , Jian Yang , Zhoujun Li , Dongdong Zhang , Zheng Cui , Furu Wei

There has been growing sentiment recently that modern large multimodal models (LMMs) have addressed most of the key challenges related to short video comprehension. As a result, both academia and industry are gradually shifting their…

Computer Vision and Pattern Recognition · Computer Science 2024-10-04 Jianrui Zhang , Mu Cai , Yong Jae Lee

Large Vision-Language Models (LVLMs) have recently played a dominant role in multimodal vision-language learning. Despite the great success, it lacks a holistic evaluation of their efficacy. This paper presents a comprehensive evaluation of…

Computer Vision and Pattern Recognition · Computer Science 2023-06-16 Peng Xu , Wenqi Shao , Kaipeng Zhang , Peng Gao , Shuo Liu , Meng Lei , Fanqing Meng , Siyuan Huang , Yu Qiao , Ping Luo

Recent advances in multimodal large language models (MLLMs) have yielded increasingly powerful models, yet their perceptual capacities remain poorly characterized. In practice, most model families scale language component while reusing…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Tejas Anvekar , Fenil Bardoliya , Pavan K. Turaga , Chitta Baral , Vivek Gupta

Distracted driving continues to be a significant cause of road traffic injuries and fatalities worldwide, even with advancements in driver monitoring technologies. Recent developments in machine learning (ML) and deep learning (DL) have…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Anthony Dontoh , Stephanie Ivey , Logan Sirbaugh , Andrews Danyo , Armstrong Aboah

Despite the remarkable multimodal capabilities of Large Vision-Language Models (LVLMs), discrepancies often occur between visual inputs and textual outputs--a phenomenon we term visual hallucination. This critical reliability gap poses…

Computer Vision and Pattern Recognition · Computer Science 2025-06-25 Tao Huang , Zhekun Liu , Rui Wang , Yang Zhang , Liping Jing

The advent of Large Language Models (LLMs) has significantly reshaped the trajectory of the AI revolution. Nevertheless, these LLMs exhibit a notable limitation, as they are primarily adept at processing textual information. To address this…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Akash Ghosh , Arkadeep Acharya , Sriparna Saha , Vinija Jain , Aman Chadha

Advances in vision language models (VLMs) have enabled the simulation of general human behavior through their reasoning and problem solving capabilities. However, prior research has not investigated such simulation capabilities in the…

Computer Vision and Pattern Recognition · Computer Science 2025-08-18 Rosiana Natalie , Wenqian Xu , Ruei-Che Chang , Rada Mihalcea , Anhong Guo

While language reasoning models excel in many tasks, visual reasoning remains challenging for current large multimodal models (LMMs). As a result, most LMMs default to verbalizing perceptual content into text, a strong limitation for tasks…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 André G. Viveiros , Nuno Gonçalves , Matthias Lindemann , André Martins