English
Related papers

Related papers: MARINER: A 3E-Driven Benchmark for Fine-Grained Pe…

200 papers

MLLMs (Multimodal Large Language Models) have showcased remarkable capabilities, but their performance in high-stakes, domain-specific scenarios like surgical settings, remains largely under-explored. To address this gap, we develop…

Computer Vision and Pattern Recognition · Computer Science 2025-10-03 Gui Wang , Yang Wennuo , Xusen Ma , Zehao Zhong , Zhuoru Wu , Ende Wu , Rong Qu , Wooi Ping Cheah , Jianfeng Ren , Linlin Shen

Vocabulary-free fine-grained image recognition aims to distinguish visually similar categories within a meta-class without a fixed, human-defined label set. Existing solutions for this problem are limited by either the usage of a large and…

Computer Vision and Pattern Recognition · Computer Science 2026-02-27 Dmitry Demidov , Zaigham Zaheer , Zongyan Han , Omkar Thawakar , Rao Anwer

By combining natural language understanding, generation capabilities, and breadth of knowledge of large language models with image perception, recent large vision language models (LVLMs) have shown unprecedented visual reasoning…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Siming Yan , Min Bai , Weifeng Chen , Xiong Zhou , Qixing Huang , Li Erran Li

Recent advancements in Large Vision-Language Models (LVLMs) have demonstrated remarkable multimodal perception capabilities, garnering significant attention. While numerous evaluation studies have emerged, assessing LVLMs both holistically…

Computer Vision and Pattern Recognition · Computer Science 2026-05-01 Hong-Tao Yu , Yuxin Peng , Serge Belongie , Xiu-Shen Wei

Vision-Language Models (VLMs) excel at understanding single images, aided by high-quality instruction datasets. However, multi-image reasoning remains underexplored in the open-source community due to two key challenges: (1) scaling…

Computer Vision and Pattern Recognition · Computer Science 2025-02-18 Andrew Li , Rahul Thapa , Rahul Chalamala , Qingyang Wu , Kezhen Chen , James Zou

Most existing underwater instance segmentation approaches are constrained by close-vocabulary prediction, limiting their ability to recognize novel marine categories. To support evaluation, we introduce \textbf{MARIS} (\underline{Mar}ine…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Bingyu Li , Feiyu Wang , Da Zhang , Zhiyuan Zhao , Junyu Gao , Xuelong Li

Recently, the remarkable success of large language models (LLMs) has achieved a profound impact on the field of artificial intelligence. Numerous advanced works based on LLMs have been proposed and applied in various scenarios. Among them,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Xizhe Xue , Yang Zhou , Dawei Yan , Lijie Tao , Junjie Li , Ying Li , Haokui Zhang , Rong Xiao

Large Multi-modality Models (LMMs) have made significant progress in visual understanding and generation, but they still face challenges in General Visual Editing, particularly in following complex instructions, preserving appearance…

Computer Vision and Pattern Recognition · Computer Science 2025-05-28 Xiangyu Zhao , Peiyuan Zhang , Kexian Tang , Xiaorong Zhu , Hao Li , Wenhao Chai , Zicheng Zhang , Renqiu Xia , Guangtao Zhai , Junchi Yan , Hua Yang , Xue Yang , Haodong Duan

Multimodal Large Language Models (MLLMs) have achieved significant advances in integrating visual and linguistic information, yet their ability to reason about complex and real-world scenarios remains limited. The existing benchmarks are…

The advancement of large language models (LLMs) has significantly broadened the scope of applications in natural language processing, with multi-modal LLMs extending these capabilities to integrate and interpret visual data. However,…

Computer Vision and Pattern Recognition · Computer Science 2024-06-19 Bingchen Zhao , Yongshuo Zong , Letian Zhang , Timothy Hospedales

Recent advances in multimodal large language models (MLLMs) have substantially expanded the capabilities of multimodal retrieval, enabling systems to align and retrieve information across visual and textual modalities. Yet, existing…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Xuan Lu , Kangle Li , Haohang Huang , Rui Meng , Wenjun Zeng , Xiaoyu Shen

While recent advancements in vision-language models have improved video understanding, diagnosing their capacity for deep, narrative comprehension remains a challenge. Existing benchmarks often test short-clip recognition or use…

Computer Vision and Pattern Recognition · Computer Science 2025-09-18 Nisarg A. Shah , Amir Ziai , Chaitanya Ekanadham , Vishal M. Patel

Recent advancements in Large Vision-Language Models (LVLMs) have significantly enhanced their ability to integrate visual and linguistic information, achieving near-human proficiency in tasks like object recognition, captioning, and visual…

Computer Vision and Pattern Recognition · Computer Science 2025-05-14 Zhikai Wang , Jiashuo Sun , Wenqi Zhang , Zhiqiang Hu , Xin Li , Fan Wang , Deli Zhao

Multimodal large language models (MLLMs) struggle with hallucinations, particularly with fine-grained queries, a challenge underrepresented by existing benchmarks that focus on coarse image-related questions. We introduce FIne-grained…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Rui Xiao , Sanghwan Kim , Yongqin Xian , Zeynep Akata , Stephan Alaniz

Any entity in the visual world can be hierarchically grouped based on shared characteristics and mapped to fine-grained sub-categories. While Multi-modal Large Language Models (MLLMs) achieve strong performance on coarse-grained visual…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Hulingxiao He , Zijun Geng , Yuxin Peng

Recent multimodal large language models (MLLMs) show strong capabilities in visual-language reasoning, yet their performance on ultra-high-resolution imagery remains largely unexplored. Existing visual question answering (VQA) benchmarks…

Computer Vision and Pattern Recognition · Computer Science 2026-01-14 Siqi Li , Xinyu Cai , Jianbiao Mei , Nianchen Deng , Pinlong Cai , Licheng Wen , Yufan Shen , Xuemeng Yang , Botian Shi , Yong Liu

Multi-modal Large Language Models (MLLMs) have shown remarkable capabilities across a wide range of vision-language tasks. However, due to the restricted input resolutions, MLLMs face significant challenges in precisely understanding and…

Computer Vision and Pattern Recognition · Computer Science 2025-10-27 Lu Zhang , Jiazuo Yu , Haomiao Xiong , Ping Hu , Yunzhi Zhuge , Huchuan Lu , You He

The frontier of visual reasoning is shifting toward models like OpenAI o3, which can intelligently create and operate tools to transform images for problem-solving, also known as thinking-\textit{with}-images in chain-of-thought. Yet…

Computer Vision and Pattern Recognition · Computer Science 2025-11-07 Ming Li , Jike Zhong , Shitian Zhao , Haoquan Zhang , Shaoheng Lin , Yuxiang Lai , Chen Wei , Konstantinos Psounis , Kaipeng Zhang

Underwater image enhancement has been attracting much attention due to its significance in marine engineering and aquatic robotics. Numerous underwater image enhancement algorithms have been proposed in the last few years. However, these…

Computer Vision and Pattern Recognition · Computer Science 2019-11-27 Chongyi Li , Chunle Guo , Wenqi Ren , Runmin Cong , Junhui Hou , Sam Kwong , Dacheng Tao

Despite rapid advances in vision-language models (VLMs), current benchmarks for multimodal reasoning fall short in three key dimensions. First, they overwhelmingly rely on static images, failing to capture the temporal complexity of…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 Zikui Cai , Andrew Wang , Anirudh Satheesh , Ankit Nakhawa , Hyunwoo Jae , Keenan Powell , Minghui Liu , Neel Jay , Sungbin Oh , Xiyao Wang , Yongyuan Liang , Tom Goldstein , Furong Huang