中文
相关论文

相关论文: HRScene: How Far Are VLMs from Effective High-Reso…

200 篇论文

Medical image super-resolution (MedSR) is essential for improving diagnostic precision across diverse imaging modalities such as MRI, CT, X-ray, Ultrasound, and Fundus imaging. Despite rapid advances in deep learning, challenges remain in…

计算机视觉与模式识别 · 计算机科学 2026-05-06 Subhash Gurappa , Trivikram Satharasi , Yashas Hariprasad , Sundararaj Sitharama Iyengar

Scene understanding enables intelligent agents to interpret and comprehend their environment. While existing large vision-language models (LVLMs) for scene understanding have primarily focused on indoor household tasks, they face two…

计算机视觉与模式识别 · 计算机科学 2025-07-18 Penglei Sun , Yaoxian Song , Xiangru Zhu , Xiang Liu , Qiang Wang , Yue Liu , Changqun Xia , Tiefeng Li , Yang Yang , Xiaowen Chu

What information is sufficient to learn the full richness of human scene understanding? The distributional hypothesis holds that the statistical co-occurrence of language and images captures the conceptual knowledge underlying visual…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Gillian Rosenberg , Skylar Stadhard , Bruce C. Hansen , Michelle R. Greene

Vision-Language Models (VLMs) are becoming increasingly popular in the medical domain, bridging the gap between medical images and clinical language. Existing VLMs demonstrate an impressive ability to comprehend medical images and text…

After setting the performance benchmarks for image, video, speech and audio processing, deep convolutional networks have been core to the greatest advances in image recognition tasks in recent times. This raises the question of whether…

计算机视觉与模式识别 · 计算机科学 2017-03-17 Grigorios Kalliatakis , Shoaib Ehsan , Maria Fasli , Ales Leonardis , Juergen Gall , Klaus D. McDonald-Maier

Vision language models (VLM) have demonstrated remarkable performance across various downstream tasks. However, understanding fine-grained visual-linguistic concepts, such as attributes and inter-object relationships, remains a significant…

计算机视觉与模式识别 · 计算机科学 2024-04-02 Wujian Peng , Sicheng Xie , Zuyao You , Shiyi Lan , Zuxuan Wu

High-resolution (HR) magnetic resonance imaging (MRI) provides detailed anatomical information that is critical for diagnosis in the clinical application. However, HR MRI typically comes at the cost of long scan time, small spatial…

计算机视觉与模式识别 · 计算机科学 2020-03-10 Yuhua Chen , Anthony G. Christodoulou , Zhengwei Zhou , Feng Shi , Yibin Xie , Debiao Li

Current vision-language models (VLMs) in medicine are primarily designed for categorical question answering (e.g., "Is this normal or abnormal?") or qualitative descriptive tasks. However, clinical decision-making often relies on…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Yongcheng Yao , Yongshuo Zong , Raman Dutt , Yongxin Yang , Sotirios A Tsaftaris , Timothy Hospedales

Camera-based 3D semantic scene completion (SSC) provides dense geometric and semantic perception for autonomous driving. However, images provide limited information making the model susceptible to geometric ambiguity caused by occlusion and…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Meng Wang , Huilong Pi , Ruihui Li , Yunchuan Qin , Zhuo Tang , Kenli Li

High-resolution (HR) 3D magnetic resonance imaging (MRI) can provide detailed anatomical structural information, enabling precise segmentation of regions of interest for various medical image analysis tasks. Due to the high demands of…

图像与视频处理 · 电气工程与系统科学 2024-10-15 Zhiyun Song , Yinjie Zhao , Xiaomin Li , Manman Fei , Xiangyu Zhao , Mengjun Liu , Cunjian Chen , Chung-Hsing Yeh , Qian Wang , Guoyan Zheng , Songtao Ai , Lichi Zhang

Multimodal Large Language Models (MLLMs) have achieved significant advances in integrating visual and linguistic information, yet their ability to reason about complex and real-world scenarios remains limited. The existing benchmarks are…

Synthetic Aperture Radar (SAR) is a crucial remote sensing technology, enabling all-weather, day-and-night observation with strong surface penetration for precise and continuous environmental monitoring and analysis. However, SAR image…

计算机视觉与模式识别 · 计算机科学 2025-04-07 Yimin Wei , Aoran Xiao , Yexian Ren , Yuting Zhu , Hongruixuan Chen , Junshi Xia , Naoto Yokoya

Enabling agents to understand and interact with complex 3D scenes is a fundamental challenge for embodied artificial intelligence systems. While Multimodal Large Language Models (MLLMs) have achieved significant progress in 2D image…

计算机视觉与模式识别 · 计算机科学 2025-09-23 Haoyuan Li , Rui Liu , Hehe Fan , Yi Yang

For collecting high-quality high-resolution (HR) MR image, we propose a novel image reconstruction network named IREM, which is trained on multiple low-resolution (LR) MR images and achieve an arbitrary up-sampling rate for HR image…

图像与视频处理 · 电气工程与系统科学 2021-06-30 Qing Wu , Yuwei Li , Lan Xu , Ruiming Feng , Hongjiang Wei , Qing Yang , Boliang Yu , Xiaozhao Liu , Jingyi Yu , Yuyao Zhang

Recent Multimodal Large Language Models (MLLMs) demonstrate strong high-level visual reasoning on tasks such as visual question answering and image captioning. Yet existing benchmarks largely overlook their ability to capture fine-grained…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Rynaa Grover , Jayant Sravan Tamarapalli , Sahiti Yerramilli , Nilay Pande

The emergence of large-scale large language models, with GPT-4 as a prominent example, has significantly propelled the rapid advancement of artificial general intelligence and sparked the revolution of Artificial Intelligence 2.0. In the…

计算机视觉与模式识别 · 计算机科学 2023-07-31 Yuan Hu , Jianlong Yuan , Congcong Wen , Xiaonan Lu , Xiang Li

The development of Large Vision-Language Models (LVLMs) is striving to catch up with the success of Large Language Models (LLMs), yet it faces more challenges to be resolved. Very recent works enable LVLMs to localize object-level visual…

计算机视觉与模式识别 · 计算机科学 2024-03-20 Zhipeng Huang , Zhizheng Zhang , Zheng-Jun Zha , Yan Lu , Baining Guo

Over the past few years, the advancement of Multimodal Large Language Models (MLLMs) has captured the wide interest of researchers, leading to numerous innovations to enhance MLLMs' comprehension. In this paper, we present AdaptVision, a…

计算机视觉与模式识别 · 计算机科学 2024-09-02 Yonghui Wang , Wengang Zhou , Hao Feng , Houqiang Li

Remote sensing has become a vital tool across sectors such as urban planning, environmental monitoring, and disaster response. While the volume of data generated has increased significantly, traditional vision models are often constrained…

计算机视觉与模式识别 · 计算机科学 2025-10-17 Jia Yun Chua , Argyrios Zolotas , Miguel Arana-Catania

The revolutionary capabilities of large language models (LLMs) have paved the way for multimodal large language models (MLLMs) and fostered diverse applications across various specialized domains. In the remote sensing (RS) field, however,…

计算机视觉与模式识别 · 计算机科学 2024-07-17 Dilxat Muhtar , Zhenshi Li , Feng Gu , Xueliang Zhang , Pengfeng Xiao