中文
相关论文

相关论文: Fine-tuning Vision Language Models with Graph-base…

200 篇论文

In recent years, artificial intelligence (AI) systems have come to the forefront. These systems, mostly based on Deep learning (DL), achieve excellent results in areas such as image processing, natural language processing, or speech…

计算机视觉与模式识别 · 计算机科学 2023-04-18 Frantisek Sefcik , Wanda Benesova

Vision-Language Models (VLMs) offer the ability to generate high-level, interpretable descriptions of complex activities from images and videos, making them valuable for situational awareness (SA) applications. In such settings, the focus…

计算机视觉与模式识别 · 计算机科学 2026-01-19 Pavana Pradeep , Krishna Kant , Suya Yu

Background: The lack of explanations for the decisions made by algorithms such as deep learning has hampered their acceptance by the clinical community despite highly accurate results on multiple problems. Recently, attribution methods have…

图像与视频处理 · 电气工程与系统科学 2021-03-26 Amitojdeep Singh , J. Jothi Balaji , Mohammed Abdul Rasheed , Varadharajan Jayakumar , Rajiv Raman , Vasudevan Lakshminarayanan

We present a knowledge augmentation strategy for assessing the diagnostic groups and gait impairment from monocular gait videos. Based on a large-scale pre-trained Vision Language Model (VLM), our model learns and improves visual, textual,…

计算机视觉与模式识别 · 计算机科学 2024-10-16 Diwei Wang , Kun Yuan , Candice Muller , Frédéric Blanc , Nicolas Padoy , Hyewon Seo

Recent advances in machine learning models have greatly increased the performance of automated methods in medical image analysis. However, the internal functioning of such models is largely hidden, which hinders their integration in…

计算机视觉与模式识别 · 计算机科学 2023-07-20 Tatiana Fountoukidou , Raphael Sznitman

Diabetic Retinopathy DR is a severe complication of diabetes. Damaged or abnormal blood vessels can cause loss of vision. The need for massive screening of a large population of diabetic patients has generated an interest in a…

图像与视频处理 · 电气工程与系统科学 2025-01-22 Ameya Uppina , S Navaneetha Krishnan , Talluri Krishna Sai Teja , Nikhil N Iyer , Joe Dhanith P R

The analysis of vision-based deep neural networks (DNNs) is highly desirable but it is very challenging due to the difficulty of expressing formal specifications for vision tasks and the lack of efficient verification procedures. In this…

机器学习 · 计算机科学 2024-04-12 Ravi Mangal , Nina Narodytska , Divya Gopinath , Boyue Caroline Hu , Anirban Roy , Susmit Jha , Corina Pasareanu

Optical coherence tomography angiography (OCTA) is a non-invasive imaging technique widely used to study vascular structures and micro-circulation dynamics in the retina and choroid. OCTA has been widely used in clinics for diagnosing…

图像与视频处理 · 电气工程与系统科学 2025-08-26 Kejie Chen , Guanbing Gao , Xiaochun Yang , Wenbo Wang , Jing Na

Integrating large language models (LLMs) with knowledge graphs derived from domain-specific data represents an important advancement towards more powerful and factual reasoning. As these models grow more capable, it is crucial to enable…

人工智能 · 计算机科学 2024-04-19 Stefan Dernbach , Khushbu Agarwal , Alejandro Zuniga , Michael Henry , Sutanay Choudhury

Diabetic Retinopathy (DR) is an art and science of recording and classifying the retinal images of a diabetic patient. DR classification deals with classifying retinal fundus image into five stages on the basis of severity of diabetes. One…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Nishi Doshi , Urvi Oza , Pankaj Kumar

Recently, diabetic retinopathy (DR) screening utilizing ultra-wide optical coherence tomography angiography (UW-OCTA) has been used in clinical practices to detect signs of early DR. However, developing a deep learning-based DR analysis…

图像与视频处理 · 电气工程与系统科学 2022-10-19 Gitaek Kwon , Eunjin Kim , Sunho Kim , Seongwon Bak , Minsung Kim , Jaeyoung Kim

Despite significant advancements, large multimodal models (LMMs) still struggle to bridge the gap between low-level visual perception -- focusing on shapes, sizes, and layouts -- and high-level language reasoning, such as semantics and…

计算与语言 · 计算机科学 2025-06-13 Zhenhailong Wang , Joy Hsu , Xingyao Wang , Kuan-Hao Huang , Manling Li , Jiajun Wu , Heng Ji

Automatic clinical diagnosis of retinal diseases has emerged as a promising approach to facilitate discovery in areas with limited access to specialists. We propose a novel visual-assisted diagnosis hybrid model based on the support vector…

计算机视觉与模式识别 · 计算机科学 2018-07-05 C. -H. Huck Yang , Jia-Hong Huang , Fangyu Liu , Fang-Yi Chiu , Mengya Gao , Weifeng Lyu , I-Hung Lin M. D. , Jesper Tegner

Diabetic retinopathy (DR) grading plays a critical role in early clinical intervention and vision preservation. Recent explorations predominantly focus on visual lesion feature extraction through data processing and domain decoupling…

人工智能 · 计算机科学 2025-12-01 Chunzheng Zhu , Yangfang Lin , Jialin Shao , Jianxin Lin , Yijun Wang

Visual grounding (VG) is the capability to identify the specific regions in an image associated with a particular text description. In medical imaging, VG enhances interpretability by highlighting relevant pathological features…

With the prevalence of Diabetes, the Diabetes Mellitus Retinopathy (DR) is becoming a major health problem across the world. The long-term medical complications arising due to DR have a significant impact on the patient as well as the…

This paper explores training medical vision-language models (VLMs) -- where the visual and language inputs are embedded into a common space -- with a particular focus on scenarios where training data is limited, as is often the case in…

计算机视觉与模式识别 · 计算机科学 2023-04-03 Rhydian Windsor , Amir Jamaludin , Timor Kadir , Andrew Zisserman

Radiology Report Generation (RRG) through Vision-Language Models (VLMs) promises to reduce documentation burden, improve reporting consistency, and accelerate clinical workflows. However, their clinical adoption remains limited by the lack…

计算机视觉与模式识别 · 计算机科学 2026-02-18 Marco Salmè , Federico Siciliano , Fabrizio Silvestri , Paolo Soda , Rosa Sicilia , Valerio Guarrasi

Abnormalities in retinal fundus images may indicate certain pathologies such as diabetic retinopathy, hypertension, stroke, glaucoma, retinal macular edema, venous occlusion, and atherosclerosis, making the study and analysis of retinal…

图像与视频处理 · 电气工程与系统科学 2024-05-28 Yuzhuo Chen , Zetong Chen , Yuanyuan Liu

Large Language Models (LLMs) and their multimodal variants (LVLMs) hold immense promise for scientific and engineering applications, particularly in processing visual information like scientific diagrams. However, their practical deployment…

计算机视觉与模式识别 · 计算机科学 2025-08-26 Minghao Zhou , Rafael Souza , Yaqian Hu , Luming Che