中文
相关论文

相关论文: Diagnostic Accuracy of Open-Source Vision-Language…

200 篇论文

Automatic dietary assessment based on food images remains a challenge, requiring precise food detection, segmentation, and classification. Vision-Language Models (VLMs) offer new possibilities by integrating visual and textual reasoning. In…

Vision language models (VLMs) have shown significant promise in visual grounding for images as well as videos. In medical imaging research, VLMs represent a bridge between object detection and segmentation, and report understanding and…

计算机视觉与模式识别 · 计算机科学 2026-03-23 Andrew Seohwan Yu , Mohsen Hariri , Kunio Nakamura , Mingrui Yang , Xiaojuan Li , Vipin Chaudhary

The application of artificial intelligence (AI) in medical imaging has revolutionized diagnostic practices, enabling advanced analysis and interpretation of radiological data. This study presents a comprehensive evaluation of…

图像与视频处理 · 电气工程与系统科学 2025-07-22 Zhijin He , Alan B. McMillan

Purpose To evaluate the reasoning capabilities of large language models (LLMs) in performing root cause analysis (RCA) of radiation oncology incidents using narrative reports from the Radiation Oncology Incident Learning System (RO-ILS),…

Vision-Language Models (VLMs) in visual question answering (VQA) offer a unique opportunity to enhance intra-operative decision-making, promote intuitive interactions, and significantly advancing surgical education. However, the development…

计算机视觉与模式识别 · 计算机科学 2025-02-21 Runlong He , Danyal Z. Khan , Evangelos B. Mazomenos , Hani J. Marcus , Danail Stoyanov , Matthew J. Clarkson , Mobarakol Islam

Large language models (LLMs) have demonstrated immense capabilities in understanding textual data and are increasingly being adopted to help researchers accelerate scientific discovery through knowledge extraction (information retrieval),…

计算机视觉与模式识别 · 计算机科学 2025-05-30 Robinson Umeike , Neil Getty , Fangfang Xia , Rick Stevens

Although Vision Language Models (VLMs) have seen tremendous progress across all kinds of use cases, they still fall behind in answering questions regard-ing diagrams compared to photos. Although progress has been made in the area of bar…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Artem Naboichenko , René Peinl

Vision-language foundation models (VLMs) have shown impressive performance in guiding image generation through text, with emerging applications in medical imaging. In this work, we are the first to investigate the question: 'Can fine-tuned…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Amar Kumar , Anita Kriz , Barak Pertzov , Tal Arbel

Multimodal language models (MLMs) show promise for clinical decision support and diagnostic reasoning, raising the prospect of end-to-end automated medical image interpretation. However, clinicians are highly selective in adopting AI tools;…

Large language models (LLMs) constitute a breakthrough state-of-the-art Artificial Intelligence technology which is rapidly evolving and promises to aid in medical diagnosis. However, the correctness and the accuracy of their returns has…

计算与语言 · 计算机科学 2024-02-07 Dimitrios P. Panagoulias , Maria Virvou , George A. Tsihrintzis

The Vision-Language Foundation model is increasingly investigated in the fields of computer vision and natural language processing, yet its exploration in ophthalmology and broader medical applications remains limited. The challenge is the…

计算机视觉与模式识别 · 计算机科学 2024-08-20 Jiawei Du , Jia Guo , Weihang Zhang , Shengzhu Yang , Hanruo Liu , Huiqi Li , Ningli Wang

Coronary angiography (CAG) is the gold-standard imaging modality for evaluating coronary artery disease, but its interpretation and subsequent treatment planning rely heavily on expert cardiologists. To enable AI-based decision support, we…

Recent advances in open-source vision-language models (VLMs) offer new opportunities for understanding complex and subjective multimodal phenomena such as sarcasm. In this work, we evaluate seven state-of-the-art VLMs - BLIP2, InstructBLIP,…

机器学习 · 计算机科学 2025-10-15 Saroj Basnet , Shafkat Farabi , Tharindu Ranasinghe , Diptesh Kanoji , Marcos Zampieri

Foundation models are large-scale versatile systems trained on vast quantities of diverse data to learn generalizable representations. Their adaptability with minimal fine-tuning makes them particularly promising for medical imaging, where…

Multimodal scientific reasoning remains a significant challenge for large language models (LLMs), particularly in chemistry, where problem-solving relies on symbolic diagrams, molecular structures, and structured visual data. Here, we…

计算与语言 · 计算机科学 2025-12-18 Yiming Cui , Xin Yao , Yuxuan Qin , Xin Li , Shijin Wang , Guoping Hu

Medical vision--language models (VLMs) are usually evaluated on intact image--question pairs, but trustworthy clinical use requires a stronger property: a model must recognise when the evidential basis for an answer has failed. We study…

计算机视觉与模式识别 · 计算机科学 2026-05-25 Hanqi Jiang , Junhao Chen , Mingyu Kang , Hyeokjae Kwon , Yi Pan , Lifeng Chen , Weihang You , Haozhen Gong , Ruiyu Yan , Jinglei Lv , Lin Zhao , Hui Ren , Quanzheng Li , Tianming Liu , Xiang Li

The development of multi-label deep learning models for retinal disease classification is often hindered by the scarcity of large, expertly annotated clinical datasets due to patient privacy concerns and high costs. The recent release of…

计算机视觉与模式识别 · 计算机科学 2025-08-25 Jerry Cao-Xue , Tien Comlekoglu , Keyi Xue , Guanliang Wang , Jiang Li , Gordon Laurie

Foundation models and vision-language pre-training have significantly advanced Vision-Language Models (VLMs), enabling multimodal processing of visual and linguistic data. However, their application in domain-specific agricultural tasks,…

计算机视觉与模式识别 · 计算机科学 2026-02-23 Khang Nguyen Quoc , Phuong D. Dao , Luyl-Da Quach

Objective: Ophthalmological pathologies such as glaucoma, diabetic retinopathy and age-related macular degeneration are major causes of blindness and vision impairment. There is a need for novel decision support tools that can simplify and…

图像与视频处理 · 电气工程与系统科学 2023-06-08 Or Abramovich , Hadas Pizem , Jan Van Eijgen , Ilan Oren , Joshua Melamed , Ingeborg Stalmans , Eytan Z. Blumenthal , Joachim A. Behar

Multimodal Large Language Models (MLLMs) have shown significant potential in medical image analysis. However, their capabilities in interpreting fundus images, a critical skill for ophthalmology, remain under-evaluated. Existing benchmarks…

计算机视觉与模式识别 · 计算机科学 2025-03-04 Qijie Wei , Kaiheng Qian , Xirong Li