中文
相关论文

相关论文: HemBLIP: A Vision-Language Model for Interpretable…

200 篇论文

A weakness in the wall of a cerebral artery causing a dilation or ballooning of the blood vessel is known as a cerebral aneurysm. Optimal treatment requires fast and accurate diagnosis of the aneurysm. HemeLB is a fluid dynamics solver for…

Medical image analysis is essential in modern healthcare. Deep learning has redirected research focus toward complex medical multimodal tasks, including report generation and visual question answering. Traditional task-specific models often…

计算机视觉与模式识别 · 计算机科学 2025-09-15 Yiming Shi , Shaoshuai Yang , Xun Zhu , Haoyu Wang , Xiangling Fu , Miao Li , Ji Wu

Building trustworthy clinical AI systems requires not only accurate predictions but also transparent, biologically grounded explanations. We present \texttt{DiagnoLLM}, a hybrid framework that integrates Bayesian deconvolution, eQTL-guided…

人工智能 · 计算机科学 2025-11-18 Bowen Xu , Xinyue Zeng , Jiazhen Hu , Tuo Wang , Adithya Kulkarni

Vision Language Models (VLMs) have demonstrated significant potential in various downstream tasks, including Image/Video Generation, Visual Question Answering, Multimodal Chatbots, and Video Understanding. However, these models often…

计算机视觉与模式识别 · 计算机科学 2025-03-17 Ahmad Mustafa Anis , Hasnain Ali , Saquib Sarfraz

Vision-language foundation models have shown great promise in computational pathology but remain primarily data-driven, lacking explicit integration of medical knowledge. We introduce KEEP (KnowledgE-Enhanced Pathology), a foundation model…

图像与视频处理 · 电气工程与系统科学 2026-01-28 Xiao Zhou , Luoyi Sun , Dexuan He , Wenbin Guan , Ge Wang , Ruifen Wang , Lifeng Wang , Xiaojun Yuan , Xin Sun , Ya Zhang , Kun Sun , Yanfeng Wang , Weidi Xie

LLMs have demonstrated remarkable capabilities in linguistic reasoning and are increasingly adept at vision-language tasks. The integration of image tokens into transformers has enabled direct visual input and output, advancing research…

计算机视觉与模式识别 · 计算机科学 2026-04-06 Jonghun Kim , Sinyoung Ra , Hyunjin Park

Interpretability of deep learning is widely used to evaluate the reliability of medical imaging models and reduce the risks of inaccurate patient recommendations. For models exceeding human performance, e.g. predicting RNA structure from…

Modern Vision-Language Models (VLMs) exhibit unprecedented capabilities in cross-modal semantic understanding between visual and textual modalities. Given the intrinsic need for multi-modal integration in clinical applications, VLMs have…

图像与视频处理 · 电气工程与系统科学 2025-06-24 Haoneng Lin , Cheng Xu , Jing Qin

The lack of large and diverse training data on Computer-Aided Diagnosis (CAD) in breast cancer detection has been one of the concerns that impedes the adoption of the system. Recently, pre-training with large-scale image text datasets via…

图像与视频处理 · 电气工程与系统科学 2024-05-24 Shantanu Ghosh , Clare B. Poynton , Shyam Visweswaran , Kayhan Batmanghelich

Vision-language models (VLMs) have shown considerable potential in digital pathology, yet their effectiveness remains limited for fine-grained, disease-specific classification tasks such as distinguishing between glomerular subtypes. The…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Zhenhao Guo , Rachit Saluja , Tianyuan Yao , Quan Liu , Yuankai Huo , Benjamin Liechty , David J. Pisapia , Kenji Ikemura , Mert R. Sabuncu , Yihe Yang , Ruining Deng

The rapid advancement of artificial intelligence (AI) in healthcare imaging has revolutionized diagnostic medicine and clinical decision-making processes. This work presents an intelligent multimodal framework for medical image analysis…

图像与视频处理 · 电气工程与系统科学 2026-04-20 Samer Al-Hamadani

Vision-language models (VLMs) are increasingly adapted through domain-specific fine-tuning, yet it remains unclear whether this improves reasoning beyond superficial visual cues, particularly in high-stakes domains like medicine. We…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Oliver McLaughlin , Daniel Shubin , Carsten Eickhoff , Ritambhara Singh , William Rudman , Michal Golovanevsky

This paper explores training medical vision-language models (VLMs) -- where the visual and language inputs are embedded into a common space -- with a particular focus on scenarios where training data is limited, as is often the case in…

计算机视觉与模式识别 · 计算机科学 2023-04-03 Rhydian Windsor , Amir Jamaludin , Timor Kadir , Andrew Zisserman

Object categories are typically organized into a multi-granularity taxonomic hierarchy. When classifying categories at different hierarchy levels, traditional uni-modal approaches focus primarily on image features, revealing limitations in…

计算机视觉与模式识别 · 计算机科学 2024-12-03 Peng Xia , Xingtong Yu , Ming Hu , Lie Ju , Zhiyong Wang , Peibo Duan , Zongyuan Ge

Large vision-language models (LVLMs) demonstrate strong performance in dermatology; however, evaluating diagnostic reasoning for rare conditions remains largely unexplored. Existing benchmarks focus on common diseases and assess only final…

计算机视觉与模式识别 · 计算机科学 2026-03-20 Yang Liu , Jiyao Yang , Hongjin Zhao , Xiaoyong Li , Yanzhe Ji , Xingjian Li , Runmin Jiang , Tianyang Wang , Saeed Anwar , Dongwoo Kim , Yue Yao , Zhenyue Qin , Min Xu

Multimodal Large Language Models (MLLMs) have become a powerful tool for integrating visual and textual information. Despite their exceptional performance on visual understanding benchmarks, measuring their ability to reason abstractly…

计算机视觉与模式识别 · 计算机科学 2026-02-26 Nilay Yilmaz , Maitreya Patel , Yiran Lawrence Luo , Tejas Gokhale , Chitta Baral , Suren Jayasuriya , Yezhou Yang

We suggest a low cost, non invasive healthcare system that measures haemoglobin levels in patients and can be used as a preliminary diagnostic test for anaemia. A combination of image processing, machine learning and deep learning…

机器学习 · 计算机科学 2020-12-01 Sarah , S. Sidhartha Narayan , Irfaan Arif , Hrithwik Shalu , Juned Kadiwala

The significant advancements in visual understanding and instruction following from Multimodal Large Language Models (MLLMs) have opened up more possibilities for broader applications in diverse and universal human-centric scenarios.…

计算机视觉与模式识别 · 计算机科学 2024-10-10 Keliang Li , Zaifei Yang , Jiahe Zhao , Hongze Shen , Ruibing Hou , Hong Chang , Shiguang Shan , Xilin Chen

Medical image representations can be learned through medical vision-language contrastive learning (mVLCL) where medical imaging reports are used as weak supervision through image-text alignment. These learned image representations can be…

计算机视觉与模式识别 · 计算机科学 2025-02-10 Mingjian Li , Mingyuan Meng , Michael Fulham , David Dagan Feng , Lei Bi , Jinman Kim

Deep learning enables the modelling of high-resolution histopathology whole-slide images (WSI). Weakly supervised learning of tile-level data is typically applied for tasks where labels only exist on the patient or WSI level (e.g. patient…

图像与视频处理 · 电气工程与系统科学 2025-06-09 Abhinav Sharma , Bojing Liu , Mattias Rantalainen