English
Related papers

Related papers: Knowledge-Driven Vision-Language Model for Plexus …

200 papers

Glomeruli are histological structures of the kidney cortex formed by interwoven blood capillaries, and are responsible for blood filtration. Glomerular lesions impair kidney filtration capability, leading to protein loss and metabolic waste…

Image and Video Processing · Electrical Eng. & Systems 2019-07-02 Paulo Chagas , Luiz Souza , Ikaro Araújo , Nayze Aldeman , Angelo Duarte , Michele Angelo , Washington LC dos-Santos , Luciano Oliveira

Automated 3D CT diagnosis empowers clinicians to make timely, evidence-based decisions by enhancing diagnostic accuracy and workflow efficiency. While multimodal large language models (MLLMs) exhibit promising performance in visual-language…

Computer Vision and Pattern Recognition · Computer Science 2025-06-12 Yanzhao Shi , Xiaodan Zhang , Junzhong Ji , Haoning Jiang , Chengxin Zheng , Yinong Wang , Liangqiong Qu

The conventional pretraining-and-finetuning paradigm, while effective for common diseases with ample data, faces challenges in diagnosing data-scarce occupational diseases like pneumoconiosis. Recently, large language models (LLMs) have…

Image and Video Processing · Electrical Eng. & Systems 2024-07-02 Meiyue Song , Zhihua Yu , Jiaxin Wang , Jiarui Wang , Yuting Lu , Baicun Li , Xiaoxu Wang , Qinghua Huang , Zhijun Li , Nikolaos I. Kanellakis , Jiangfeng Liu , Jing Wang , Binglu Wang , Juntao Yang

Colorectal diseases, including inflammatory conditions and neoplasms, require quick, accurate care to be effectively treated. Traditional diagnostic pipelines require extensive preparation and rely on separate, individual evaluations on…

Computer Vision and Pattern Recognition · Computer Science 2025-09-09 Krithik Ramesh , Ritvik Koneru

While multimodal large language models (MLLMs) have made significant strides in natural image understanding, their ability to perceive and reason over hyperspectral image (HSI) remains underexplored, which is a vital modality in remote…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Xinyu Zhang , Zurong Mai , Qingmei Li , Zjin Liao , Yibin Wen , Yuhang Chen , Xiaoya Fan , Chan Tsz Ho , Bi Tianyuan , Haoyuan Liang , Ruifeng Su , Zihao Qian , Juepeng Zheng , Jianxi Huang , Yutong Lu , Haohuan Fu

Large multimodal language models (LMMs) have achieved significant success in general domains. However, due to the significant differences between medical images and text and general web content, the performance of LMMs in medical scenarios…

Computer Vision and Pattern Recognition · Computer Science 2023-06-23 Weihao Gao , Zhuo Deng , Zhiyuan Niu , Fuju Rong , Chucheng Chen , Zheng Gong , Wenze Zhang , Daimin Xiao , Fang Li , Zhenjie Cao , Zhaoyi Ma , Wenbin Wei , Lan Ma

It is feasible to recognize the presence and seriousness of eye disease by investigating the progressions in retinal biological structure. Fundus examination is a diagnostic procedure to examine the biological structure and anomaly of the…

Image and Video Processing · Electrical Eng. & Systems 2022-07-19 Amit Bhati , Neha Gour , Pritee Khanna , Aparajita Ojha

Oral mucosal diseases such as leukoplakia, oral lichen planus, and recurrent aphthous ulcers exhibit diverse and overlapping visual features, making diagnosis challenging for non-specialists. While vision-language models (VLMs) have shown…

Quantitative Methods · Quantitative Biology 2025-10-17 Jia Zhang , Bodong Du , Yitong Miao , Dongwei Sun , Xiangyong Cao

Different convolutional neural network (CNN) models have been tested for their application in histological image analyses. However, these models are prone to overfitting due to their large parameter capacity, requiring more data or valuable…

Global context information is vital in visual understanding problems, especially in pixel-level semantic segmentation. The mainstream methods adopt the self-attention mechanism to model global context information. However, pixels belonging…

Computer Vision and Pattern Recognition · Computer Science 2020-10-21 Yanwen Chong , Congchong Nie , Yulong Tao , Xiaoshu Chen , Shaoming Pan

In recent years, the incidence of vision-threatening eye diseases has risen dramatically, necessitating scalable and accurate screening solutions. This paper presents a comprehensive study on deep learning architectures for the automated…

Computer Vision and Pattern Recognition · Computer Science 2025-12-12 Mohammad Sadegh Gholizadeh , Amir Arsalan Rezapour

Accurate and interpretable gait analysis plays a crucial role in the early detection of Parkinsons disease (PD),yet most existing approaches remain limited by single-modality inputs, low robustness, and a lack of clinical transparency. This…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Manar Alnaasan , Md Selim Sarowar , Sungho Kim

Accurate diagnosis of skin diseases remains a significant challenge due to the complex and diverse visual features present in dermatoscopic images, often compounded by a lack of interpretability in existing purely visual diagnostic models.…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Kexin Yu , Zihan Xu , Jialei Xie , Carter Adams

Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in understanding common visual elements, largely due to their large-scale datasets and advanced training strategies. However, their effectiveness in medical…

Classifying chest radiographs is a time-consuming and challenging task, even for experienced radiologists. This provides an area for improvement due to the difficulty in precisely distinguishing between conditions such as pleural effusion,…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Maria Efimovich , Jayden Lim , Vedant Mehta , Ethan Poon

Deep learning for medical imaging is hampered by task-specific models that lack generalizability and prognostic capabilities, while existing 'universal' approaches suffer from simplistic conditioning and poor medical semantic understanding.…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Numan Saeed , Tausifa Jan Saleem , Fadillah Maani , Muhammad Ridzuan , Hu Wang , Mohammad Yaqub

Visual Language Models such as CLIP excel in image recognition due to extensive image-text pre-training. However, applying the CLIP inference in zero-shot classification, particularly for medical image diagnosis, faces challenges due to: 1)…

Computer Vision and Pattern Recognition · Computer Science 2025-01-22 Jiaxiang Liu , Tianxiang Hu , Jiawei Du , Ruiyuan Zhang , Joey Tianyi Zhou , Zuozhu Liu

Preventable or undiagnosed visual impairment and blindness affect billion of people worldwide. Automated multi-disease detection models offer great potential to address this problem via clinical decision support in diagnosis. In this work,…

Image and Video Processing · Electrical Eng. & Systems 2021-03-30 Dominik Müller , Iñaki Soto-Rey , Frank Kramer

Pneumonia remains a leading cause of morbidity and mortality worldwide. Chest X-ray (CXR) imaging is a fundamental diagnostic tool, but traditional analysis relies on time-intensive expert evaluation. Recently, deep learning has shown…

Image and Video Processing · Electrical Eng. & Systems 2024-01-05 Sandeep Angara , Nishith Reddy Mannuru , Aashrith Mannuru , Sharath Thirunagaru

Lung disease is common throughout the world. These include chronic obstructive pulmonary disease, pneumonia, asthma, tuberculosis, fibrosis, etc. Timely diagnosis of lung disease is essential. Many image processing and machine learning…

Image and Video Processing · Electrical Eng. & Systems 2021-01-13 Subrato Bharati , Prajoy Podder , M. Rubaiyat Hossain Mondal