English
Related papers

Related papers: Fairness-Aware Fine-Tuning of Vision-Language Mode…

200 papers

Liver transplantation often faces fairness challenges across subgroups defined by sensitive attributes such as age group, gender, and race/ethnicity. Machine learning models for outcome prediction can introduce additional biases. Therefore,…

Machine Learning · Computer Science 2024-08-28 Can Li , Dejian Lai , Xiaoqian Jiang , Kai Zhang

Medical reports with substantial information can be naturally complementary to medical images for computer vision tasks, and the modality gap between vision and language can be solved by vision-language matching (VLM). However, current…

Image and Video Processing · Electrical Eng. & Systems 2023-05-23 Chen Wenting , Liu Jie , Yuan Yixuan

Large Multimodal Models (LMMs) have shown significant progress in various complex vision tasks with the solid linguistic and reasoning capacity inherited from large language models (LMMs). Low-rank adaptation (LoRA) offers a promising…

Computer Vision and Pattern Recognition · Computer Science 2025-04-04 Liang Mi , Weijun Wang , Wenming Tu , Qingfeng He , Rui Kong , Xinyu Fang , Yazhu Dong , Yikang Zhang , Yunchun Li , Meng Li , Haipeng Dai , Guihai Chen , Yunxin Liu

Existing Medical Large Vision-Language Models (Med-LVLMs), encapsulating extensive medical knowledge, demonstrate excellent capabilities in understanding medical images. However, there remain challenges in visual localization in medical…

Computation and Language · Computer Science 2025-06-03 Yucheng Zhou , Lingran Song , Jianbing Shen

Background/Aims: Standard Automated Perimetry (SAP) is the gold standard to monitor visual field (VF) loss in glaucoma management, but is prone to intra-subject variability. We developed and validated a deep learning (DL) regression model…

Image and Video Processing · Electrical Eng. & Systems 2021-06-08 Ruben Hemelings , Bart Elen , João Barbosa Breda , Erwin Bellon , Matthew B Blaschko , Patrick De Boever , Ingeborg Stalmans

Vision Language Models (VLMs) integrate visual and text modalities to enable multimodal understanding and generation. These models typically combine a Vision Transformer (ViT) as an image encoder and a Large Language Model (LLM) for text…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Krishna Teja Chitty-Venkata , Murali Emani , Venkatram Vishwanath

Deep learning algorithms have demonstrated remarkable efficacy in various medical image analysis (MedIA) applications. However, recent research highlights a performance disparity in these algorithms when applied to specific subgroups, such…

Computer Vision and Pattern Recognition · Computer Science 2024-07-10 Zikang Xu , Jun Li , Qingsong Yao , Han Li , Mingyue Zhao , S. Kevin Zhou

MLLMs MLLMs are beginning to appear in clinical workflows, but their ability to perform complex medical reasoning remains unclear. We present Med-CMR, a fine-grained Medical Complex Multimodal Reasoning benchmark. Med-CMR distinguishes from…

Large Language Models (LLMs) are highly resource-intensive to fine-tune due to their enormous size. While low-rank adaptation is a prominent parameter-efficient fine-tuning approach, it suffers from sensitivity to hyperparameter choices,…

Equity in AI for healthcare is crucial due to its direct impact on human well-being. Despite advancements in 2D medical imaging fairness, the fairness of 3D models remains underexplored, hindered by the small sizes of 3D fairness datasets.…

Computer Vision and Pattern Recognition · Computer Science 2024-04-15 Yan Luo , Muhammad Osama Khan , Yu Tian , Min Shi , Zehao Dou , Tobias Elze , Yi Fang , Mengyu Wang

Machine learning (ML) algorithms play a critical role in decision-making across various domains, such as healthcare, finance, education, and law enforcement. However, concerns about fairness and bias in these systems have raised significant…

Machine Learning · Computer Science 2025-07-25 Ahmed Rashed , Abdelkrim Kallich , Mohamed Eltayeb

Large language models (LLMs) are increasingly integrated into high-stakes decision-making. Inspired by the theory of \emph{inattentional blindness} in human cognition, we investigate whether LLMs, trained on human-preferred corpora that…

Computation and Language · Computer Science 2026-05-20 Yuanqing Cai , Ziyi Huang , Minhao Liu , Lixin Duan , Wen Li , Yanru Zhang

Glaucoma is one of the primary causes of vision loss around the world, necessitating accurate and efficient detection methods. Traditional manual detection approaches have limitations in terms of cost, time, and subjectivity. Recent…

Image and Video Processing · Electrical Eng. & Systems 2023-11-03 Aized Amin Soofi , Fazal-e-Amin

Large Language Models (LLMs) have shown strong performance in text-based healthcare tasks. However, their utility in image-based applications remains unexplored. We investigate the effectiveness of LLMs for medical imaging tasks,…

Computer Vision and Pattern Recognition · Computer Science 2025-09-16 Felicia Liu , Jay J. Yoo , Farzad Khalvati

Training and fine-tuning large language models (LLMs) come with challenges related to memory and computational requirements due to the increasing size of the model weights and the optimizer states. Various techniques have been developed to…

Machine Learning · Computer Science 2025-12-09 Yehonathan Refael , Jonathan Svirsky , Boris Shustin , Wasim Huleihel , Ofir Lindenbaum

Fine-grained glomerular subtyping is central to kidney biopsy interpretation, but clinically valuable labels are scarce and difficult to obtain. Existing computational pathology approaches instead tend to evaluate coarse diseased…

Optical coherence tomography (OCT) based measurements of retinal layer thickness, such as the retinal nerve fibre layer (RNFL) and the ganglion cell with inner plexiform layer (GCIPL) are commonly used for the diagnosis and monitoring of…

Computer Vision and Pattern Recognition · Computer Science 2019-10-25 Stefan Maetschke , Bhavna Antony , Hiroshi Ishikawa , Gadi Wollstein , Joel S. Schuman , Rahil Garnavi

Automated diagnosis based on color fundus photography is essential for large-scale glaucoma screening. However, existing deep learning models are typically data-driven and lack explicit integration of retinal anatomical knowledge, which…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Yuzhuo Zhou , Chi Liu , Sheng Shen , Zongyuan Ge , Fengshi Jing , Shiran Zhang , Yu Jiang , Anli Wang , Wenjian Liu , Feilong Yang , Tianqing Zhu , Xiaotong Han

Vision-Language Models (VLMs) inherit significant social biases from their training data, notably in gender representation. Current fairness interventions often adopt a difference-unaware perspective that enforces uniform treatment across…

Artificial Intelligence · Computer Science 2025-12-02 Yujie Lin , Jiayao Ma , Qingguo Hu , Derek F. Wong , Jinsong Su

Despite significant advancements in adapting Large Language Models (LLMs) for radiology report generation (RRG), clinical adoption remains challenging due to difficulties in accurately mapping pathological and anatomical features to their…

Computer Vision and Pattern Recognition · Computer Science 2025-08-22 Qilong Xing , Zikai Song , Youjia Zhang , Na Feng , Junqing Yu , Wei Yang