English
Related papers

Related papers: MDF-MLLM: Deep Fusion Through Cross-Modal Feature …

200 papers

Red-lesions, microaneurysms (MAs) and hemorrhages (HMs), are the early signs of diabetic retinopathy (DR). The automatic detection of MAs and HMs on retinal fundus images is a challenging task. Most of the existing methods detect either…

Image and Video Processing · Electrical Eng. & Systems 2025-05-02 Norah Asiri , Muhammad Hussain , Fadwa Al Adel

In recent years, the rapid evolution of large vision-language models (LVLMs) has driven a paradigm shift in multimodal fake news detection (MFND), transforming it from traditional feature-engineering approaches to unified, end-to-end…

Artificial Intelligence · Computer Science 2026-01-23 Wei Ai , Yilong Tan , Yuntao Shou , Tao Meng , Haowen Chen , Zhixiong He , Keqin Li

Recent advancements in Large Multimodal Models (LMMs) have attracted interest in their generalization capability with only a few samples in the prompt. This progress is particularly relevant to the medical domain, where the quality and…

Computation and Language · Computer Science 2024-05-06 Seonhee Cho , Choonghan Kim , Jiho Lee , Chetan Chilkunda , Sujin Choi , Joo Heung Yoon

Accurate segmentation of skin lesion from dermoscopic images is a crucial part of computer-aided diagnosis of melanoma. It is challenging due to the fact that dermoscopic images from different patients have non-negligible lesion variation,…

Computer Vision and Pattern Recognition · Computer Science 2020-02-21 Xiaohong Wang , Xudong Jiang , Henghui Ding , Jun Liu

Diabetic retinopathy is the most important complication of diabetes. Early diagnosis of retinal lesions helps to avoid visual loss or blindness. Due to high-resolution and small-size lesion regions, applying existing methods, such as…

Computer Vision and Pattern Recognition · Computer Science 2019-01-21 Zizheng Yan , Xiaoguang Han , Changmiao Wang , Yuda Qiu , Zixiang Xiong , Shuguang Cui

In recent years, deep learning has shown promise in predicting hypertension (HTN) from fundus images. However, most prior research has primarily focused on analyzing a single type of data, which may not capture the full complexity of HTN…

Image and Video Processing · Electrical Eng. & Systems 2024-03-26 Mohammed Baharoon , Hessa Almatar , Reema Alduhayan , Tariq Aldebasi , Badr Alahmadi , Yahya Bokhari , Mohammed Alawad , Ahmed Almazroa , Abdulrhman Aljouie

There is growing interest in integrating high-fidelity visual synthesis capabilities into large language models (LLMs) without compromising their strong reasoning capabilities. Existing methods that directly train LLMs or bridge LLMs and…

Computer Vision and Pattern Recognition · Computer Science 2025-08-11 Han Lin , Jaemin Cho , Amir Zadeh , Chuan Li , Mohit Bansal

Shoulder disorders, such as frozen shoulder (a.k.a., adhesive capsulitis), are common conditions affecting the health of people worldwide, and have a high incidence rate among the elderly and workers engaged in repetitive shoulder tasks. In…

Computer Vision and Pattern Recognition · Computer Science 2025-10-13 Jindong Hong , Wencheng Zhang , Shiqin Qiao , Jianhai Chen , Jianing Qiu , Chuanyang Zheng , Qian Xu , Yun Ji , Qianyue Wen , Weiwei Sun , Hao Li , Huizhen Li , Huichao Wang , Kai Wu , Meng Li , Yijun He , Lingjie Luo , Jiankai Sun

This study investigates the effects of including patients' clinical information on the performance of deep learning (DL) classifiers for disease location in chest X-ray images. Although current classifiers achieve high performance using…

Image and Video Processing · Electrical Eng. & Systems 2023-12-29 Chihcheng Hsieh , Isabel Blanco Nobre , Sandra Costa Sousa , Chun Ouyang , Margot Brereton , Jacinto C. Nascimento , Joaquim Jorge , Catarina Moreira

Multimodal Large Language Models (MLLMs) demonstrate exceptional semantic reasoning but struggle with 3D spatial perception when restricted to pure RGB inputs. Despite leveraging implicit geometric priors from 3D reconstruction models,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Jiaxin Zhang , Junjun Jiang , Haijie Li , Youyu Chen , Kui Jiang , Dave Zhenyu Chen

Large Vision-Language Models (LVLMs) have shown strong performance across multimodal tasks. However, they often produce hallucinations -- text that is inconsistent with visual input, due to the limited ability to verify information in…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Haonan Ge , Yiwei Wang , Ming-Hsuan Yang , Yujun Cai

MLLMs MLLMs are beginning to appear in clinical workflows, but their ability to perform complex medical reasoning remains unclear. We present Med-CMR, a fine-grained Medical Complex Multimodal Reasoning benchmark. Med-CMR distinguishes from…

Foundation models (FMs) have demonstrated strong performance across diverse pathology tasks. While there are similarities in the pre-training objectives of FMs, there is still limited understanding of their complementarity, redundancy in…

Computer Vision and Pattern Recognition · Computer Science 2025-12-15 Brennan Flannery , Thomas DeSilvio , Jane Nguyen , Satish E. Viswanath

Optical Coherence Tomography (OCT) is a novel and effective screening tool for ophthalmic examination. Since collecting OCT images is relatively more expensive than fundus photographs, existing methods use multi-modal learning to complement…

Image and Video Processing · Electrical Eng. & Systems 2023-08-02 Lehan Wang , Weihang Dai , Mei Jin , Chubin Ou , Xiaomeng Li

Artificial intelligence applied to retinal images offers significant potential for recognizing signs and symptoms of retinal conditions and expediting the diagnosis of eye diseases and systemic disorders. However, developing generalized…

Image and Video Processing · Electrical Eng. & Systems 2024-08-19 Boa Jang , Youngbin Ahn , Eun Kyung Choe , Chang Ki Yoon , Hyuk Jin Choi , Young-Gon Kim

Existing Medical Large Vision-Language Models (Med-LVLMs), encapsulating extensive medical knowledge, demonstrate excellent capabilities in understanding medical images. However, there remain challenges in visual localization in medical…

Computation and Language · Computer Science 2025-06-03 Yucheng Zhou , Lingran Song , Jianbing Shen

Vision-language alignment in multi-modal large language models (MLLMs) relies on supervised fine-tuning (SFT) or reinforcement learning (RL). To align multi-modal large language models (MLLMs) in the post-training stage, supervised…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Xin Jin , Siyuan Li , Siyong Jian , Kai Yu , Huan Wang

The challenge of Multimodal Deformable Image Registration (MDIR) lies in the conversion and alignment of features between images of different modalities. Generative models (GMs) cannot retain the necessary information enough from the source…

Computer Vision and Pattern Recognition · Computer Science 2024-08-21 Mingrui Ma , Weijie Wang , Jie Ning , Jianfeng He , Nicu Sebe , Bruno Lepri

Vision-language models (VLMs) have shown considerable potential in digital pathology, yet their effectiveness remains limited for fine-grained, disease-specific classification tasks such as distinguishing between glomerular subtypes. The…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Zhenhao Guo , Rachit Saluja , Tianyuan Yao , Quan Liu , Yuankai Huo , Benjamin Liechty , David J. Pisapia , Kenji Ikemura , Mert R. Sabuncu , Yihe Yang , Ruining Deng

Diabetic Retinopathy (DR) is a major cause of global blindness, necessitating early and accurate diagnosis. While deep learning models have shown promise in DR detection, their black-box nature often hinders clinical adoption due to a lack…

Computer Vision and Pattern Recognition · Computer Science 2025-08-22 Masato Ito , Kaito Tanaka , Keisuke Matsuda , Aya Nakayama
‹ Prev 1 4 5 6 7 8 10 Next ›