English
Related papers

Related papers: MDF-MLLM: Deep Fusion Through Cross-Modal Feature …

200 papers

Deep learning (DL) is the state-of-the-art methodology in various medical image segmentation tasks. However, it requires relatively large amounts of manually labeled training data, which may be infeasible to generate in some applications.…

Image and Video Processing · Electrical Eng. & Systems 2021-03-22 Long Xie , Laura E. M. Wisse , Jiancong Wang , Sadhana Ravikumar , Trevor Glenn , Anica Luther , Sydney Lim , David A. Wolk , Paul A. Yushkevich

Multimodal MRIs play a crucial role in clinical diagnosis and treatment. Feature disentanglement (FD)-based methods, aiming at learning superior feature representations for multimodal data analysis, have achieved significant success in…

Computer Vision and Pattern Recognition · Computer Science 2025-02-28 Tianling Liu , Hongying Liu , Fanhua Shang , Lequan Yu , Tong Han , Liang Wan

Current video understanding models excel at recognizing "what" is happening but fall short in high-level cognitive tasks like causal reasoning and future prediction, a limitation rooted in their lack of commonsense world knowledge. To…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 L'ea Dubois , Klaus Schmidt , Chengyu Wang , Ji-Hoon Park , Lin Wang , Santiago Munoz

Our research is motivated by the urgent global issue of a large population affected by retinal diseases, which are evenly distributed but underserved by specialized medical expertise, particularly in non-urban areas. Our primary objective…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Deependra Singh , Saksham Agarwal , Subhankar Mishra

Accurate medical image analysis can greatly assist clinical diagnosis, but its effectiveness relies on high-quality expert annotations Obtaining pixel-level labels for medical images, particularly fundus images, remains costly and…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Xi Luo , Shixin Xu , Ying Xie , JianZhong Hu , Yuwei He , Yuhui Deng , Huaxiong Huang

Multimodal large language models (MLLMs) have achieved impressive performance on visual perception and reasoning tasks with RGB imagery, yet they remain fragile under common degradations, such as fog, blur, or low-light conditions. Infrared…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Abrar Majeedi , Zhiyuan Ruan , Ziyi Zhao , Hongcheng Wang , Jianglin Lu , Yin Li

Neuroblastoma (NB), a leading cause of childhood cancer mortality, exhibits significant histopathological variability, necessitating precise subtyping for accurate prognosis and treatment. Traditional diagnostic methods rely on subjective…

Computer Vision and Pattern Recognition · Computer Science 2025-03-20 Huangwei Chen , Yifei Chen , Zhenyu Yan , Mingyang Ding , Chenlei Li , Zhu Zhu , Feiwei Qin

This study investigates a hybrid method for text classification that integrates deep feature extraction from large language models, multi-scale fusion through feature pyramids, and structured modeling with graph neural networks to enhance…

Computation and Language · Computer Science 2025-11-11 Xiangchen Song , Yulin Huang , Jinxu Guo , Yuchen Liu , Yaxuan Luan

Major depressive disorder is a serious and heterogeneous psychiatric disorder that needs accurate diagnosis. Resting-state functional MRI (rsfMRI), which captures multiple perspectives on brain structure, function, and connectivity, is…

Machine Learning · Computer Science 2023-08-21 Yunsong Luo , Wenyu Chen , Ling Zhan , Jiang Qiu , Tao Jia

Universal lesion detection (ULD) on computed tomography (CT) images is an important but underdeveloped problem. Recently, deep learning-based approaches have been proposed for ULD, aiming to learn representative features from annotated CT…

Computer Vision and Pattern Recognition · Computer Science 2019-12-17 Zihao Li , Shu Zhang , Junge Zhang , Kaiqi Huang , Yizhou Wang , Yizhou Yu

Robot vision has greatly benefited from advancements in multimodal fusion techniques and vision-language models (VLMs). We adopt a task-oriented perspective to systematically review the applications and advancements of multimodal fusion…

Large language models (LLMs) have demonstrated significant progress in multilingual language understanding and generation. However, due to the imbalance in training data, their capabilities in non-English languages are limited. Recent…

Computation and Language · Computer Science 2025-03-06 Wenshuai Huo , Xiaocheng Feng , Yichong Huang , Chengpeng Fu , Baohang Li , Yangfan Ye , Zhirui Zhang , Dandan Tu , Duyu Tang , Yunfei Lu , Hui Wang , Bing Qin

Multimodal medical analysis combining image and tabular data has gained increasing attention. However, effective fusion remains challenging due to cross-modal discrepancies in feature dimensions and modality contributions, as well as the…

Computer Vision and Pattern Recognition · Computer Science 2025-09-17 Congjing Yu , Jing Ye , Yang Liu , Xiaodong Zhang , Zhiyong Zhang

Visual deep learning (VDL) systems have shown significant success in real-world applications like image recognition, object detection, and autonomous driving. To evaluate the reliability of VDL, a mainstream approach is software testing,…

Software Engineering · Computer Science 2024-12-24 Liwen Wang , Yuanyuan Yuan , Ao Sun , Zongjie Li , Pingchuan Ma , Daoyuan Wu , Shuai Wang

Federated learning (FL) enables the collaborative training of deep neural networks across decentralized data archives (i.e., clients) without sharing the local data of the clients. Most of the existing FL methods assume that the data…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Barış Büyüktaş , Gencer Sumbul , Begüm Demir

The impressive performance of Large Language Model (LLM) has prompted researchers to develop Multi-modal LLM (MLLM), which has shown great potential for various multi-modal tasks. However, current MLLM often struggles to effectively address…

Computer Vision and Pattern Recognition · Computer Science 2024-12-24 Yeyuan Wang , Dehong Gao , Bin Li , Rujiao Long , Lei Yi , Xiaoyan Cai , Libin Yang , Jinxia Zhang , Shanqing Yu , Qi Xuan

Diabetic retinopathy and diabetic macular edema are significant complications of diabetes that can lead to vision loss. Early detection through ultra-widefield fundus imaging enhances patient outcomes but presents challenges in image…

Image and Video Processing · Electrical Eng. & Systems 2024-09-20 Philippe Zhang , Pierre-Henri Conze , Mathieu Lamard , Gwenolé Quellec , Mostafa El Habib Daho

Precise segmentation of retinal arteries and veins carries the diagnosis of systemic cardiovascular conditions. However, standard convolutional architectures often yield topologically disjointed segmentations, characterized by gaps and…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Iftekhar Ahmed , Shakib Absar , Aftar Ahmad Sami , Shadman Sakib , Debojyoti Biswas , Seraj Al Mahmud Mostafa

Diabetic Retinopathy (DR) is one of the major causes of visual impairment and blindness across the world. It is usually found in patients who suffer from diabetes for a long period. The major focus of this work is to derive optimal…

Image and Video Processing · Electrical Eng. & Systems 2020-06-02 J. D. Bodapati , N. Veeranjaneyulu , S. N. Shareef , S. Hakak , M. Bilal , P. K. R. Maddikunta , O. Jo

The fusion of language and vision in large vision-language models (LVLMs) has revolutionized deep learning-based object detection by enhancing adaptability, contextual reasoning, and generalization beyond traditional architectures. This…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Ranjan Sapkota , Manoj Karkee