English
Related papers

Related papers: HyPCA-Net: Advancing Multimodal Fusion in Medical …

200 papers

Multimodal learning has gained much success in recent years. However, current multimodal fusion methods adopt the attention mechanism of Transformers to implicitly learn the underlying correlation of multimodal features. As a result, the…

Computer Vision and Pattern Recognition · Computer Science 2025-11-27 Thanh-Dat Truong , Christophe Bobda , Nitin Agarwal , Khoa Luu

To improve the prediction of cancer survival using whole-slide images and transcriptomics data, it is crucial to capture both modality-shared and modality-specific information. However, multimodal frameworks often entangle these…

Computer Vision and Pattern Recognition · Computer Science 2025-06-30 Aniek Eijpe , Soufyan Lakbir , Melis Erdal Cesur , Sara P. Oliveira , Sanne Abeln , Wilson Silva

In vision and linguistics; the main input modalities are facial expressions, speech patterns, and the words uttered. The issue with analysis of any one mode of expression (Visual, Verbal or Vocal) is that lot of contextual information can…

Computer Vision and Pattern Recognition · Computer Science 2022-11-29 Kunjal Panchal

Survival prediction is crucial for cancer patients as it provides early prognostic information for treatment planning. Recently, deep survival models based on deep learning and medical images have shown promising performance for survival…

Image and Video Processing · Electrical Eng. & Systems 2023-10-03 Mingyuan Meng , Lei Bi , Michael Fulham , Dagan Feng , Jinman Kim

A key challenge in learning from multimodal biological data is missing modalities, where data from one or more modalities are absent for some patients. Existing approaches either exclude patients with missing modalities, impute missing…

Machine Learning · Computer Science 2026-05-19 Sina Tabakhi , Chen , Chen , Haiping Lu

Recent studies have focused on utilizing multi-modal data to develop robust models for facial Action Unit (AU) detection. However, the heterogeneity of multi-modal data poses challenges in learning effective representations. One such…

Computer Vision and Pattern Recognition · Computer Science 2023-08-24 Xiang Zhang , Huiyuan Yang , Taoyue Wang , Xiaotian Li , Lijun Yin

Medical Visual Question Answering (Med-VQA) answers clinical questions using medical images, aiding diagnosis. Designing the MedVQA system holds profound importance in assisting clinical diagnosis and enhancing diagnostic accuracy. Building…

Computer Vision and Pattern Recognition · Computer Science 2025-04-11 Junkai Zhang , Bin Li , Shoujun Zhou , Yue Du

In this study, we introduce a multi-modal approach that efficiently integrates multi-scale clinical and dermoscopy features within a single network, thereby substantially reducing model parameters. The proposed method includes three novel…

Image and Video Processing · Electrical Eng. & Systems 2024-03-31 Peng Tang , Tobias Lasser

Transformer-based methods have demonstrated excellent performance on super-resolution visual tasks, surpassing conventional convolutional neural networks. However, existing work typically restricts self-attention computation to…

Computer Vision and Pattern Recognition · Computer Science 2024-05-09 Shu-Chuan Chu , Zhi-Chao Dou , Jeng-Shyang Pan , Shaowei Weng , Junbao Li

The use of complex attention modules has improved the performance of the Visual Question Answering (VQA) task. This work aims to learn an improved multi-modal representation through dense interaction of visual and textual modalities. The…

Computer Vision and Pattern Recognition · Computer Science 2023-03-01 Aakansha Mishra , Ashish Anand , Prithwijit Guha

The 3D scene understanding is mainly considered as a crucial requirement in computer vision and robotics applications. One of the high-level tasks in 3D scene understanding is semantic segmentation of RGB-Depth images. With the availability…

Computer Vision and Pattern Recognition · Computer Science 2019-12-30 Fahimeh Fooladgar , Shohreh Kasaei

Multimodal medical image fusion integrates complementary information from different imaging modalities to enhance diagnostic accuracy and treatment planning. While deep learning methods have advanced performance, existing approaches face…

Image and Video Processing · Electrical Eng. & Systems 2025-08-06 Meng Zhou , Farzad Khalvati

Recently, dense connections have attracted substantial attention in computer vision because they facilitate gradient flow and implicit deep supervision during training. Particularly, DenseNet, which connects each layer to every other layer…

Computer Vision and Pattern Recognition · Computer Science 2019-03-05 Jose Dolz , Karthik Gopinath , Jing Yuan , Herve Lombaert , Christian Desrosiers , Ismail Ben Ayed

Cancer has relational information residing at varying scales, modalities, and resolutions of the acquired data, such as radiology, pathology, genomics, proteomics, and clinical records. Integrating diverse data types can improve the…

Machine Learning · Computer Science 2024-07-29 Asim Waqas , Aakash Tripathi , Ravi P. Ramachandran , Paul Stewart , Ghulam Rasool

Crowd counting research has made significant advancements in real-world applications, but it remains a formidable challenge in cross-modal settings. Most existing methods rely solely on the optical features of RGB images, ignoring the…

Computer Vision and Pattern Recognition · Computer Science 2022-11-15 Youjia Zhang , Soyun Choi , Sungeun Hong

The UNet architecture, based on Convolutional Neural Networks (CNN), has demonstrated its remarkable performance in medical image analysis. However, it faces challenges in capturing long-range dependencies due to the limited receptive…

Image and Video Processing · Electrical Eng. & Systems 2023-07-28 Liang Xu , Mingxiao Chen , Yi Cheng , Pengfei Shao , Shuwei Shen , Peng Yao , Ronald X. Xu

Cancer survival prediction requires integrating pathological Whole Slide Images (WSIs) and genomic profiles, a challenging task due to the inherent heterogeneity and the complexity of modeling both inter- and intra-modality interactions.…

Image and Video Processing · Electrical Eng. & Systems 2025-07-08 Mingxin Liu , Chengfei Cai , Jun Li , Pengbo Xu , Jinze Li , Jiquan Ma , Jun Xu

Hybrid quantum and classical learning aims to couple quantum feature maps with the robustness of classical neural networks, yet most architectures treat the quantum circuit as an isolated feature extractor and merge its measurements with…

Machine Learning · Computer Science 2025-12-23 Azadeh Alavi , Fatemeh Kouchmeshki , Abdolrahman Alavi

Accurate medical image segmentation requires effective modeling of both long-range dependencies and fine-grained boundary details. While transformers mitigate the issue of insufficient semantic information arising from the limited receptive…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Yanxin Li , Hui Wan , Libin Lan

Oral cancer is frequently diagnosed at later stages due to its similarity to other lesions. Existing research on computer aided diagnosis has made progress using deep learning; however, most approaches remain limited by small, imbalanced…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Joy Naoum , Revana Salama , Ali Hamdi