English
Related papers

Related papers: Architecture-Agnostic Modality-Isolated Gated Fusi…

200 papers

In this work, we consider the task of pairwise cross-modality image registration, which may benefit from exploiting additional images available only at training time from an additional modality that is different to those being registered.…

Computer Vision and Pattern Recognition · Computer Science 2022-07-27 Qianye Yang , David Atkinson , Yunguan Fu , Tom Syer , Wen Yan , Shonit Punwani , Matthew J. Clarkson , Dean C. Barratt , Tom Vercauteren , Yipeng Hu

Current imaging-based prostate cancer diagnosis requires both MR T2-weighted (T2w) and diffusion-weighted imaging (DWI) sequences, with additional sequences for potentially greater accuracy improvement. However, measuring diffusion patterns…

Image and Video Processing · Electrical Eng. & Systems 2024-11-13 Weixi Yi , Yipei Wang , Natasha Thorley , Alexander Ng , Shonit Punwani , Veeru Kasivisvanathan , Dean C. Barratt , Shaheer Ullah Saeed , Yipeng Hu

We introduce PGF-Net (Progressive Gated-Fusion Network), a novel deep learning framework designed for efficient and interpretable multimodal sentiment analysis. Our framework incorporates three primary innovations. Firstly, we propose a…

Machine Learning · Computer Science 2025-08-25 Bin Wen , Tien-Ping Tan

Multi-modal learning has emerged as a crucial research direction, as integrating textual and visual information can substantially enhance performance in tasks such as classification, retrieval, and scene understanding. Despite advances with…

Computer Vision and Pattern Recognition · Computer Science 2025-07-31 Md. Mithun Hossain , Md. Shakil Hossain , Sudipto Chaki , M. F. Mridha

Prostate gland segmentation from T2-weighted MRI is a critical yet challenging task in clinical prostate cancer assessment. While deep learning-based methods have significantly advanced automated segmentation, most conventional…

Image and Video Processing · Electrical Eng. & Systems 2025-06-25 Ahmad Mustafa , Reza Rastegar , Ghassan AlRegib

Multimodal remote sensing semantic segmentation enhances scene interpretation by exploiting complementary physical cues from heterogeneous data. Although pretrained Vision Foundation Models (VFMs) provide strong general-purpose…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Haocheng Li , Juepeng Zheng , Shuangxi Miao , Ruibo Lu , Guosheng Cai , Haohuan Fu , Jianxi Huang

A key challenge in learning from multimodal biological data is missing modalities, where data from one or more modalities are absent for some patients. Existing approaches either exclude patients with missing modalities, impute missing…

Machine Learning · Computer Science 2026-05-19 Sina Tabakhi , Chen , Chen , Haiping Lu

Multimodal MR-US registration is critical for prostate cancer diagnosis. However, this task remains challenging due to significant modality discrepancies. Existing methods often fail to align critical boundaries while being overly sensitive…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Xudong Ma , Nantheera Anantrasirichai , Stefanos Bolomytis , Alin Achim

Multimodal physiological data powers clinical AI systems from intensive care units to wearable devices, but sensors routinely fail in practice. Two failure modes are common: modality missing, where an entire channel is absent, and…

Machine Learning · Computer Science 2026-05-18 Wugeng Zheng , Ziwen Kan , Tianlong Chen , Chen Chen , Song Wang

Multi-modality image fusion and segmentation play a vital role in autonomous driving and robotic operation. Early efforts focus on boosting the performance for only one task, \emph{e.g.,} fusion or segmentation, making it hard to…

Computer Vision and Pattern Recognition · Computer Science 2023-08-07 Jinyuan Liu , Zhu Liu , Guanyao Wu , Long Ma , Risheng Liu , Wei Zhong , Zhongxuan Luo , Xin Fan

Generalizability in deep neural networks plays a pivotal role in medical image segmentation. However, deep learning-based medical image analyses tend to overlook the importance of frequency variance, which is critical element for achieving…

Image and Video Processing · Electrical Eng. & Systems 2024-05-13 Ju-Hyeon Nam , Nur Suriza Syazwany , Su Jung Kim , Sang-Chul Lee

Reliable unmanned aerial vehicle (UAV) detection is critical for autonomous airspace monitoring but remains challenging when integrating sensor streams that differ substantially in resolution, perspective, and field of view. Conventional…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Ishrat Jahan , Molla E Majid , M Murugappan , Muhammad E. H. Chowdhury , N. B. Prakash , Saad Bin Abul Kashem , Balamurugan Balusamy , Amith Khandakar

Complete and high-quality multi-modal Magnetic Resonance Imaging (MRI) is essential for accurate neuro-oncological assessment, as each contrast provides complementary anatomical and pathological information. However, acquiring all…

Image and Video Processing · Electrical Eng. & Systems 2026-05-04 Zaid A. Abod , Furqan Aziz

Fusing an arbitrary number of modalities is vital for achieving robust multi-modal fusion of semantic segmentation yet remains less explored to date. Recent endeavors regard RGB modality as the center and the others as the auxiliary,…

Computer Vision and Pattern Recognition · Computer Science 2024-08-27 Xu Zheng , Yuanhuiyi Lyu , Jiazhou Zhou , Lin Wang

Automatic modulation classification (AMC) is essential for wireless communication systems in both military and civilian applications. However, existing deep learning-based AMC methods often require large labeled signals and struggle with…

Signal Processing · Electrical Eng. & Systems 2025-08-05 Haoyue Tan , Yu Li , Zhenxi Zhang , Xiaoran Shi , Feng Zhou

Recent CLIP-based few-shot semantic segmentation methods introduce class-level textual priors to assist segmentation by typically using a single prompt (e.g., a photo of class). However, these approaches often result in incomplete…

Computer Vision and Pattern Recognition · Computer Science 2025-11-20 Qiang Jiao , Bin Yan , Yi Yang , Mengrui Shi , Qiang Zhang

Inter reader variability and cross site domain shift challenge the automatic segmentation of prostate anatomy using T2 weighted MRI images. This study investigates whether transformer models can retain precision amid such heterogeneity. We…

Image and Video Processing · Electrical Eng. & Systems 2026-04-23 Shatha Abudalou , Jung Choi , Yasin Yilmaz , Yoganand Balagurunathan

T2-weighted magnetic resonance imaging (MRI) and diffusion-weighted imaging (DWI) are essential components for cervical cancer diagnosis. However, combining these channels for training deep learning models are challenging due to…

Image and Video Processing · Electrical Eng. & Systems 2023-06-21 Reza Kalantar , Sebastian Curcean , Jessica M Winfield , Gigin Lin , Christina Messiou , Matthew D Blackledge , Dow-Mu Koh

Automatic segmentation of the prostate cancer from the multi-modal magnetic resonance images is of critical importance for the initial staging and prognosis of patients. However, how to use the multi-modal image features more efficiently is…

Image and Video Processing · Electrical Eng. & Systems 2020-11-10 Guokai Zhang , Xiaoang Shen , Ye Luo , Jihao Luo , Zeju Wang , Weigang Wang , Binghui Zhao , Jianwei Lu

Accurate mitotic figure classification is crucial in computational pathology, as mitotic activity informs cancer grading and patient prognosis. Distinguishing atypical mitotic figures (AMFs), which indicate higher tumor aggressiveness, from…

Image and Video Processing · Electrical Eng. & Systems 2025-09-04 Hana Feki , Alice Blondel , Thomas Walter