English

Cross-Modal Translation and Alignment for Survival Analysis

Image and Video Processing 2023-09-25 v1 Computer Vision and Pattern Recognition Machine Learning

Abstract

With the rapid advances in high-throughput sequencing technologies, the focus of survival analysis has shifted from examining clinical indicators to incorporating genomic profiles with pathological images. However, existing methods either directly adopt a straightforward fusion of pathological features and genomic profiles for survival prediction, or take genomic profiles as guidance to integrate the features of pathological images. The former would overlook intrinsic cross-modal correlations. The latter would discard pathological information irrelevant to gene expression. To address these issues, we present a Cross-Modal Translation and Alignment (CMTA) framework to explore the intrinsic cross-modal correlations and transfer potential complementary information. Specifically, we construct two parallel encoder-decoder structures for multi-modal data to integrate intra-modal information and generate cross-modal representation. Taking the generated cross-modal representation to enhance and recalibrate intra-modal representation can significantly improve its discrimination for comprehensive survival analysis. To explore the intrinsic crossmodal correlations, we further design a cross-modal attention module as the information bridge between different modalities to perform cross-modal interactions and transfer complementary information. Our extensive experiments on five public TCGA datasets demonstrate that our proposed framework outperforms the state-of-the-art methods.

Cite

@article{arxiv.2309.12855,
  title  = {Cross-Modal Translation and Alignment for Survival Analysis},
  author = {Fengtao Zhou and Hao Chen},
  journal= {arXiv preprint arXiv:2309.12855},
  year   = {2023}
}

Comments

Accepted by ICCV2023

R2 v1 2026-06-28T12:29:26.875Z