Related Experiment Videos
Bridging the gap between performance and interpretability: An explainable disentangled multimodal framework for
Aniek Eijpe1, Soufyan Lakbir1, Melis Erdal Cesur2
1AI Technology for Life, Department of Information and Computing Sciences, Department of Biology, Utrecht University, Princetonplein 5, Utrecht, 3584 CC, The Netherlands.
None:
While multimodal survival prediction models are increasingly accurate, their complexity often reduces interpretability, limiting insight into how different data sources influence predictions. To address this, we introduce DIMAFx, an explainable multimodal framework for cancer survival prediction that produces disentangled, interpretable modality-specific and modality-shared representations from histopathology whole-slide images and transcriptomics data. Across four TCGA cancer cohorts, DIMAFx achieves survival prediction performance competitive with the state of the art and consistently stronger representation disentanglement. Leveraging its interpretable design, SHapley Additive exPlanations, and pathologist-in-the-loop annotations, DIMAFx facilitates systematic investigation of key multimodal interactions and the biological information encoded in the multimodal, disentangled representations. In breast cancer survival prediction, the most predictive features contain modality-shared information, including one capturing solid tumor morphology contextualized primarily by late estrogen response, where higher-grade morphology aligned with pathway downregulation was associated with increased risk, consistent with known breast cancer biology. Key modality-specific features capture microenvironmental signals from interacting adipose and stromal morphologies. These results show that DIMAFx substantially narrows the gap between performance and interpretability, supporting the application of such models in precision oncology.