Related Experiment Videos
K-STAMM: a knowledge-enhanced spatial - temporal attention model with multimodal fusion for pneumonia prediction
S Anbukkarasi1, S Hemalatha2, Arunkumar Balakrishnan3
1Manipal Institute of Technology Bengaluru, Manipal Academy of Higher Education, Manipal, India. anbukkarasi.s@manipal.edu.
Scientific Reports
|April 9, 2026
Summary
This study introduces K-STAMM, a novel model for pneumonia prediction that effectively integrates diverse clinical data. K-STAMM enhances prediction accuracy by incorporating biomedical knowledge and advanced attention mechanisms for multimodal fusion.
Area of Science:
- Artificial Intelligence in Medicine
- Biomedical Informatics
- Machine Learning for Healthcare
Background:
- Accurate pneumonia prediction is challenging due to the heterogeneity of clinical data, including electronic health records (EHRs), medical imaging, and clinical text.
- Existing multimodal transformer models struggle with aligning different data types, maintaining temporal order, and integrating structured medical knowledge.
Purpose of the Study:
- To develop K-STAMM, a knowledge-augmented spatiotemporal attention model for effective multimodal fusion in pneumonia prediction.
- To address limitations in multimodal alignment, temporal regularity, and knowledge incorporation faced by current models.
Main Methods:
- K-STAMM integrates biomedical knowledge from the Unified Medical Language System using embedding-based representations for semantically enriched feature learning.
- Employs attention-based spatial modeling of structured EHR data and temporal sequence modeling to capture disease progression at irregular intervals.
- Utilizes a cross-modal fusion mechanism to harmonize chest X-ray images, clinical text, and knowledge embeddings into a single patient representation.
Main Results:
- K-STAMM achieved superior performance over baseline models on MIMIC-IV and MIMIC-CXR datasets, with an AUROC of 0.953, AUPRC of 0.962, and F1-score of 0.910.
- Ablation studies validated the significant contributions of knowledge augmentation, temporal attention, and multimodal fusion components.
- The model provides a scalable and interpretable framework for multimodal clinical prediction.
Conclusions:
- K-STAMM offers a robust and interpretable solution for multimodal clinical prediction, particularly for pneumonia.
- The integration of biomedical knowledge and advanced attention mechanisms significantly improves prediction accuracy and handles data heterogeneity.
- The framework demonstrates potential for broader applications in clinical decision support systems.