Interpretable multimodal learning for integrating neuroimaging and genetic data in Alzheimer's disease
Kun Zhao1, Siyuan Dai1, Yingying Zhang2
1Electrical & Computer Engineering, University of Pittsburgh, Pittsburgh, PA, United States.
Introduction:
Early detection of Alzheimer's disease (AD) requires models that combine brain structure changes with genetic risk, but existing methods struggle to align these different data types.
Methods:
We present R-GenIMA, an interpretable multimodal large language model that pairs a region-of-interest vision transformer with genetic prompting to jointly analyze structural MRI and single nucleotide polymorphisms (SNPs). Each brain region becomes a visual token and SNP profiles are encoded as structured text, letting the model link regional atrophy to genetic factors through cross-modal attention. Tested on the ADNI cohort, R-GenIMA performs well in classifying four groups: normal cognition, subjective memory concerns, mild cognitive impairment, and AD.
Results:
Beyond accuracy, it produces biologically meaningful explanations, identifying stage-specific brain regions and genes. The model consistently highlighted known AD risk genes (APOE, BIN1, CLU, RBFOX1) and revealed stage-specific patterns: striatal involvement in subjective decline, frontotemporal changes in early impairment, and broad network disruption in AD.
Discussion:
These results show that interpretable multimodal AI can integrate imaging and genetics to reveal disease mechanisms, providing a foundation for clinical tools that enable earlier risk assessment and inform precision treatment in Alzheimer's disease.

