Related Experiment Video
Updated: Aug 5, 2026

Mapping Alzheimer's Disease Variants to Their Target Genes Using Computational Analysis of Chromatin Configuration
Published on: January 9, 2020
Mapping Alzheimer's disease heterogeneity through exploratory unsupervised learning
Catarina Xavier1, Ana Paula Correia2, Joana Lopes2
1i3S - Instituto de Investigação e Inovação em Saúde, Universidade do Porto, Porto, Portugal.
Introduction:
Alzheimer's disease (AD) is becoming one of the most pressing health challenges of the century, affecting circa 55 million people worldwide and expected to triple this number by 2050. Besides the growing prevalence, AD remains difficult to diagnose due to its long preclinical phase and substantial symptomatic heterogeneity. Despite the existence of international guidelines for AD diagnosis emphasizing the use of biomarkers (such as neuroimaging and cerebrospinal molecular biomarkers), most research studies are still using data disregarding their diagnostic quality, blurring the overall statistical outcomes. Aiming to circumvent this effect, biomarker confirmed samples from the Alzheimer's Disease Neuroimaging Initiative (ADNI) were used, which integrate genetic and neuroimaging data for all participants.
Methods:
Unsupervised machine learning clustering methods were tested to explore whether curated genetic markers can reveal underlying substructure within sporadic AD (sAD). Here two specific sets of single nucleotide polymorphisms (SNPs) were selected: (i) candidate SNPs highlighted in previous studies that identified the existence of different subtypes in sAD and (ii) SNPs found associated with sAD in genome wide association studies (GWAS) that used at least 50% of samples with biomarker-confirmed diagnosis.
Results:
This strategy minimized background noise and enabled the construction of high quality SNP sets for analysis. Different clustering algorithms were evaluated with agglomerative hierarchical clustering consistently yielding the most robust performance. Across SNP sets, most trials revealed a reproducible binary structure within sAD samples, suggesting the presence of genetically distinguishable subgroups.
Discussion:
These findings align with emerging multimodal evidence supporting biological heterogeneity in AD. Overall, this work demonstrates that curated SNP panels combined with unsupervised learning can uncover meaningful substructure in sAD, reinforcing the value of integrating high quality genetic data into subtype research.
Related Concept Videos
Alzheimer's Disease: Overview
The clinical diagnosis of AD hinges on the presence of memory and other cognitive impairments. Biomarkers, such as changes in Aβ and tau...
Alzheimer Disease l: Introduction
Dementia l: Introduction

