Related Experiment Video
Updated: Aug 8, 2025

Supervised Machine Learning for Semi-Quantification of Extracellular DNA in Glomerulonephritis
Published on: June 18, 2020
Identifying subtypes of chronic kidney disease with machine learning: development, internal validation and prognostic
Ashkan Dashtban1, Mehrdad A Mizani2, Laura Pasea1
1Institute of Health Informatics, University College London, London, UK.
Insights
Machine learning identified five chronic kidney disease (CKD) subtypes, improving risk prediction. These subtypes, including cardiometabolic and early-onset, show distinct mortality and admission risks, aiding targeted interventions.
Area of Science:
- Nephrology
- Data Science
- Public Health
Background:
- Chronic kidney disease (CKD) is linked to high multimorbidity, polypharmacy, and mortality.
- Current CKD classification and risk models inadequately capture disease complexity and outcomes.
- Improved CKD subtype definition is crucial for predicting outcomes and guiding interventions.
Purpose of the Study:
- To define distinct subtypes of chronic kidney disease (CKD) using machine learning.
- To evaluate the internal validity and prognostic accuracy of identified CKD subtypes.
- To assess the relationship between CKD subtypes and medication use.
Main Methods:
- Analysis of 350,067 incident and 195,422 prevalent CKD cases from electronic health records (2006-2020).
- Application of seven unsupervised machine learning methods to identify CKD subtypes using 66 variables.
- Evaluation of subtypes for internal validity, prognostic accuracy (5-year mortality/admissions), and medication patterns.
Main Results:
- Five distinct CKD subtypes were identified: Early-onset, Late-onset, Cancer, Metabolic, and Cardiometabolic.
- A predictive model achieved 95% accuracy in classifying CKD subtypes.
- The Cardiometabolic subtype exhibited the highest 5-year mortality (43.3%) and admission (29.5%) risks; Early-onset had the lowest.
Conclusions:
- This study, the largest CKD analysis using machine learning, identified five distinct subtypes.
- These subtypes demonstrate differential risks for mortality, hospital admissions, and new chronic diseases.
- The identified CKD subtypes offer significant relevance for etiological studies, therapeutic development, and risk prediction.
Background:
Although chronic kidney disease (CKD) is associated with high multimorbidity, polypharmacy, morbidity and mortality, existing classification systems (mild to severe, usually based on estimated glomerular filtration rate, proteinuria or urine albumin-creatinine ratio) and risk prediction models largely ignore the complexity of CKD, its risk factors and its outcomes. Improved subtype definition could improve prediction of outcomes and inform effective interventions.
Methods:
We analysed individuals ≥18 years with incident and prevalent CKD (n = 350,067 and 195,422 respectively) from a population-based electronic health record resource (2006-2020; Clinical Practice Research Datalink, CPRD). We included factors (n = 264 with 2670 derived variables), e.g. demography, history, examination, blood laboratory values and medications. Using a published framework, we identified subtypes through seven unsupervised machine learning (ML) methods (K-means, Diana, HC, Fanny, PAM, Clara, Model-based) with 66 (of 2670) variables in each dataset. We evaluated subtypes for: (i) internal validity (within dataset, across methods); (ii) prognostic validity (predictive accuracy for 5-year all-cause mortality and admissions); and (iii) medications (new and existing by British National Formulary chapter).
Findings:
After identifying five clusters across seven approaches, we labelled CKD subtypes: 1. Early-onset, 2. Late-onset, 3. Cancer, 4. Metabolic, and 5. Cardiometabolic. Internal validity: We trained a high performing model (using XGBoost) that could predict disease subtypes with 95% accuracy for incident and prevalent CKD (Sensitivity: 0.81-0.98, F1 score:0.84-0.97). Prognostic validity: 5-year all-cause mortality, hospital admissions, and incidence of new chronic diseases differed across CKD subtypes. The 5-year risk of mortality and admissions in the overall incident CKD population were highest in cardiometabolic subtype: 43.3% (42.3-42.8%) and 29.5% (29.1-30.0%), respectively, and lowest in the early-onset subtype: 5.7% (5.5-5.9%) and 18.7% (18.4-19.1%).
Medications:
Across CKD subtypes, the distribution of prescription medication classes at baseline varied, with highest medication burden in cardiometabolic and metabolic subtypes, and higher burden in prevalent than incident CKD.
Interpretation:
In the largest CKD study using ML, to-date, we identified five distinct subtypes in individuals with incident and prevalent CKD. These subtypes have relevance to study of aetiology, therapeutics and risk prediction.
Funding:
AstraZeneca UK Ltd, Health Data Research UK.
More Related Videos
09:00TBase - an Integrated Electronic Health Record and Research Database for Kidney Transplant Recipients
Published on: April 13, 2021
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018
Related Concept Videos
Chronic Kidney Disease I: Introduction
Chronic Kidney Disease III: Interprofessional Care
Chronic Kidney Disease IV: Nursing Management
Chronic Kidney Disease II: Clinical Manifestations
Acute Kidney Injury I: Introduction
Acute Kidney Injury IV: Diagnostic Studies and Prevention