Related Experiment Video
Updated: Dec 2, 2025

A Familial Hypercholesterolemia Human Liver Chimeric Mouse Model Using Induced Pluripotent Stem Cell-derived Hepatocytes
Published on: September 15, 2018
Performance and clinical utility of supervised machine-learning approaches in detecting familial
Ralph K Akyea1, Nadeem Qureshi1, Joe Kai1
1Primary Care Stratified Medicine, Division of Primary Care, University of Nottingham, Nottingham, UK.
Insights
Machine learning models significantly improve the detection of familial hypercholesterolaemia (FH), an inherited cholesterol disorder. Ensemble learning offers the best balance of accuracy and clinical utility for identifying FH cases in primary care.
Area of Science:
- Cardiovascular Medicine
- Medical Informatics
- Genetics
Background:
- Familial hypercholesterolaemia (FH) is a common genetic disorder causing elevated LDL cholesterol, leading to premature heart disease.
- Most FH cases remain undiagnosed, missing opportunities for early intervention and prevention.
- Machine learning (ML) shows promise for FH detection in electronic health records, but clinical utility needs further assessment.
Purpose of the Study:
- To evaluate the performance and clinical utility of various ML algorithms for enhancing FH detection in a large primary care population.
- To compare the predictive accuracy and case-finding workload of different ML models in identifying FH.
Main Methods:
- A retrospective cohort study analyzed 4,027,775 UK primary care records (1999-2019).
- Five ML algorithms (logistic regression, random forest, gradient boosting, neural networks, ensemble learning) were assessed for FH detection.
- Performance metrics included AUC, calibration slope, likelihood ratios, and expected case-review workload.
Main Results:
- Four ML approaches (excluding logistic regression) demonstrated high predictive accuracy (AUC > 0.89).
- Ensemble learning achieved the highest positive likelihood ratio (45.5) and a low case-review workload (0.73%).
- ML models identified novel predictive features, such as raised triglycerides, which can decrease FH likelihood.
Conclusions:
- ML models offer high accuracy for FH detection, presenting opportunities to increase diagnosis rates.
- Different ML models vary significantly in their clinical case-finding workload and efficiency.
- Ensemble learning appears most promising for efficient and accurate FH case identification in primary care settings.
Abstract:
Familial hypercholesterolaemia (FH) is a common inherited disorder, causing lifelong elevated low-density lipoprotein cholesterol (LDL-C). Most individuals with FH remain undiagnosed, precluding opportunities to prevent premature heart disease and death. Some machine-learning approaches improve detection of FH in electronic health records, though clinical impact is under-explored. We assessed performance of an array of machine-learning approaches for enhancing detection of FH, and their clinical utility, within a large primary care population. A retrospective cohort study was done using routine primary care clinical records of 4,027,775 individuals from the United Kingdom with total cholesterol measured from 1 January 1999 to 25 June 2019. Predictive accuracy of five common machine-learning algorithms (logistic regression, random forest, gradient boosting machines, neural networks and ensemble learning) were assessed for detecting FH. Predictive accuracy was assessed by area under the receiver operating curves (AUC) and expected vs observed calibration slope; with clinical utility assessed by expected case-review workload and likelihood ratios. There were 7928 incident diagnoses of FH. In addition to known clinical features of FH (raised total cholesterol or LDL-C and family history of premature coronary heart disease), machine-learning (ML) algorithms identified features such as raised triglycerides which reduced the likelihood of FH. Apart from logistic regression (AUC, 0.81), all four other ML approaches had similarly high predictive accuracy (AUC > 0.89). Calibration slope ranged from 0.997 for gradient boosting machines to 1.857 for logistic regression. Among those screened, high probability cases requiring clinical review varied from 0.73% using ensemble learning to 10.16% using deep learning, but with positive predictive values of 15.5% and 2.8% respectively. Ensemble learning exhibited a dominant positive likelihood ratio (45.5) compared to all other ML models (7.0-14.4). Machine-learning models show similar high accuracy in detecting FH, offering opportunities to increase diagnosis. However, the clinical case-finding workload required for yield of cases will differ substantially between models.
More Related Videos
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018
12:40Using Human Induced Pluripotent Stem Cell-derived Hepatocyte-like Cells for Drug Discovery
Published on: May 19, 2018
Related Concept Videos
Lipid-Lowering Drugs: Statins and Miscellaneous Agents
Atherosclerosis II: Clinical Manifestations and Diagnostic Tests
Atherosclerosis III: Management