Related Experiment Video
Updated: Jul 14, 2025

A Familial Hypercholesterolemia Human Liver Chimeric Mouse Model Using Induced Pluripotent Stem Cell-derived Hepatocytes
Published on: September 15, 2018
Generation and validation of a classification model to diagnose familial hypercholesterolaemia in adults
João Albuquerque1, Ana Margarida Medeiros2, Ana Catarina Alves2
1Departamento de Biomedicina, Unidade de Bioquímica, Faculdade de Medicina, Universidade do Porto, 4200-319, Porto, Portugal; Centro de Estatística e Aplicações, Faculdade de Ciências, Universidade de Lisboa, 1749-016, Lisboa, Portugal; Grupo de Investigação Cardiovascular, Departamento de Promoção da Saúde e Prevenção de Doenças Não Transmissíveis, Instituto Nacional de Saúde Doutor Ricardo Jorge, 1649-016, Lisboa, Portugal.
Insights
Early diagnosis of familial hypercholesterolaemia (FH) reduces cardiovascular disease risk. A new logistic regression model, trained on multiple cohorts, accurately identifies FH cases across diverse populations, outperforming traditional criteria.
Area of Science:
- Cardiovascular disease research
- Medical diagnostics
- Machine learning in healthcare
Background:
- Early diagnosis of familial hypercholesterolaemia (FH) significantly reduces cardiovascular disease (CVD) risk.
- Current screening methods often rely on single-cohort studies, limiting generalizability.
- Logistic regression (LR) and machine learning show promise for FH screening.
Purpose of the Study:
- To develop and validate a logistic regression (LR) based algorithm for familial hypercholesterolaemia (FH) screening.
- To assess the algorithm's performance across multiple national cohorts and an external dataset.
- To compare the LR model's efficacy against traditional clinical criteria.
Main Methods:
- Developed a logistic regression (LR) algorithm using data from three national FH cohorts (Portugal, Brazil, Sweden).
- Validated the LR model on independent samples from these cohorts and an external Italian dataset.
- Assessed discriminatory ability using AUROC and AUPRC; compared performance with Dutch Lipid Clinic Network (DLCN) criteria.
Main Results:
- The LR model demonstrated higher AUROC and AUPRC values on testing sets compared to the training set.
- Significantly more correct classifications were achieved with the LR model versus DLCN criteria across Brazilian, Swedish, and Italian test sets.
- Improved accuracy, G mean, and F1 score were observed for all testing sets using the LR model.
Conclusions:
- The LR model exhibits superior classification ability compared to DLCN criteria, identifying similar numbers of FH cases with fewer false positives.
- The model shows excellent generalization across diverse populations, indicating its potential as an effective FH screening tool.
- This multi-cohort developed algorithm offers a robust approach for widespread FH screening.
Background And Aims:
The early diagnosis of familial hypercholesterolaemia is associated with a significant reduction in cardiovascular disease (CVD) risk. While the recent use of statistical and machine learning algorithms has shown promising results in comparison with traditional clinical criteria, when applied to screening of potential FH cases in large cohorts, most studies in this field are developed using a single cohort of patients, which may hamper the application of such algorithms to other populations. In the current study, a logistic regression (LR) based algorithm was developed combining observations from three different national FH cohorts, from Portugal, Brazil and Sweden. Independent samples from these cohorts were then used to test the model, as well as an external dataset from Italy.
Methods:
The area under the receiver operating characteristics (AUROC) and precision-recall (AUPRC) curves was used to assess the discriminatory ability among the different samples. Comparisons between the LR model and Dutch Lipid Clinic Network (DLCN) clinical criteria were performed by means of McNemar tests, and by the calculation of several operating characteristics.
Results:
AUROC and AUPRC values were generally higher for all testing sets when compared to the training set. Compared with DLCN criteria, a significantly higher number of correctly classified observations were identified for the Brazilian (p < 0.01), Swedish (p < 0.01), and Italian testing sets (p < 0.01). Higher accuracy (Acc), G mean and F1 score values were also observed for all testing sets.
Conclusions:
Compared to DLCN criteria, the LR model revealed improved ability to correctly classify observations, and was able to retain a similar number of FH cases, with less false positive retention. Generalization of the LR model was very good across all testing samples, suggesting it can be an effective screening tool if applied to different populations.
Related Concept Videos
Atherosclerosis II: Clinical Manifestations and Diagnostic Tests
Cholesterol: Significance and Regulation
Considering cholesterol and...
Atherosclerosis III: Management
Pedigree Analysis

