Related Experiment Video
Updated: Oct 3, 2025

Quantitative Analysis and Characterization of Atherosclerotic Lesions in the Murine Aortic Sinus
Published on: December 7, 2013
Construction of genetic classification model for coronary atherosclerosis heart disease using three machine learning
Wenjuan Peng1, Yuan Sun1, Ling Zhang2
1Department of Epidemiology and Health Statistics, School of Public Health, Capital Medical University, and Beijing Municipal Key Laboratory of Clinical Epidemiology, No. 10, Xi Toutiao You Anmenwai, Fengtai District, Beijing, 100069, China.
This study identified 12 gene signatures for diagnosing coronary artery disease (CAD) early. The support vector machine (SVM) model demonstrated the best performance in classifying CAD using gene expression data.
Area of Science:
- Genomics
- Bioinformatics
- Cardiovascular Disease Research
Background:
- Early-stage coronary artery disease (CAD) often presents asymptomatically, leading to missed diagnoses.
- Gene expression levels fluctuate during disease development, offering potential for diagnostic biomarkers.
- Current diagnostic methods for CAD require innovation to improve early detection.
Purpose of the Study:
- To construct genetic classification models for CAD using gene expression data.
- To identify key genes and gene modules associated with CAD pathogenesis.
- To provide new insights into the molecular mechanisms underlying CAD.
Main Methods:
- Utilized R software and downloaded three CAD-related gene expression datasets (GSE12288, GSE7638, GSE66360) from the Gene Expression Omnibus database.
- Employed Limma package for identifying differentially expressed genes (DEGs) and WGCNA for recognizing CAD-related gene modules and hub genes.
- Applied recursive feature elimination to select optimal feature genes (OFGs) and constructed classification models using Support Vector Machine (SVM), Random Forest (RF), and Logistic Regression (LR), validated with ROC curve analysis.
Main Results:
- Identified 374 DEGs, eight gene modules, 33 hub genes, and 12 OFGs (including HTR4, KISS1, CA12, CAMK2B, KLK2, DDC, CNGB1, DERL1, BCL6, LILRA2, HCK, MTF2).
- The Support Vector Machine (SVM) model achieved the highest accuracy (75.58%) and an Area Under the Curve (AUC) of 0.813.
- Random Forest (RF) and Logistic Regression (LR) models showed accuracies of 63.57% and 63.95%, respectively, with AUCs of 0.727 and 0.783.
Conclusions:
- The study identified 12 gene signatures crucial to the pathogenic mechanisms of CAD.
- The SVM model exhibited superior performance in classifying CAD compared to RF and LR models.
- These findings suggest the potential of gene expression-based classifiers for early CAD diagnosis and understanding its pathogenesis.
More Related Videos
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018
06:19Constructing and Visualizing Models using Mime-based Machine-learning Framework
Published on: July 22, 2025
Related Concept Videos
Coronary Artery Disease I: Introduction
Coronary Artery Disease IV: Preventive Measures
Coronary Artery Disease II: Pathophysiology
Atherosclerosis I: Introduction
Atherosclerosis II: Clinical Manifestations and Diagnostic Tests
Cardiomyopathy I: Introduction and Classification