Related Experiment Video
Updated: May 21, 2025

Author Spotlight: Advancing Cardiovascular Imaging - Introducing the Spatially Weighted Calcium Score for Early Disease Detection
Published on: September 22, 2023
Assessment of atherosclerosis risk in an insufficient sample size based on K-Means BS and TW-gcForest
Yudong Zhang1, Wenjun Liu1, Lidan He1
1School of Mathematics and Statistics, Nanjing University of Information Science and Technology, Nanjing, China.
Background:
Insufficient data is a common issue encountered in studies of atherosclerosis risk assessment. However, when the sample size is insufficient, commonly used classification algorithms often fail to achieve superior performance, thereby limiting the application of atherosclerosis data in patient risk assessment.
Purpose:
In cases where the sample size is inadequate, the use of an algorithmic model can allow for an effective evaluation of the risk of atherosclerosis in patients.
Methods:
We propose an oversampling technique called K-Means-Borderline-SMOTE (K-Means BS) and a classification algorithm named triple-weighted gcForest (TW-gcForest). Our proposed K-Means BS generates diverse synthetic samples by imposing strict constraints and avoids generating similar synthetic samples. TW-gcForest is mainly designed to address the problem of unfair forest weight allocation and sliding window weight allocation in standard gcForest. We perform numerical simulations on two datasets to demonstrate the robustness of the two methods.
Results:
Numerical simulations show that standard gcForest achieves the high-performance for atherosclerosis risk assessment on the K-Means BS synthetic dataset. However, TW-gcForest exhibits superior performance to the standard gcForest on the original dataset, as well as the SMOTE and K-Means BS synthetic datasets.
Conclusion:
Our approach can effectively improve accuracy, precision, recall, F1 score, and AUC compared with traditional algorithms.
Related Concept Videos
Comparing the Survival Analysis of Two or More Groups
One-Way ANOVA: Equal Sample Sizes
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
One-Way ANOVA: Unequal Sample Sizes
Survival Tree
Building a Survival Tree
Constructing a...
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Statistical Methods for Analyzing Epidemiological Data

