相关实验视频
Updated: Jun 27, 2025

13:51
Cross-Modal Multivariate Pattern Analysis
Published on: November 9, 2011
19.9K
对于多感官跨数据集的心脏病预测中的性能差异缓解.
Mahmudul Hasan1,2, Md Abdus Sahid1, Md Palash Uddin1,2
1Department of Computer Science and Engineering, Hajee Mohammad Danesh Science and Technology University, Dinajpur, Bangladesh.
PeerJ. Computer science
|April 25, 2024
概括
这项研究解决了用于心脏病预测的机器学习模型中的数据集间差异问题. 有效的预处理显著提高模型性能,即使在不同的数据集.
科学领域:
- 心脏病学 心脏病学
- 计算机科学 计算机科学
- 生物医学信息学 生物医学信息学
背景情况:
- 心脏病是全球死亡的主要原因,需要改进早期预测方法.
- 现有的心脏病机器学习 (ML) 模型通常在未见的数据集上表现不佳,原因是数据集间的差异问题.
- 当在一个数据集上训练的模型在另一个不同的数据集上进行测试时,这种差异就会出现.
研究的目的:
- 为了减轻心脏病预测中的数据集间差异问题,使用ML.
- 系统地评估预处理技术对多个数据集模型性能的影响.
- 在跨数据集情景中确定特征选择和分类的最佳策略.
主要方法:
- 利用了五个不同的心脏病数据集,探索了所有训练和测试组合.
- 实施了全面的预处理管道,包括SMOTE-Tomek用于不平衡处理,随机森林 (RF) 用于特征选择,以及主要组件分析 (PCA) 用于特征提取.
- 包含缺失值赋值 (RF回归),日志转换,异常值去除,规范化和数据平衡.
- 评估了八种不同的分类器:支持向量机,K-最近邻居,决策树,射频,极端梯度提升,高斯天真贝斯,物流回归和多层感知器.
主要成果:
- 随机森林 (RF) 在数据集内部和数据集间的设置中,在特征选择和分类方面表现出卓越的表现.
- 在某些配置中,RF可达到高达100%的准确性,在跨数据集特征选择过程中达到96%的准确性.
- 拟议的预处理管道显著改善了ML模型的性能,减少了数据集间的差异,而不需要复杂的模型开发.
结论:
- 有效的数据预处理对于克服心脏病预测模型中的数据集间差异至关重要.
- 随机森林是一个非常有效的工具,用于特征选择和分类在这个领域.
- 解决数据集间的差异,通过结合不同的数据集,为创建强大的,可泛化的ML模型打开了道路.
相关概念视频
Sensitivity, Specificity, and Predicted Value
298
In healthcare diagnostics, laboratory tests play a crucial role in identifying and diagnosing a wide range of medical conditions. However, interpreting test results is not always straightforward. An abnormal test result does not always confirm the presence of a disease, just as a normal result does not guarantee its absence. To assess the reliability of these diagnostic tools, healthcare practitioners rely on two key statistical indicators: sensitivity and specificity.
Sensitivity is the...
Sensitivity is the...
298
Improving Translational Accuracy
10.2K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
10.2K

