从CBC到清晰:在不平衡的数据集中可解释地检测β-血病载体
Saim Chishti1,2, Faryal Nosheen2, Joddat Fatima1,3
1Center of Excellence in AI (CoE-AI), Bahria University, Islamabad, Pakistan.
PloS one
|September 18, 2025
概括
一个新的基于主导的粗略设置方法模型准确地识别了使用完整血清数据而不是类正常化的β-thalassemia携带者. 这种机器学习模型在训练和未见数据上都表现出强的表现,提供了一个有前途的诊断工具.
科学领域:
- 医学诊断 医学诊断 医学诊断
- 计算生物学 计算生物学
- 医疗保健中的机器学习
背景情况:
- thalassemia是一种普遍存在的遗传性血液疾病,特别是在东南亚.
- 机器学习 (ML) 模型越来越多地用于疾病预测和分类.
- 现有的用于检测β-thalassemia载体的ML模型通常需要数据规范化,改变原始数据集.
研究的目的:
- 开发和评估一种新的ML模型,用于beta-thalassemia携带者分类.
- 在没有类规范化的情况下评估模型的性能.
- 将拟议的模型与现有方法进行比较.
主要方法:
- 提出了一个基于主导的粗略设置方法 (DRSA) 模型.
- 该DRSA模型是使用完整血清 (CBC) 数据进行训练和测试的.
- 分类在没有平衡数据集类 (正常,异常) 的情况下进行.
主要成果:
- 拟议的DRSA模型在分类患者方面实现了91%的准确性.
- 该模型表现出强大的概括性,在未见的数据上准确率为89%.
- 性能与使用数据规范化的现有ML方法相当或优于现有ML方法.
结论:
- DRSA模型提供了一种有效的方法,用于使用CBC数据检测β-thalassemia载体.
- 在没有数据规范化的情况下进行分类,可以保持原始数据集的完整性.
- 这种方法提供了一个强大的和准确的替代方案,以选的血病.
相关概念视频
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...


