超越预测准确性:在生物医学机器学习中统计验证特征的重要性.
Souichi Oka1, Nobuko Inoue1, Yoshiyasu Takefuji2
1Science Park Corporation, 3-24-9 Iriya-Nishi Zama-shi, Kanagawa 252-0029, Japan.
Computer methods and programs in biomedicine
|September 28, 2025
概括
医学机器学习 (ML) 的高预测准确度并不能保证对潜在的生理机制的准确识别. 使用ML和统计方法的综合方法对于对呼吸道恶化等疾病驱动因素的可靠见解至关重要.
科学领域:
- 医学机器学习 医学机器学习
- 计算生物学 计算生物学
- 呼吸系统医学 呼吸系统医学
背景情况:
- 医疗机器学习 (ML) 模型可以实现高预测准确度的任务,如呼吸道恶化分类.
- 然而,高预测性能并不等同于发现真正的生理机制.
- 从诸如随机森林 (RF) 这样的复杂模型中解释特征的重要性可能会产生偏见,可能会误解潜在的疾病驱动因素.
研究的目的:
- 解决医学机器学习中关于特征重要性的解释性挑战.
- 要突出仅仅依赖于机械洞察力的预测准确性的局限性.
- 倡导一种协同方法,将ML与互补的统计方法相结合,以实现可靠的科学发现.
主要方法:
- 分析多特征融合模型 (KNN,SVM,RF,BiLSTM) 用于呼吸道恶化预测.
- 批评依赖模型的解释方法,如SHapley增量解释 (SHAP) 对于特征的重要性.
- 建议将公正的统计方法 (例如,非参数相关性,相互信息) 与ML集成.
主要成果:
- 在ML模型中的高预测精度不能可靠地表明对生理机制的真正特征重要性.
- 特定模型的解释技术,如SHAP在RF上,可以在特征重要性排名中引入偏差.
- 不同的ML模型通常会产生不同的特征重要性排名,使得在没有基本真相的情况下复杂化验证.
结论:
- 仅仅依靠预测指标和依赖模型的解释,就有可能误解复杂的医学现象.
- 一个强大的分析策略,将ML的预测能力与补充的统计方法相结合,是必不可少的.
- 这种协同方法确保了对呼吸道恶化等疾病的真正驱动因素的可靠科学见解.
相关概念视频
Sensitivity, Specificity, and Predicted Value
1.2K
In healthcare diagnostics, laboratory tests play a crucial role in identifying and diagnosing a wide range of medical conditions. However, interpreting test results is not always straightforward. An abnormal test result does not always confirm the presence of a disease, just as a normal result does not guarantee its absence. To assess the reliability of these diagnostic tools, healthcare practitioners rely on two key statistical indicators: sensitivity and specificity.
Sensitivity is the...
Sensitivity is the...
1.2K
Biostatistics: Overview
732
Biostatistics plays a crucial role in understanding and analyzing data in healthcare and biology. Biostatisticians conduct experiments, gather evidence, and draw meaningful conclusions using statistical methods and techniques. Different variables form the foundation of biostatistical analysis, allowing researchers to understand and interpret data effectively. These variables are classified into different types, each serving a specific purpose in statistical analysis.
Discrete variables are...
Discrete variables are...
732
Improving Translational Accuracy
14.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.1K
Improving Translational Accuracy
3.5K
3.5K
Receiver Operating Characteristic Plot
471
A ROC (Receiver Operating Characteristic) plot is a graphical tool used to assess the performance of a binary classification model by illustrating the trade-off between sensitivity (true positive rate) and specificity (false positive rate). By plotting sensitivity against 1 - specificity across various threshold settings, the ROC curve shows how well the model distinguishes between classes, with a curve closer to the top-left corner indicating a more accurate model. The area under the ROC curve...
471
Statistical Significance
21.0K
Once data is collected from both the experimental and the control groups, a statistical analysis is conducted to find out if there are meaningful differences between the two groups. A statistical analysis determines how likely any difference found is due to chance (and thus not meaningful). In psychology, group differences are considered meaningful, or significant, if the odds that these differences occurred by chance alone are 5 percent or less. Stated another way, if we repeated this...
21.0K


