稳定机器学习以获得可重现和可解释的结果:对特定学科的见解进行新型验证方法
Gideon Vos1, Liza van Eijk2, Zoltan Sarnyai3
1College of Science and Engineering, James Cook University, James Cook Dr, Townsville, 4811, QLD, Australia.
Computer methods and programs in biomedicine
|June 26, 2025
概括
本研究引入了一种新的验证方法,以稳定机器学习 (ML) 模型的性能和特征的重要性. 该方法提高了医学研究中群体和特定学科ML模型的可重现性和可解释性.
科学领域:
- 医学研究 医学研究
- 机器学习应用程序 机器学习应用程序
- 计算生物学是一种计算生物学.
背景情况:
- 机器学习 (ML) 增强了医学研究,但一般模型与个体的生物变异性作斗争.
- 专题特定的ML模型提供了精确性,但面临实际和财务挑战.
- 随机ML模型由于随机种子变异而表现出可重复性问题.
研究的目的:
- 引入一种新的验证方法,以提高ML模型的可解释性.
- 稳定预测性表现和特征的重要性在群体和特定主题层面.
- 在随机机器学习模型中解决可重现性挑战.
主要方法:
- 在九个不同的数据集上利用随机森林 (RF) 模型.
- 重复的实验,每人多达400次试验,随机种子变异.
- 聚合特征重要性排名,以确定一致重要的特征.
- 开发了特定于群体和主题的特征重要性集.
主要成果:
- 随机ML模型显示,由于随机种子,准确度和特征重要性存在显著的变化.
- 新的验证方法大大降低了功能排名的变化.
- 实现了一致的模型性能指标和对关键主题特定特征的强有力的识别.
- 提高了ML模型预测的可解释性和稳定性.
结论:
- 具体学科模型是有价值的,但往往是不切实际的.
- 拟议的验证技术提高了特征选择稳定性,预测准确性和可解释性.
- 确保可重现的准确度指标和可靠的特征排名用于随机ML模型.
- 增加ML模型的稳定性和临床适用性.
相关概念视频
Data Validation
259
Method validation is a crucial process in analytical chemistry designed to confirm that a given method consistently produces reliable and high-quality results. This process is essential when a method is applied to different sample matrices or when procedural modifications are made, ensuring that the results meet acceptable standards across various applications.
Key parameters for method validation include:
Key parameters for method validation include:
259
Improving Translational Accuracy
11.9K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.9K
Reliability and Validity
13.2K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
13.2K
Distribution Reliability and Automation
167
Distribution reliability in electrical power systems is critical for ensuring an uninterrupted power supply to consumers at minimal cost. According to IEEE Standard Terms, reliability is the probability that a device will function without failure over a specified time period or amount of usage. For electric power distribution, this translates to maintaining continuous power supply and addressing customer concerns over power outages. Several indices, as defined by IEEE Standard 1366-2012, are...
167
Survival Tree
166
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
166
Randomized Experiments
7.3K
The randomization process involves assigning study participants randomly to experimental or control groups based on their probability of being equally assigned. Randomization is meant to eliminate selection bias and balance known and unknown confounding factors so that the control group is similar to the treatment group as much as possible. A computer program and a random number generator can be used to assign participants to groups in a way that minimizes bias.
Simple randomization
Simple...
Simple randomization
Simple...
7.3K


