照亮机器学习模型增强弹性的道路,抵御缺失标签的阴影
IEEE journal of biomedical and health informatics
|April 8, 2025
概括
这项研究引入了一种使用一类半监督异常检测 (OCSSAD) 的自动化方法,以过生物医学数据集中的 inliers. 这种方法显著提高了机器学习模型在医学和生命科学研究中的弹性.
科学领域:
- 生物医学数据分析
- 机器学习在医疗保健中的应用
- 生命科学中的数据质量
背景情况:
- 生物医学数据集容易受到初始污染 (缺失或错误的标签),损害了监督分类模型的敏感性.
- 虽然异常值经常被删除,但生物医学研究数据集中的内值在很大程度上没有得到解决.
- 现有的方法在各种生物医学数据类型中与初始污染的独特挑战作斗争.
研究的目的:
- 开发和评估一种自动化的方法来过生物医学培训数据集中的内置值.
- 提高在医学和生命科学中使用的机器学习模型的稳定性和可靠性.
- 解决临床和生物医学数据中早期污染的研究不足的问题.
主要方法:
- 升级一级半监督异常检测 (OCSSAD) 模型用于内置识别.
- 在六个数据库中对五种OCSSAD和两种集合方法进行基准测试,污染水平各不相同.
- 利用隔离森林来过训练集以提高模型的弹性.
主要成果:
- 在验证过程中,OCSSAD模型平均达到78±17%的马修斯相关系数 (MCC).
- 在先前过的数据上训练的机器学习模型显示,平均弹性从69±11%显著增加到95±1%.
- 提出的方法提高了各种机器学习模型的性能,包括神经网络和梯度增强方法.
结论:
- 开发的自动化方法有效地过了inliers,提高了机器学习模型在生物医学应用中的弹性.
- 解决先进污染对于提高机器学习在医学和生命科学中的准确性和可靠性至关重要.
- 这种方法为提高敏感研究领域的数据质量和模型性能提供了多功能解决方案.
相关概念视频
Improving Translational Accuracy
8.5K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
8.5K
Prediction Intervals
2.2K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.2K
Residuals and Least-Squares Property
7.2K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.2K
Distribution Reliability and Automation
94
Distribution reliability in electrical power systems is critical for ensuring an uninterrupted power supply to consumers at minimal cost. According to IEEE Standard Terms, reliability is the probability that a device will function without failure over a specified time period or amount of usage. For electric power distribution, this translates to maintaining continuous power supply and addressing customer concerns over power outages. Several indices, as defined by IEEE Standard 1366-2012, are...
94
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
34
Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
34


