解决健康数据的不平衡:使用深度学习的合成少数群体过量抽样
Alex X Wang1, Viet-Tuan Le2, Hau Nguyen Trung2
1School of Mathematics and Statistics, Victoria University of Wellington, Kelburn Parade, Wellington 6012, New Zealand.
Computers in biology and medicine
|February 21, 2025
概括
本研究引入了一种先进的深度学习方法,以解决医疗保健数据中的阶级不平衡,通过生成合成阳性样本和完善多数类数据来提高机器学习模型公平性和患者安全性.
科学领域:
- 机器学习 机器学习
- 人工智能的人工智能
- 医疗保健信息学 医疗保健信息学
背景情况:
- 医疗保健数据中的阶级不平衡导致有偏见的机器学习模型,损害了患者安全和医疗保健服务.
- 像SMOTE这样的传统过量采样方法在处理复杂数据,异质类型和多类场景方面存在局限性.
研究的目的:
- 提出一种新的深度学习方法来解决医疗保健数据集中的阶级不平衡问题.
- 提高机器学习模型在临床应用中的性能和公平性.
主要方法:
- 一个辅助引导的有条件变量自编码器 (ACVAE) 具有对比学习被开发用于合成数据生成.
- 采用了一种组合技术,将ACVAE用于过量抽样的阳性病例和编辑的中枢神经位移最近邻居 (ECDNN) 用于多数类低样本.
主要成果:
- 对12个健康数据集的实验证明了拟议的ACVAE-ECDNN组合方法的有效性.
- 与传统的过量采样技术相比,该方法在各种指标上显著改善了模型性能.
结论:
- 基于深度学习的合成过量采样为医疗保健数据中的类不平衡提供了强大的解决方案.
- 拟议的方法提高了数据集的平衡性和信息性,从而为医疗保健提供了更可靠的机器学习模型.
更多相关视频
06:55Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
14.3K
12:18A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
7.4K
相关概念视频
Sampling Methods: Overview
269
A sample refers to a smaller subset representative of a larger population. In analytical chemistry, studying or analyzing an entire population is often impractical or impossible. Therefore, samples are used to draw inferences and generalize the whole population. The sampling method selects individuals or items from a population to create a sample. Standard sampling methods include random, judgemental, systematic, stratified, and cluster sampling.
In analytical chemistry, the choice of...
In analytical chemistry, the choice of...
269
Improving Translational Accuracy
8.5K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
8.5K
Downsampling
126
When considering a sampled sequence with zero values between sampling instants, one can replace it by taking every N-th value of the sequence. At these integer multiples of N, the original and sampled sequences coincide. This process, known as decimation, involves extracting every N-th sample from a sequence, thereby creating a more efficient sequence.
The Fourier transform of the decimated sequence reveals a combination of scaled and shifted versions of the original spectrum. This...
The Fourier transform of the decimated sequence reveals a combination of scaled and shifted versions of the original spectrum. This...
126
Systematic Sampling Method
9.9K
Sampling is a technique to select a portion (or subset) of the larger population and study that portion (the sample) to gain information about the population. Data are the result of sampling from a population. The sampling method ensures that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
Systematic sampling is one of the simplest methods...
Systematic sampling is one of the simplest methods...
9.9K
Bias in Epidemiological Studies
131
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:
131
Random Sampling Method
10.9K
Sampling is a technique to select a portion (or subset) of the larger population and study that portion (the sample) to gain information about the population. Data are the result of sampling from a population. The sampling method ensures that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest. Among the various sampling methods used by...
10.9K
