生成模型在分配转移下提高了医疗分类器的公平性
Ira Ktena1, Olivia Wiles2, Isabela Albuquerque3
1Google DeepMind, London, UK. iraktena@google.com.
Nature medicine
|April 10, 2024
概括
生成型人工智能,特别是扩散模型,可以创建合成医疗数据,以提高机器学习模型的公平性和稳定性. 这种方法提高了代表性不足的群体的诊断准确性,特别是在现实世界,分布之外的场景中.
科学领域:
- 医疗成像医学成像
- 机器学习 机器学习
- 生成型的人工智能
背景情况:
- 域泛化是医疗保健机器学习的一个关键挑战,由于开发和部署之间的数据差异,模型表现不佳.
- 培训数据中特定群体或条件的代表性不足导致模型性能和公平性降低.
- 获得和标记广泛的临床数据往往是不可行的,因为成本和条件的稀有性.
研究的目的:
- 调查生成人工智能,特别是扩散模型的使用,以创建合成数据,以提高机器学习模型的稳定性和公平性.
- 解决医疗机器学习中对标签效率数据增强的未满足需求.
- 在各种医学成像任务中评估学习增强的有效性.
主要方法:
- 利用扩散模型以标签有效的方式学习现实的数据增强.
- 丰富的培训数据集与合成示例来解决代表性不足的条件和子组.
- 在三个不同的医疗成像数据集上评估模型性能和公平性:组织病理学,胸部X射线和皮肤学图像.
主要成果:
- 由扩散模型产生的学习增强在所有测试的医学成像任务中显著提高了模型的稳定性.
- 这种方法提高了统计学公平性,特别是改善了代表性不足的群体的诊断准确性,特别是在分布之外的环境中.
- 合成数据增强在组织病理学,胸部X射线和皮肤病学图像分析中被证明是有效的.
结论:
- 通过扩散模型,生成人工智能提供了一种可引导和标签效率高的方法,用于医疗机器学习创建合成数据.
- 合成数据增强增强了模型的稳定性和公平性,这对于现实世界的临床部署至关重要.
- 这种方法在缓解医疗保健人工智能中数据代表性不足所造成的绩效差距方面表现有希望.
相关概念视频
Types of Skewness
11.6K
If the frequency distribution of a data set is more inclined towards smaller or larger values, the distribution is said to be skewed. If data values are skewed to the right, then the distribution is called positively skewed. Conversely, if the plot is skewed to the left, the distribution is called negatively skewed.
For instance, in the middle of a pandemic, the geographical distribution of vaccine coverage may be positively skewed towards populations in the global north countries. However,...
For instance, in the middle of a pandemic, the geographical distribution of vaccine coverage may be positively skewed towards populations in the global north countries. However,...
11.6K
Improving Translational Accuracy
10.3K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
10.3K
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
69
Pharmacokinetic models are mathematical constructs that represent and predict the time course of drug concentrations in the body, providing meaningful pharmacokinetic parameters. These models are categorized into compartment, physiological, and distributed parameter models.
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
69
Bias in Epidemiological Studies
254
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:
254
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Distributions to Estimate Population Parameter
4.1K
The accurate values of population parameters such as population proportion, population mean, and population standard deviation (or variance) are usually unknown. These are fixed values that can only be estimated from the data collected from the samples. The estimates of each of these parameters are sample proportion, the sample mean, and sample standard deviation (or variance). To obtain the values of these sample statistics, data are required that have particular distribution and central...
4.1K


