消除偏差:通过杂的群体测量来改善差异评估的方法
Solvejg Wastvedt1, Joshua Snoke2, Denis Agniel3
1Department of Statistics & Data Science, NORC at the University of Chicago, Chicago, IL 60603, United States.
Biometrics
|December 28, 2024
概括
新的统计方法解决了当种族数据缺失时,医疗保健算法中的种族差异. 这种方法量化了从归算的种族概率的偏差,使得更公平的临床决策支持.
科学领域:
- 医疗信息学 医疗信息学
- 生物统计学 生物统计学
- 健康 公平 卫生 公平
背景情况:
- 临床决策支持算法可能会加剧医疗保健中的种族和民族差异.
- 缺少或质量差的种族/种族数据阻碍了这些算法的准确偏差评估.
研究的目的:
- 开发新的统计方法来评估算法偏差,使用概率性种族/种族成员身份.
- 量化由归算组概率中的错误引起的统计偏差.
- 为决策者提供工具,以评估临床决策支持算法的差异.
主要方法:
- 用于算法性能评估的种族/民族群体成员的利用概率.
- 开发了一种灵敏度分析方法,以在计算概率的不同误差水平下估计统计偏差.
- 在常见公平度指标上的统计偏差的理论界限.
- 应用方法用于一个案例研究,使用假定的种族/种族数据来研究骨质疏松症治疗差异.
主要成果:
- 展示了一种评估算法性能和量化偏差的方法,即使具有不完整的种族/种族数据.
- 提出了一个灵敏度分析框架,以了解归算错误对偏差的影响.
- 根据特定的公平性指标,建立了偏差估计的理论保证.
- 在使用归算数据的骨质疏松症治疗算法中说明了潜在的差异.
结论:
- 当种族/种族数据不完美时,新的统计方法可以对算法偏见和健康差异进行可靠的评估.
- 灵敏度分析方法允许理解一系列潜在的差异.
- 有关在临床环境中实施机器学习的知情决策得到了促进,促进了健康公平.
相关概念视频
Bias in Epidemiological Studies
88
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:
88
Strategies for Assessing and Addressing Confounding
62
Confounding is a critical issue in epidemiological studies, often leading to misleading conclusions about associations between exposures and outcomes. It occurs when the relationship between the exposure and the outcome is mixed with the effects of other factors that influence the outcome. Given that, addressing confounding is of high importance for drawing accurate inferences in research.
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
62
Bias
3.6K
Bias refers to any tendency that prevents a question from being considered unprejudiced. In research, bias occurs when one outcome or answer is selected or encouraged over others in sampling or testing. Bias can occur during any research phase, including study design, data collection, analysis, and publication.
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
3.6K
Empirical Method to Interpret Standard Deviation
5.1K
The empirical rule, also known as the three-sigma rule, allows a statistician to interpret the standard deviation in a normally distributed dataset. The rule states that 68% of the data lies within one standard deviation from the mean, 95% lies within two standard deviations from the mean, and 99.7% lies within three standard deviations from the mean. Additionally, this rule is also called the 68-95-99.7 rule.
This rule is used widely in statistics to calculate the proportion of data values...
This rule is used widely in statistics to calculate the proportion of data values...
5.1K
Estimating Population Mean with Unknown Standard Deviation
7.5K
In practice, we rarely know the population standard deviation. In the past, when the sample size was large, this did not present a problem to statisticians. They used the sample standard deviation s as an estimate for σ and proceeded as before to calculate a confidence interval with close enough results. However, statisticians ran into problems when the sample size was small. A small sample size caused inaccuracies in the confidence interval.
William S. Gosset (1876–1937) of the...
William S. Gosset (1876–1937) of the...
7.5K
Testing a Claim about Mean: Known Population SD
2.7K
A complete procedure of testing the hypothesis about a population mean is explained here.
Estimating a population mean requires the samples to be distributed normally. The data should be collected from the randomly selected samples having no sampling bias. The sample size needed to be higher than 30, and most importantly, the population standard deviation should be already known.
In most realistic situations, the population standard deviation is often unknown, but in rare circumstances, when it...
Estimating a population mean requires the samples to be distributed normally. The data should be collected from the randomly selected samples having no sampling bias. The sample size needed to be higher than 30, and most importantly, the population standard deviation should be already known.
In most realistic situations, the population standard deviation is often unknown, but in rare circumstances, when it...
2.7K


