在医疗保健数据集中检测,描述和缓解隐含和明确的种族偏见:算法开发和验证研究
Faris Gulamali1, Ashwin Shreekant Sawant1, Lora Liharska1
1Icahn School of Medicine at Mount Sinai, 1468 Madison Avenue, New York, NY, 10029, United States, 1 2122416500.
Journal of medical Internet research
|September 4, 2025
概括
通过指导数据收集和重新标记, 一种名为AEquity的新指标有效地减少了医疗数据中的算法偏差. 这种方法在各种数据集和算法中表现优于现有方法,提高了人工智能诊断的公平性.
科学领域:
- 医疗保健中的人工智能
- 算法公平性
- 数据科学
背景情况:
- 越来越多的人采用医疗保健算法,
- 现有的偏差缓解方法侧重于模型修改,对数据级干预的努力有限.
- 医疗保健数据集容易产生偏差,这可能会影响诊断和预后算法的性能.
研究的目的:
- 介绍AEquity,一个使用学习曲线近似的新型指标,以识别和减轻医疗保健数据中的偏差.
- 展示AEquity在指导数据集收集和重新标记方面的有效性,以提高算法公平性.
- 评估AEquity在各种数据集,算法和公平性指标中的稳定性.
主要方法:
- 开发基于学习曲线近似的AEquity指标,用于偏差检测和缓解.
- 应用于胸部X射线数据集,医疗保健成本利用数据和国家健康和营养检查调查 (NHANES).
- 与先进的方法比较,例如平衡的经验风险最小化和校准.
主要成果:
- 以公平为导向的数据采集使胸部X- 射线图的偏差降低了29% - 96. 5% (AUC).
- 在多个公平度指标 (例如,FNR减少33.3%) 中观察到对交叉群体 (医疗补助的黑人患者) 的显著偏差减少.
- AEquity的表现优于平衡的经验风险最小化和校准,在各种人工智能模型 (CNN,变压器等) 中表现强. ) 的情况.
结论:
- 在医疗保健数据层面减轻算法偏差的一个强大而有效的工具.
- 该指标在不同数据集,人口群体和机器学习架构中具有广泛的适用性.
- 提供一个有希望的以数据为中心的方法来提高医疗保健AI的公平性.
更多相关视频
07:35Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
7.6K
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
565
相关概念视频
Stereotypes, Prejudice, and Discrimination
91.4K
Humans are very diverse and although we share many similarities, we also have many differences. The social groups we belong to help form our identities (Tajfel, 1974). These differences may be difficult for some people to reconcile, which may lead to prejudice toward people who are different. Prejudice is a negative attitude and feeling toward an individual based solely on one’s membership in a particular social group (Allport, 1954; Brown, 2010). Prejudice is common against people who...
91.4K
Bias in Epidemiological Studies
666
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:
666
Bias
4.8K
Bias refers to any tendency that prevents a question from being considered unprejudiced. In research, bias occurs when one outcome or answer is selected or encouraged over others in sampling or testing. Bias can occur during any research phase, including study design, data collection, analysis, and publication.
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
4.8K
Data Validation
5.3K
Data validation is an essential part of a comprehensive assessment. Validation is confirming or verifying and opening the door to gathering more assessment data as it clarifies vague or unclear data. The process of checking and verifying the collected information is called data validation. The primary purpose of data validation is to ensure data is as free from error, bias, and misinterpretation as possible.
Nursing assessment guides are generally based on holistic models rather than medical...
Nursing assessment guides are generally based on holistic models rather than medical...
5.3K
Blind Procedures
12.1K
Ideally, the people who observe and record the children’s behavior are unaware of who was assigned to the experimental or control group, in order to control for experimenter bias. Experimenter bias refers to the possibility that a researcher’s expectations might skew the results of the study. Remember, conducting an experiment requires a lot of planning, and the people involved in the research project have a vested interest in supporting their hypotheses. If the observers knew which...
12.1K
Surveys
15.3K
Often, psychologists develop surveys as a means of gathering data. Surveys are lists of questions to be answered by research participants, and can be delivered as paper-and-pencil questionnaires, administered electronically, or conducted verbally. Generally, the survey itself can be completed in a short time, and the ease of administering a survey makes it easy to collect data from a large number of people.
15.3K
