关于基因病理学数据集中的潜在偏差因素的调查
Farnaz Kheiri1, Shahryar Rahnamayan2, Masoud Makrehchi3
1Department of Electrical, Computer and Software Engineering, Ontario Tech University, Oshawa, Canada. farnaz.kheiri@ontariotechu.net.
Scientific reports
|April 2, 2025
概括
深度神经网络 (DNN) 在癌症基因组图谱 (TCGA) 数据集中显示偏差. 在TCGA数据上训练的模型识别采集部位特征,而不仅仅是癌症特征,导致膨胀的性能指标.
科学领域:
- 医学成像分析分析 医学成像分析
- 计算病理学计算病理学
- 人工智能在瘤学中的应用
背景情况:
- 深度神经网络 (DNN) 是数字病理学的强大工具,用于疾病诊断和预后.
- 广泛用于训练DNN的癌症基因组图谱 (TCGA) 数据集可能包含特定站点的偏差.
- 之前的研究表明,DNN可以在分类TCGA数据采集站点时实现高精度,这表明它依赖于非组织学特征.
研究的目的:
- 在TCGA数据上训练的DNN中调查特定站点偏差的原因.
- 分析DNN在识别采集部位模式与癌症模式方面的表现.
- 了解数据集偏差如何影响数字病理学深度学习模型的可靠性.
主要方法:
- 分析了两个DNN模型:KimiaNet和一个自我训练的EfficientNet.
- 使用平衡精度度指标对模型性能进行评估.
- 测试模型对其分类数据采集地点与癌症类型的能力.
主要成果:
- 在TCGA数据集中,DNN显示出了区分不同数据采集站点的显著能力.
- 这种特定位置的识别能力即使在模型用于癌症模式分析时也被观察到.
- 站点分类的平衡精度出乎意料地高,表明学习特征存在偏差.
结论:
- 这些发现突出了TCGA数据集偏差的一个关键问题,影响了DNN的概括性.
- 无意中,DNN可能会学习特定网站的文物,而不是真正的病理特征.
- 需要进一步的研究来缓解偏见,并确保AI模型在数字病理学中的临床有效性.
相关概念视频
Bias
Bias refers to any tendency that prevents a question from being considered unprejudiced. In research, bias occurs when one outcome or answer is selected or encouraged over others in sampling or testing. Bias can occur during any research phase, including study design, data collection, analysis, and publication.
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
Bias in Epidemiological Studies
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:


