Related Experiment Video
Updated: Feb 11, 2026

Author Spotlight: Addressing Technical and Subjective Challenges in Measuring Classroom Attention
Published on: December 15, 2023
Evaluating Sociodemographic Biases in Artificial Intelligence-Based Glioblastoma Response Assessment Algorithms
Rachel S Lee1, Dominic LaBella2, Jikai Zhang3
1From the Duke University School of Medicine (R.S.L.), Durham, North Carolina Rachel.lee@duke.edu.
Background And Purpose:
Recent studies have demonstrated bias in various medical imaging artificial intelligence (AI) models, yet the factors underpinning these biases remain relatively unclear. This study evaluated potential sociodemographic biases in AI-based glioblastoma MRI segmentation models trained on data sets varying in size and demographic composition. We evaluated 4 nnUNet models with different training data sets: 1) the Federated Tumor Segmentation (FeTS) postoperative model trained on a large (>10,000 examinations) multinational, multi-institution data set; 2) the Brain Tumor Segmentation (BraTS) 2024 postoperative glioma model trained on a moderate size (>2000 examinations) multi-institution, North American data set; 3) a model trained on a small (>200 examinations), private, demographically homogeneous, single-institution data set; and 4) a model trained on an equally small (>200 examinations), but demographically heterogeneous data set.
Materials And Methods:
Models were evaluated for bias using an independent, manually corrected data set of 480 patients (mean age 52 ± 14) that was prospectively collected from a single high-volume academic brain tumor center. Automated FLAIR and enhancing tumor segmentations from the AI models were evaluated using Dice scores. Sociodemographic factors were collected and analyzed using beta regression to assess their influence on model performance.
Results:
The model trained exclusively on white, non-Hispanic men had the lowest overall Dice scores (0.943 for FLAIR, 0.909 for enhancement) and exhibited biases in age and smoking status. The BraTS model demonstrated the highest Dice scores (0.996 for FLAIR, 0.999 for enhancement) and had the least bias overall.
Conclusions:
Demographic bias was relatively low in glioblastoma MRI segmentation models. The model trained on the smallest and most homogeneous data set exhibited the most bias. Greater demographic heterogeneity even without increasing training data set size was associated with reduced bias. The BraTS model, trained on a moderate-sized cohort that included more diverse tumor types, performed better and demonstrated less bias than the FeTS model, despite the FeTS being trained on the largest data set.
Related Concept Videos
Confirmation Biases
Hindsight Biases
Bias
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
Intelligence
Trial and Error and Algorithm
Correspondence Bias

